Open source, for your terminal

A coding agent that
turns smart context
into effective action.

Caudra brings steerable subagents, built-in token savings, a full workbench, and a session history you can always return to. It all ships in one native binary, so there are no plugins to assemble.

Leave a session working toward a goal while you are away. When you are at the keyboard, guide any subagent, ask a side question, or open the code beside the conversation.

curl -fsSL https://caudra.ai/install.sh | sh
IllustrationThe order in which work starts during one streamed response. Bars show order, not measured time.Work starts while the model writes

Measured by the maintainer in daily use

  • 20B+tokens a month through Caudra
  • <2 GBfor over a month of complete local history
  • <1 sto load or save a 10k-turn session, subagents included
  • ~97%Anthropic prompt-cache hit rate, and about 94% on OpenAI
How these were measured

Why Caudra exists

The agent I wanted
for my own working day.

I spend most of my working day with coding agents, and I have tried many of them. Each had ideas I liked. I wanted those ideas together, along with a few new techniques, in one standalone binary I could bring into any environment.

Caudra is that binary, and I push more than 20 billion tokens a month through it. With another agent, my local history grew by more than 1 GB a day. Caudra keeps over a month of my complete history in less than 2 GB, and a 10k-turn session still loads in under a second.

Thorsten BornSoftware engineer and data architect, maintainer of Caudra
Local history elsewhere
>1 GB a day
Local history in Caudra
<2 GB for over a month

Built in and released together.

The core team develops, tests, and releases every capability together. You install one binary and get the whole agent, with no third-party plugins to vet or keep up to date. Lua extensions remain available as an experimental option.

Open-source credits

Caudra builds on ideas from these projects, with thanks.

MethodThese are the maintainer's own measurements from daily use, not benchmarks. Timings were measured by hand. Results depend on models, projects, and hardware.

Why the name

Caudra comes
from caudate.

Caudra, pronounced KAW-druh, is named after the caudate nucleus. This part of the brain belongs to the circuits that connect evidence and goals to action. Studies link it to learning which actions lead to which outcomes.

The name describes how Caudra works. It turns the context of your task into an action, then reads the outcome before it chooses the next one.

The brain is only the inspiration for the name. Caudra is software, and it works toward the goal you set.

Sources: Grahn, Parkinson and Owen, 2008, Lau and Glimcher, 2007, Doi et al., 2020.

  1. 01Context and intent
  2. 02Evaluate evidence
  3. 03Select an action
  4. 04Execute with tools
  5. 05Observe the outcome
Each outcome becomes evidence for the next step.

01Steer anything, any time

Guide any agent
while it works.

The input stays open while Caudra works. Queue the next task, add guidance to the current run, or replace it. Open any running subagent and guide it directly.

  1. EnterNextWaits for the current run and its background work.
  2. Ctrl+X gGuideJoins the run before its next model request.
  3. Ctrl+X xReplaceStops the run and starts your instruction.
  4. /tasksGuide a subagentReaches that subagent at its next turn boundary.
IllustrationWhere your input goes while work runs.Guide a running subagent
  1. ?The agent asksA question form waits for your choice.
  2. F2Ask /btwClarify the choices in a side thread.
  3. EscReturnThe form is unchanged, and the side thread stays out of history.
IllustrationA side question while the agent waits for your answer.Clarify a question with /btw
  • Next, Guide, and Replace

    Enter queues a prompt for after the current run. Ctrl+X g adds guidance before the next model request, and Ctrl+X x stops the run and starts yours. Queued prompts stay editable.

  • Guide a subagent directly

    Open a running subagent from /tasks and type. Your guidance stays visible until the subagent reads it at its next turn boundary.

  • Read a brief as it is written

    A subagent's chat opens while the model is still writing its instructions, so you can read the task before it starts.

  • Ask /btw on the side

    Ask about the conversation without adding to its history. When the main agent asks you a question, F2 opens /btw to clarify the choices, and Esc returns to the unchanged form.

02While you sleep

Set the finish line.
Wake to results.

/goal keeps a session working until a separate evaluator finds evidence that your condition is met. Background work reports back without polling, and your permission rules decide what runs while you are away.

  1. /goalSet the conditionFor example: tests pass and clippy is clean.
  2. 1WorkThe turn ends and background results settle.
  3. 2EvaluateA separate call looks for evidence in the transcript.
  4. ↺ContinueAn unmet goal starts another turn, up to the limit.
  5. ✓DoneA met goal clears itself.
IllustrationThe loop that /goal runs after each work turn.Work toward a goal
  • Goals that check evidence

    After each work turn, a separate model call with no tools reads the transcript. An unmet goal starts another turn. A met goal clears itself.

  • Background work wakes the agent

    Tasks and shell jobs can keep running in the background. Their reports start the next step at a safe boundary, with no polling.

  • Permissions that read the command

    Shell chains and pipelines split into separate commands for approval. Approve once, for the conversation, the project, or every project, and accept patterns Caudra suggests from repeated approvals.

  • Plan first, then hear back

    Caudra opens in Plan mode, where the agent edits only its plan and asks before commands it cannot prove read-only. Notifications tell you when a turn finishes or a prompt needs your answer.

Unattended work needs a running session, for example in tmux or Herdr. Closing it cancels background work. A goal allows 16 automatic continuations by default, and work pauses when it needs your answer. /goal looks for evidence in the transcript. It does not prove the work is correct.

Completion goals

03Built-in token savings

Token savings,
built in.

Every turn re-sends the conversation, so a noisy result costs tokens again on every later turn until compaction. Caudra keeps results small and round-trips few, starting with RTK-style shell output filtering that is on by default.

  1. $RunYou watch the raw output while the command runs.
  2. filterTrimCommand-aware rules remove routine noise from the completed output.
  3. →ReadThe model receives the filtered result.
  4. rawCompareSwitch between filtered and raw views in the transcript.
IllustrationHow shell output is filtered before the model reads it.Filter shell output
  • RTK-style output filtering

    Command-aware rules, most of them from RTK, trim build and test noise before the model reads the output, and a filtered result is always smaller than the raw one. A failure is never reported as a success, and a failing command keeps its first and last lines. Progress bars collapse to one row, and you can switch between filtered and raw views.

  • Structure before content

    file_index returns a file outline with signatures and line numbers. The code_* tools rank symbols and show callers, impact, and the tests they can trace to a change. They parse the source on the fly, so there is no index to build or maintain and nothing to configure.

  • Fewer, smaller turns

    Oversized results reach the model as a bounded head and tail, and tool_output searches the rest. batch runs independent calls in one turn, and subagents keep their exploration out of the main context.

  • Requests shaped for caching

    Caudra sends a per-conversation cache key where providers accept one, and /usage scores the cache hit rate for every model.

04A workbench beside the agent

Open the code
without leaving.

Ctrl+X w brings up a file explorer, tabbed editor, project search, and source control while the session keeps running. It lives in the terminal, so it comes along over SSH.

Workbench layout: a sidebar with files, Git, and search views on the left, editor tabs and a buffer on the right, and a status row that offers Ctrl+X Enter to send a reference.

IllustrationThe workbench layout, taken from its documentation.The workbench beside a session
  1. Ctrl+X rOpenReview the last reply, or any message from its menu.
  2. Shift+↓SelectTake the rows that need work, or drag across them.
  3. EnterNoteWrite a note on the selection.
  4. Ctrl+SSendEvery note lands in the prompt as one block.
IllustrationThe passage review keys.Review passages of a reply
  • Read the actual change

    Inspect diffs, browse the commit graph, and stage, unstage, or discard per file or folder.

  • Point at exact lines

    Ctrl+X Enter sends the file, line, or selection to the composer as a mention such as @src/api.ts:L10-L20. Caudra puts those lines in the request.

  • Review passages of a reply

    Mark the parts of an answer that need work, add a note to each, and send every note back as one prompt.

  • Replies rendered for engineers

    Tables, highlighted code, Unicode maths, and Mermaid flowcharts render in the terminal. Copying a passage gives back its Markdown source.

05Never lose the thread

Go back to any
point in the work.

Every session keeps its full history, subagent transcripts included, in compact local storage. Revert the conversation, the files, or both from a message, and preview the file changes first.

  1. ⋮Open the menuBeside any message in the main transcript.
  2. 1PreviewRevert files counts what would be created, replaced, or deleted.
  3. 2ApplyRepeat the action. A conflict found before writing aborts the whole revert.
  4. ↶UnrevertPut the files and the conversation head back.
IllustrationA file revert with its preview.Revert files from a message
  • Large sessions open fast

    Payloads are compressed, appending saves write only new rows, and the transcript lays out only what is on screen. In the maintainer's use, a 10k-turn session loads in under a second.

  • Revert with a preview

    Each tool call that may change files is recorded before and after it runs, so a revert touches only those files. The first Revert files shows what would change, and repeating it applies the revert.

  • Memory that outlasts the session

    The agent keeps tagged notes on project gotchas and decisions outside your repository. Requests carry only the tags until a note is needed, and /memory lets you read, edit, or delete them.

  • Fork and branch out

    Fork from a message to try another approach. /worktree new moves a session into a fresh Git worktree with its conversation and plan.

File revert cannot undo external side effects such as running processes, databases, network calls, or Git branch state.

Worktrees and agent status integrate with Herdr.

Sessions and revert

06Automatic steering

Keep any model
on task.

Smaller local models such as Qwen3.8-27B need more help to finish a task. Caudra nudges the model when a turn stalls, gets cut off, loops, or stops after announcing work. The rules react to what a reply did, so the same defaults work well with flagship models.

  1. 1Stop earlyThe reply ends on “I will run the tests now” and calls no tool.
  2. 2NudgeThe abandoned_turn rule asks the model to do that work now.
  3. 3Read the nudgeA dim row in the transcript holds the exact text the model received.
  4. ↺ContinueThe next reply can call its tools. By default, a third announcement in a row ends the turn as written.
IllustrationWhat happens when a reply announces work and calls no tool.Nudge a turn that stopped early
  • Nudges for stalled turns

    After an empty or cut-off reply, Caudra asks the model to continue. After “I will run the tests now” with no tool call, it asks for the work itself. The third identical tool call in a row is refused before it runs.

  • Hints when work goes in circles

    When the model repeats a tool cycle or an answer, or keeps making failed calls, a hint asks it to reconsider its approach. Hints stop at four per run by default and never reopen a finished answer.

  • Sensible defaults, tuned per model

    Every rule is on by default, and budgets cap how often each one fires. Change a threshold, a budget, or the wording of a nudge for all models, or only for one exact provider/model-id.

  • Every nudge in the transcript

    Each nudge appears as a dim row. Click it to read the exact text the model received.

In daily use, the abandoned_turn rule matched 35 of 386 final Qwen3.8-27B replies and none of 2,185 from Claude and GPT models, as measured by the maintainer. A nudge is a message to the model. It cannot run or approve a tool, and every real tool call still passes through validation and your permission rules.

Configure automatic steering

07Built-in tools

Tools that say
what they missed.

When a tool stops at a limit or leaves something out, its result says what is missing, and the model can go back for it. The file, shell, web, code, and Python tools come from Workcell, which also runs on its own as an MCP server.

  1. grepSearchA search that reaches its bounds reports how much it withheld.
  2. URLCheckPrivate addresses are refused, and every redirect is checked again.
  3. 50 KiBCutA long page ends with a line that names the limit it reached.
  4. PDFAttachA model that reads PDFs receives the file inside the tool result.
IllustrationHow a search and a fetch report what they left out.Search and fetch with stated limits
  • Searches that report what they withheld

    A file_grep or file_glob call that reaches its bounds returns what it found and reports how much it withheld, so the model can tell a missing match from a file it never searched. Directory searches skip credential files such as .env and SSH private keys.

  • Guarded web fetching

    webfetch refuses private, loopback, and link-local addresses and checks every redirect the same way. Pages arrive in the character set they declare, and a page cut at 2,000 lines or 50 KiB ends with a line that names the limit.

  • PDFs as text or as the file

    A fetched PDF arrives as its text. In attachment mode, a model that reads PDFs receives the file itself inside the tool result, within a quarter of its context window. Saved sessions keep only the URL, name, and page count.

  • Python in an isolated worker

    python_execution runs scripts in a separate worker with no file system, network, environment variables, or subprocesses. That isolation is why the default permission policy runs it without a prompt.

Attached PDFs reach Claude through Anthropic or Amazon Bedrock, and custom models on a compatible API that declare PDF support. Other models receive the extracted text with a line that says why.

Tools provided by Workcell.

Built-in tools

Bring your model

Use the models
you already pay for.

First-class support for Anthropic and OpenAI, 16 built-in providers, local models through Ollama or llama.cpp, and custom endpoints that speak a supported API.

Anthropic · OpenAI · Google · GitHub Copilot · xAI · Mistral · DeepSeek · OpenRouter · Z.AI · Ollama · llama.cpp

Providers and sign-in

Claude subscription sign-in is experimental. Anthropic's terms limit Pro and Max subscriptions to official clients.

A model for each job

Nine jobs decide which model serves each kind of work. Pin a job to a model in /model, or let it follow Chat, Plan, Fast, or Best. A local endpoint can name its own Fast and Best models in providers.toml.

Chat
The main conversation. Picking a model in /model sets it.
Plan
Main turns in Plan mode, on the Chat model until you bind it.
Subagent
Delegated tasks, on the parent agent's model until you bind it.
Compact
Summaries when a context window fills.
Title
Session names, on Fast until you bind it.
Goal
The /goal evaluator, on Fast until you bind it.
Extract
Requirements for /extract and compaction, on Fast until you bind it.
Fast
The preferred small model, for routine calls.
Best
The provider's flagship. Point Plan or a prompt profile's subagents at it.
Model jobs

Everything in the box

60 capabilities
in one binary.

Each one links to its documentation. Experimental items are marked and stay off until you switch them on.

Experimental lab

Switch on what
you want to try.

These capabilities are implemented and experimental. Each one is off by default and needs its switch under [experimental] in your global caudra.toml.

Experimental

Move execution
into a VM.

Managed sandboxes run workspace tools in a separate VM. Model connections, provider credentials, and the conversation stay on your machine, so a compromised VM cannot read credentials it never received.

  • Network policy allows listed domains and IP ranges, with TLS hostname checks or an inspecting proxy.
  • File transfers between your machine and the VM go through a review.
  • Requires e2b-libvirt infrastructure that you or your operator run.

Transferred files can still carry secrets. The separation reduces exposure and is no guarantee against compromise.

How managed sandboxes work
  • Experimental

    Durable workflows

    Scripts launch subagents in phases, keep a journal, and can pause and resume. Each agent call can name a model job, such as Fast for a wide pass and Best for the answer you read. Built-in workflows cover deep research, change review, and root-cause analysis. They are heavily inspired by Grok Build workflows and mostly compatible with them.

  • Experimental

    JEV decision engine

    Typed decisions from an endpoint you configure: permission advice, Auto mode screening, shell effect and duration predictions, sampled web and MCP content screening, tool search ranking, skill suggestions, goal prescreening, and subagent routing to the Fast or Best model. Predictions add to deterministic permission rules and can miss risks.

  • Experimental

    Cross-session messaging

    Live sessions on one machine exchange messages, publish to topics, and share work through consumer groups.

  • Experimental

    Lua extensions

    Add your own commands, tools, and interface behavior in Lua when the built-in workflow needs something specific.

  • Experimental

    Remote Workcell

    Run workspace tools on a Workcell server while the conversation stays on your machine.

Privacy

Your history stays on your machine.

Sessions, retained output, and file change records are stored locally. Cloud models and network tools still receive what you send them.

  • No tracking

    Telemetry is off unless you send it to a collector you run.

  • Offline with a local model

    Use Ollama or llama.cpp without internet access, or point a providers.toml entry at ninfer-4090, the maintainer's custom inference engine for Qwen3.8-27B on one RTX 4090. Otherwise a normal run contacts your provider, refreshes the public models.dev catalog at most once a day, and reaches Exa when the agent searches the web.

  • Commands start from a cleared environment

    Shell commands receive only a short list of variables, such as PATH, HOME, and proxy settings, so API keys set for Caudra stay out of them.

  • Private addresses refused

    webfetch checks every URL and redirect before it connects, and refuses private, loopback, and link-local addresses.

  • Updates on request

    Caudra checks for updates only when you run caudra update or turn on the startup check.

Your next working session

Start with your
next project.

Install Caudra on macOS or Linux, connect a provider, and open your repository.

curl -fsSL https://caudra.ai/install.sh | sh