Open source, for your terminal
A coding agent that
turns smart context
into effective action.
Caudra brings steerable subagents, built-in token savings, a full workbench, and a session history you can always return to. It all ships in one native binary, so there are no plugins to assemble.
Leave a session working toward a goal while you are away. When you are at the keyboard, guide any subagent, ask a side question, or open the code beside the conversation.
curl -fsSL https://caudra.ai/install.sh | shfile_grepfile_readtaskChat opensMeasured by the maintainer in daily use
- 20B+tokens a month through Caudra
- <2 GBfor over a month of complete local history
- <1 sto load or save a 10k-turn session, subagents included
- ~97%Anthropic prompt-cache hit rate, and about 94% on OpenAI
Why Caudra exists
The agent I wanted
for my own working day.
I spend most of my working day with coding agents, and I have tried many of them. Each had ideas I liked. I wanted those ideas together, along with a few new techniques, in one standalone binary I could bring into any environment.
Caudra is that binary, and I push more than 20 billion tokens a month through it. With another agent, my local history grew by more than 1 GB a day. Caudra keeps over a month of my complete history in less than 2 GB, and a 10k-turn session still loads in under a second.
- Local history elsewhere
- >1 GB a day
- Local history in Caudra
- <2 GB for over a month
Built in and released together.
The core team develops, tests, and releases every capability together. You install one binary and get the whole agent, with no third-party plugins to vet or keep up to date. Lua extensions remain available as an experimental option.
Open-source credits
Caudra builds on ideas from these projects, with thanks.
- RTKshell output filtering
- ripwirecode maps
- Plannotatorpassage review
- Herdrterminal workspaces for agents
MethodThese are the maintainer's own measurements from daily use, not benchmarks. Timings were measured by hand. Results depend on models, projects, and hardware.
Why the name
Caudra comes
from caudate.
Caudra, pronounced KAW-druh, is named after the caudate nucleus. This part of the brain belongs to the circuits that connect evidence and goals to action. Studies link it to learning which actions lead to which outcomes.
The name describes how Caudra works. It turns the context of your task into an action, then reads the outcome before it chooses the next one.
The brain is only the inspiration for the name. Caudra is software, and it works toward the goal you set.
Sources: Grahn, Parkinson and Owen, 2008, Lau and Glimcher, 2007, Doi et al., 2020.
- 01Context and intent
- 02Evaluate evidence
- 03Select an action
- 04Execute with tools
- 05Observe the outcome
01Steer anything, any time
Guide any agent
while it works.
The input stays open while Caudra works. Queue the next task, add guidance to the current run, or replace it. Open any running subagent and guide it directly.
- EnterNextWaits for the current run and its background work.
- Ctrl+X gGuideJoins the run before its next model request.
- Ctrl+X xReplaceStops the run and starts your instruction.
- /tasksGuide a subagentReaches that subagent at its next turn boundary.
- ?The agent asksA question form waits for your choice.
- F2Ask
/btwClarify the choices in a side thread. - EscReturnThe form is unchanged, and the side thread stays out of history.
Next, Guide, and Replace
Enterqueues a prompt for after the current run.Ctrl+X gadds guidance before the next model request, andCtrl+X xstops the run and starts yours. Queued prompts stay editable.Guide a subagent directly
Open a running subagent from
/tasksand type. Your guidance stays visible until the subagent reads it at its next turn boundary.Read a brief as it is written
A subagent's chat opens while the model is still writing its instructions, so you can read the task before it starts.
Ask
/btwon the sideAsk about the conversation without adding to its history. When the main agent asks you a question,
F2opens/btwto clarify the choices, andEscreturns to the unchanged form.
02While you sleep
Set the finish line.
Wake to results.
/goal keeps a session working until a separate evaluator finds evidence that your condition is met. Background work reports back without polling, and your permission rules decide what runs while you are away.
- /goalSet the conditionFor example: tests pass and clippy is clean.
- 1WorkThe turn ends and background results settle.
- 2EvaluateA separate call looks for evidence in the transcript.
- ↺ContinueAn unmet goal starts another turn, up to the limit.
- ✓DoneA met goal clears itself.
Goals that check evidence
After each work turn, a separate model call with no tools reads the transcript. An unmet goal starts another turn. A met goal clears itself.
Background work wakes the agent
Tasks and shell jobs can keep running in the background. Their reports start the next step at a safe boundary, with no polling.
Permissions that read the command
Shell chains and pipelines split into separate commands for approval. Approve once, for the conversation, the project, or every project, and accept patterns Caudra suggests from repeated approvals.
Plan first, then hear back
Caudra opens in Plan mode, where the agent edits only its plan and asks before commands it cannot prove read-only. Notifications tell you when a turn finishes or a prompt needs your answer.
03Built-in token savings
Token savings,
built in.
Every turn re-sends the conversation, so a noisy result costs tokens again on every later turn until compaction. Caudra keeps results small and round-trips few, starting with RTK-style shell output filtering that is on by default.
- $RunYou watch the raw output while the command runs.
- filterTrimCommand-aware rules remove routine noise from the completed output.
- →ReadThe model receives the filtered result.
- rawCompareSwitch between filtered and raw views in the transcript.
RTK-style output filtering
Command-aware rules, most of them from RTK, trim build and test noise before the model reads the output, and a filtered result is always smaller than the raw one. A failure is never reported as a success, and a failing command keeps its first and last lines. Progress bars collapse to one row, and you can switch between filtered and raw views.
Structure before content
file_indexreturns a file outline with signatures and line numbers. Thecode_*tools rank symbols and show callers, impact, and the tests they can trace to a change. They parse the source on the fly, so there is no index to build or maintain and nothing to configure.Fewer, smaller turns
Oversized results reach the model as a bounded head and tail, and
tool_outputsearches the rest.batchruns independent calls in one turn, and subagents keep their exploration out of the main context.Requests shaped for caching
Caudra sends a per-conversation cache key where providers accept one, and
/usagescores the cache hit rate for every model.
04A workbench beside the agent
Open the code
without leaving.
Ctrl+X w brings up a file explorer, tabbed editor, project search, and source control while the session keeps running. It lives in the terminal, so it comes along over SSH.
Workbench layout: a sidebar with files, Git, and search views on the left, editor tabs and a buffer on the right, and a status row that offers Ctrl+X Enter to send a reference.
- Ctrl+X rOpenReview the last reply, or any message from its menu.
- Shift+↓SelectTake the rows that need work, or drag across them.
- EnterNoteWrite a note on the selection.
- Ctrl+SSendEvery note lands in the prompt as one block.
Read the actual change
Inspect diffs, browse the commit graph, and stage, unstage, or discard per file or folder.
Point at exact lines
Ctrl+X Entersends the file, line, or selection to the composer as a mention such as@src/api.ts:L10-L20. Caudra puts those lines in the request.Review passages of a reply
Mark the parts of an answer that need work, add a note to each, and send every note back as one prompt.
Replies rendered for engineers
Tables, highlighted code, Unicode maths, and Mermaid flowcharts render in the terminal. Copying a passage gives back its Markdown source.
05Never lose the thread
Go back to any
point in the work.
Every session keeps its full history, subagent transcripts included, in compact local storage. Revert the conversation, the files, or both from a message, and preview the file changes first.
- ⋮Open the menuBeside any message in the main transcript.
- 1PreviewRevert files counts what would be created, replaced, or deleted.
- 2ApplyRepeat the action. A conflict found before writing aborts the whole revert.
- ↶UnrevertPut the files and the conversation head back.
Large sessions open fast
Payloads are compressed, appending saves write only new rows, and the transcript lays out only what is on screen. In the maintainer's use, a 10k-turn session loads in under a second.
Revert with a preview
Each tool call that may change files is recorded before and after it runs, so a revert touches only those files. The first Revert files shows what would change, and repeating it applies the revert.
Memory that outlasts the session
The agent keeps tagged notes on project gotchas and decisions outside your repository. Requests carry only the tags until a note is needed, and
/memorylets you read, edit, or delete them.Fork and branch out
Fork from a message to try another approach.
/worktree newmoves a session into a fresh Git worktree with its conversation and plan.
06Automatic steering
Keep any model
on task.
Smaller local models such as Qwen3.8-27B need more help to finish a task. Caudra nudges the model when a turn stalls, gets cut off, loops, or stops after announcing work. The rules react to what a reply did, so the same defaults work well with flagship models.
- 1Stop earlyThe reply ends on “I will run the tests now” and calls no tool.
- 2NudgeThe
abandoned_turnrule asks the model to do that work now. - 3Read the nudgeA dim row in the transcript holds the exact text the model received.
- ↺ContinueThe next reply can call its tools. By default, a third announcement in a row ends the turn as written.
Nudges for stalled turns
After an empty or cut-off reply, Caudra asks the model to continue. After “I will run the tests now” with no tool call, it asks for the work itself. The third identical tool call in a row is refused before it runs.
Hints when work goes in circles
When the model repeats a tool cycle or an answer, or keeps making failed calls, a hint asks it to reconsider its approach. Hints stop at four per run by default and never reopen a finished answer.
Sensible defaults, tuned per model
Every rule is on by default, and budgets cap how often each one fires. Change a threshold, a budget, or the wording of a nudge for all models, or only for one exact
provider/model-id.Every nudge in the transcript
Each nudge appears as a dim row. Click it to read the exact text the model received.
07Built-in tools
Tools that say
what they missed.
When a tool stops at a limit or leaves something out, its result says what is missing, and the model can go back for it. The file, shell, web, code, and Python tools come from Workcell, which also runs on its own as an MCP server.
- grepSearchA search that reaches its bounds reports how much it withheld.
- URLCheckPrivate addresses are refused, and every redirect is checked again.
- 50 KiBCutA long page ends with a line that names the limit it reached.
- PDFAttachA model that reads PDFs receives the file inside the tool result.
Searches that report what they withheld
A
file_greporfile_globcall that reaches its bounds returns what it found and reports how much it withheld, so the model can tell a missing match from a file it never searched. Directory searches skip credential files such as.envand SSH private keys.Guarded web fetching
webfetchrefuses private, loopback, and link-local addresses and checks every redirect the same way. Pages arrive in the character set they declare, and a page cut at 2,000 lines or 50 KiB ends with a line that names the limit.PDFs as text or as the file
A fetched PDF arrives as its text. In attachment mode, a model that reads PDFs receives the file itself inside the tool result, within a quarter of its context window. Saved sessions keep only the URL, name, and page count.
Python in an isolated worker
python_executionruns scripts in a separate worker with no file system, network, environment variables, or subprocesses. That isolation is why the default permission policy runs it without a prompt.
Bring your model
Use the models
you already pay for.
First-class support for Anthropic and OpenAI, 16 built-in providers, local models through Ollama or llama.cpp, and custom endpoints that speak a supported API.
Anthropic · OpenAI · Google · GitHub Copilot · xAI · Mistral · DeepSeek · OpenRouter · Z.AI · Ollama · llama.cpp
Providers and sign-in- ChatGPTSign in with your ChatGPT subscription.
- GitHub CopilotReuse an existing Copilot sign-in.
- xAISign in with your xAI account.
- ClaudeSign in with your Claude subscription.Experimental
Claude subscription sign-in is experimental. Anthropic's terms limit Pro and Max subscriptions to official clients.
A model for each job
Nine jobs decide which model serves each kind of work. Pin a job to a model in /model, or let it follow Chat, Plan, Fast, or Best. A local endpoint can name its own Fast and Best models in providers.toml.
- Chat
- The main conversation. Picking a model in
/modelsets it. - Plan
- Main turns in Plan mode, on the Chat model until you bind it.
- Subagent
- Delegated tasks, on the parent agent's model until you bind it.
- Compact
- Summaries when a context window fills.
- Title
- Session names, on Fast until you bind it.
- Goal
- The
/goalevaluator, on Fast until you bind it. - Extract
- Requirements for
/extractand compaction, on Fast until you bind it. - Fast
- The preferred small model, for routine calls.
- Best
- The provider's flagship. Point Plan or a prompt profile's subagents at it.
Everything in the box
60 capabilities
in one binary.
Each one links to its documentation. Experimental items are marked and stay off until you switch them on.
Agents
Tools
- File read and search, with one editor per model
- File outlines with
file_index - Code maps for 37 languages and formats
- Shell with output filtering
- Isolated Python for computation
- Parallel calls with
batch - Retained output search
- Web search and fetch, with PDFs as text or as the file
- Image generation with a ChatGPT login
- Tool-call JSON repair
- Tools loaded on demand
Context
Sessions
Interface
Safety
Experimental lab
Switch on what
you want to try.
These capabilities are implemented and experimental. Each one is off by default and needs its switch under [experimental] in your global caudra.toml.
Move execution
into a VM.
Managed sandboxes run workspace tools in a separate VM. Model connections, provider credentials, and the conversation stay on your machine, so a compromised VM cannot read credentials it never received.
- Network policy allows listed domains and IP ranges, with TLS hostname checks or an inspecting proxy.
- File transfers between your machine and the VM go through a review.
- Requires e2b-libvirt infrastructure that you or your operator run.
Transferred files can still carry secrets. The separation reduces exposure and is no guarantee against compromise.
How managed sandboxes work- Experimental
Durable workflows
Scripts launch subagents in phases, keep a journal, and can pause and resume. Each agent call can name a model job, such as Fast for a wide pass and Best for the answer you read. Built-in workflows cover deep research, change review, and root-cause analysis. They are heavily inspired by Grok Build workflows and mostly compatible with them.
- Experimental
JEV decision engine
Typed decisions from an endpoint you configure: permission advice, Auto mode screening, shell effect and duration predictions, sampled web and MCP content screening, tool search ranking, skill suggestions, goal prescreening, and subagent routing to the Fast or Best model. Predictions add to deterministic permission rules and can miss risks.
- Experimental
Cross-session messaging
Live sessions on one machine exchange messages, publish to topics, and share work through consumer groups.
- Experimental
Lua extensions
Add your own commands, tools, and interface behavior in Lua when the built-in workflow needs something specific.
- Experimental
Remote Workcell
Run workspace tools on a Workcell server while the conversation stays on your machine.
Privacy
Your history stays on your machine.
Sessions, retained output, and file change records are stored locally. Cloud models and network tools still receive what you send them.
No tracking
Telemetry is off unless you send it to a collector you run.
Offline with a local model
Use Ollama or llama.cpp without internet access, or point a
providers.tomlentry at ninfer-4090, the maintainer's custom inference engine for Qwen3.8-27B on one RTX 4090. Otherwise a normal run contacts your provider, refreshes the public models.dev catalog at most once a day, and reaches Exa when the agent searches the web.Commands start from a cleared environment
Shell commands receive only a short list of variables, such as
PATH,HOME, and proxy settings, so API keys set for Caudra stay out of them.Private addresses refused
webfetchchecks every URL and redirect before it connects, and refuses private, loopback, and link-local addresses.Updates on request
Caudra checks for updates only when you run
caudra updateor turn on the startup check.
Your next working session
Start with your
next project.
Install Caudra on macOS or Linux, connect a provider, and open your repository.
curl -fsSL https://caudra.ai/install.sh | shRead the quick startInspect the scriptWindows and other options