Agent Codemode Explained
Code Mode: Moving Agent Control Flow Out of the LLM
Most agents still use tools one call at a time.
The model decides to call a tool. The runtime executes it. The result is appended to the conversation. The model runs again and decides what to do next.
For a short task this works well.
For a task that needs twenty API calls, filtering, retries, a loop, or several dependent operations, the design becomes expensive:
LLM
↓
tool A
↓
LLM
↓
tool B
↓
LLM
↓
tool C
↓
LLM
Each edge through the LLM costs inference time. Tool results also accumulate in the context even when most of the data is only needed to decide the next call.
Code Mode changes this boundary.
Instead of asking the model to emit every tool call, give it one tool that executes a small program:
LLM
↓
program
├─ tool A
├─ tool B
├─ tool C
└─ local filtering
↓
LLM
The program becomes the control plane for a section of the agent run.
This idea now appears in Cloudflare Code Mode, Anthropic’s programmatic tool calling, PydanticAI Code Mode, Hugging Face smolagents, and, as of September 29, 2026, Pi.
They share the same basic observation:
Tool calling is an unusually bad programming language.
JSON tool calls can express:
call X with Y
A normal programming language can express:
call X for every Y
call X and Z concurrently
retry X if it fails
filter the result
call Z only if condition Q is true
store this value and reuse it later
return only these five fields
The difference is larger than syntax.
Code Mode changes where computation happens.
The normal tool loop
Suppose an agent needs to find all unhealthy services and inspect their recent deployments.
With normal tool calling it might do:
LLM → list_services()
LLM ← 300 services
LLM → get_health(service_1)
LLM ← result
LLM → get_health(service_2)
LLM ← result
...
A capable model can issue several calls in parallel, but the model still has to construct the calls and receive their results.
A better tool might support bulk operations. But now every tool designer has to predict every future composition:
get_health_for_services(...)
get_unhealthy_services(...)
get_unhealthy_services_with_deployments(...)
This is where Code Mode becomes interesting.
The model can instead write something equivalent to:
services = await list_services()
health = await gather(
*[get_health({"id": s["id"]}) for s in services]
)
bad = [
s for s, h in zip(services, health)
if h["status"] != "healthy"
]
deployments = await gather(
*[get_deployments({"service_id": s["id"]}) for s in bad]
)
return [
{
"service": s["name"],
"deployment": d[0]
}
for s, d in zip(bad, deployments)
]
The LLM sees the task once and writes the control flow once.
The intermediate 300 service records and hundreds of health responses do not need to become conversation history.
That is the important part.
Code Mode is not mainly a prettier way to call tools.
It is a way to move intermediate computation out of model context.
Cloudflare: code as a compact API plan
Cloudflare has probably made the clearest argument for Code Mode.
Its problem is unusually visible. The Cloudflare API has more than 2,500 endpoints. Exposing every API endpoint as an MCP tool means shipping a huge collection of JSON schemas into the model context before the agent does anything useful.
Cloudflare’s Code Mode MCP server instead exposes essentially two operations: search for an API surface, then execute code against the discovered typed API.
Cloudflare reports that its whole API can be exposed in roughly 1,000 tokens this way. Its comparison estimates 1.17 million input tokens if the equivalent API were represented as ordinary MCP tool definitions.
This solves the first Code Mode problem:
too many tools
There is a second problem:
too much intermediate data
Cloudflare describes generated code as a compact plan. Calls, filtering and transformations happen inside the execution environment, and only selected output returns to the model.
Their current interface exposes primitives such as:
codemode.search(...)
codemode.describe(...)
codemode.step(...)
codemode.run(...)
search() performs progressive discovery. describe() loads the detailed interface only when needed. step() records work that needs replay semantics. run() executes reusable snippets.
This is an important refinement.
A naive Code Mode implementation still puts every available function declaration in the system prompt:
one code tool
+
10,000 function declarations
That fixes intermediate results but not tool-definition cost.
Cloudflare instead makes the API itself lazy.
The model starts with a small discovery interface and pulls schemas when it needs them.
Conceptually:
┌──────── search
LLM → code│
├──────── describe
│
├──────── API A
├──────── API B
└──────── API C
The tool catalog becomes something closer to a filesystem or symbol table than a prompt.
That is likely the right abstraction for very large tool spaces.
Cloudflare’s sandbox matters
Once the model can write arbitrary JavaScript, eval() is not an acceptable runtime.
Cloudflare executes generated code in isolated Workers. Its MCP design blocks direct outbound access by default; generated code reaches the outside world through capabilities supplied by the host. Credentials remain outside the generated program.
The distinction is useful:
code execution capability
≠
external authority
The sandbox can be computationally expressive while having almost no ambient authority.
The program may know how to write:
await dns.deleteZone(...)
but the actual authority still belongs to the host connector.
This means authorization does not have to be implemented by trusting generated code.
It remains at the tool boundary.
Cloudflare’s durable runtime goes further. It records execution history and can pause generated code for approval, then replay completed work and continue the program.
That turns Code Mode from “run some generated JavaScript” into an agent runtime primitive.
Anthropic: programmatic tool calling
Anthropic independently arrived at almost exactly the same model.
Their current term is programmatic tool calling.
Claude writes Python inside its code execution environment. Tool calls from the program cross back to the host. When a tool returns, execution resumes. Intermediate results remain inside the execution environment instead of being inserted into Claude’s context.
The execution model looks like this:
Claude generates Python
│
▼
sandbox
│
├── tool()
│ │
│ └──── host executes tool
│
├── process result locally
│
├── tool()
│
└── print small result
│
▼
Claude continues
This is not equivalent to asking Claude to generate a Python script and running the script after the model finishes.
The interesting part is the suspended host call.
The Python program is effectively an orchestration language around external capabilities.
Anthropic reports useful measurements here. On its BrowseComp and DeepSearchQA experiments, programmatic tool calling improved performance by an average of 11% while using 24% fewer input tokens. On an internal 75-tool project-management benchmark it reports roughly 38% lower billed input tokens with unchanged task accuracy. But on τ²-bench tasks dominated by one or two sequential calls, it cost about 8% more.
That last number is important.
Code Mode is not free.
For:
call tool
return answer
generating a program and starting a code environment is extra work.
Code Mode becomes useful when the code replaces enough model-mediated control flow to pay for itself.
Pi: Code Mode moves into the harness
Pi is a more interesting implementation for coding agents because Code Mode is not only an API optimization.
It is integrated into the agent harness.
A large Pi change was merged on September 29, 2026. The implementation runs model-generated JavaScript in a QuickJS VM compiled to WebAssembly. The script can call Pi tools as async functions. Nested calls do not enter the LLM context individually; only explicit script output and the final return value do.
The standalone package describes the capability directly:
const source = await tools.read({ path: "package.json" })
text(JSON.parse(source).name)
Every execution gets a QuickJS VM in a worker. The VM has its own WASM memory and no general route back into the Node host except the explicitly installed host-call bridge.
This makes Pi’s implementation worth studying.
Pi already had very powerful tools:
read
write
edit
bash
For coding work, bash alone is close to universal.
So why add Code Mode?
Because universality is not the same as a good agent interface.
Consider:
cat package.json | jq ...
versus:
const p = await tools.read({path: "package.json"})
const pkg = JSON.parse(p)
Both work.
But the second form can combine normal agent tools, MCP tools and future host capabilities under one runtime, without turning every operation into shell text.
More importantly, nested calls remain visible to the harness.
Pi routes nested calls through the normal validation and tool hooks. They carry a parentToolCallId, and nested calls are recorded on the parent tool result.
This gives a useful structure:
codemode call
│
├── read
├── grep
├── MCP: search
└── bash
instead of hiding the entire operation behind an opaque process.
That matters for observability, permissions and evaluation.
Pi’s tool exposure model
Pi also exposes a design problem that smaller Code Mode demos can avoid:
Which tools should the model see directly?
Pi now distinguishes different forms of tool exposure and supports two useful Code Mode configurations.
In codemode.mode = "on", normal tools may remain directly visible while Code Mode can also call them.
In codemode.mode = "only", the model sees Code Mode as the primary interface and calls ordinary tools through generated code instead.
This gives two different agent architectures.
Hybrid
LLM
├── read
├── edit
├── bash
└── codemode
├── read
├── edit
└── MCP tools
Code-only
LLM
└── codemode
├── read
├── edit
├── bash
├── MCP
└── extensions
The second one is more radical.
The model-facing action language stops being a collection of JSON tools.
It becomes JavaScript.
Pi also places a token budget on inline declarations. Tools that do not fit can be found dynamically through tool search.
So Pi is converging on the same two-layer structure as Cloudflare:
small language runtime
+
lazy capability discovery
This looks increasingly less like an agent feature and more like an operating-system interface.
PydanticAI: Python as the capability language
PydanticAI implements the same pattern using Python and Monty.
Eligible tools disappear from the direct model-facing tool list and become functions available inside one run_code tool.
The generated Python can use loops, conditions, variables and asyncio.gather. Intermediate data stays in the sandbox.
For example:
rows = await search("failed builds")
details = await asyncio.gather(
*[get_build(row["id"]) for row in rows]
)
return [
x for x in details
if x["status"] == "failed"
]
PydanticAI’s implementation is particularly interesting because Monty is not CPython running in a container.
Monty is a restricted Python interpreter written in Rust for AI-generated programs.
By default it has no normal filesystem, environment variables or network access. Host capabilities have to be passed in deliberately.
That gives a capability model similar to Pi’s QuickJS sandbox:
generated program
│
├── computation
├── variables
├── branching
└── injected host functions
The model gets the useful parts of Python without automatically inheriting the authority of the host Python process.
PydanticAI also supports persistent REPL state during an agent run. Variables, imports and helper functions can survive across run_code calls.
That changes Code Mode again.
It is no longer just:
LLM → program → result
It can become:
LLM → persistent computational workspace
That distinction will matter for long-running agents.
smolagents: code as the action format
Hugging Face’s smolagents predates some of these systems and takes the idea further at the agent-loop level.
Its CodeAgent uses Python code as the normal action representation. ToolCallingAgent is the alternative that emits conventional structured tool calls.
The important observation from smolagents is simple:
JSON is a serialization format.
A programming language is a control-flow language.
If the task is:
search A
search B
take the intersection
fetch each result
rank them
representing that as a program is natural.
Representing it as five separate JSON actions requires the LLM itself to become the control-flow interpreter.
That is the hidden cost of ordinary agent loops.
Code Mode is a tiny compiler boundary
One way to understand these systems is to stop thinking about “code execution.”
Think of the LLM as producing a program for a small virtual machine.
The traditional agent loop is:
observation
↓
LLM
↓
one action
↓
observation
Code Mode becomes:
observation
↓
LLM
↓
program
↓
runtime
↓
many actions
↓
compressed observation
The LLM has effectively compiled several future decisions into one artifact.
This means Code Mode is useful only when those future decisions can be expressed algorithmically.
Suppose a result requires semantic judgment:
Read this incident report.
Understand what probably caused the outage.
Choose what to investigate next.
A JavaScript if statement cannot replace the next LLM call.
But many agent decisions are not semantic reasoning:
for every item
if status == failed
take first five
sort by date
retry twice
call these concurrently
extract this field
Today we often waste LLM calls performing those operations.
Code Mode removes them from the model loop.
The real token saving has two parts
Discussions of Code Mode often combine two independent optimizations.
They should be separated.
1. Tool definition compression
Instead of:
tool A schema
tool B schema
tool C schema
...
tool 2500 schema
give the model:
search(query)
describe(symbol)
execute(code)
Then discover interfaces lazily.
This is the part that lets Cloudflare expose thousands of endpoints in a small fixed context.
Anthropic described a similar MCP approach where tool definitions are discovered through files instead of loading the whole MCP catalog into context. In one example, it reduced tool-related context from roughly 150,000 tokens to about 2,000.
2. Intermediate result compression
Instead of:
tool result
→ LLM
→ tool result
→ LLM
→ tool result
→ LLM
do:
tool result
→ program
→ tool result
→ program
→ filter
→ final small result
→ LLM
These solve different scaling problems.
A serious Code Mode implementation should probably do both.
Code Mode does not remove the agent loop
A tempting conclusion is:
Why not put the entire agent in code?
Because generated code only knows what the model knew when it wrote the program.
Consider:
result = await investigate()
If result contains something genuinely surprising, deciding what it means may require another model inference.
A normal program can branch on structure:
if result["status"] == "failed":
It cannot reliably branch on arbitrary semantics:
if result implies that the database migration
is probably unrelated to the outage:
You can put another model behind a function call:
if await classify(result, question):
but now the model has returned through another door.
So the likely architecture is not:
Code Mode replaces Agent Loop
It is:
Agent Loop
│
├── semantic decision
│
└── Code Mode
├── deterministic control flow
├── data processing
├── bulk tool calls
└── cheap decisions
The boundary between these two layers is one of the more interesting agent-design problems.
Code Mode also changes tool design
Today’s tool APIs are partly shaped by the weakness of model tool calling.
We create high-level tools such as:
search_and_summarize_incidents
find_recent_failed_deployments
get_customer_context
because asking the model to compose eight primitive tools is expensive and unreliable.
Once tools can be composed inside code, smaller primitives become practical:
search
fetch
read
write
query
execute
The generated program provides the composition.
This does not mean every tool should become low level.
Every tool boundary is also a:
permission boundary
validation boundary
observability boundary
failure boundary
But Code Mode changes the trade-off.
Without Code Mode, coarse tools save LLM turns.
With Code Mode, coarse tools need to justify themselves for another reason.
The security boundary is not the language sandbox
Running generated code in QuickJS, Monty or a Dynamic Worker is necessary.
It is not sufficient.
The dangerous operation is usually not:
while (true) {}
It is:
await tools.deleteDatabase(...)
The generated program should therefore be treated as untrusted orchestration around trusted capabilities.
The useful architecture is:
generated code
│
sandbox boundary
│
▼
capability bridge
│
┌──────────────┼──────────────┐
▼ ▼ ▼
read tool GitHub tool DB tool
│ │ │
policy policy policy
│ │ │
execute execute execute
Permissions belong below the code.
Pi’s nested calls pass through normal tool validation and hooks.
Cloudflare explicitly says generated code does not replace authorization and that permissions should still be enforced in upstream handlers.
Anthropic similarly warns that its tool caller configuration should not itself be treated as a security boundary.
These implementations are converging on the same rule:
sandbox computation; authorize capabilities.
Observability becomes harder
Direct tool calling gives a convenient trace:
model
tool
model
tool
model
tool
Code Mode compresses the model trace:
model
codemode
model
That is good for tokens.
It can be bad for debugging if the runtime treats the code invocation as one opaque tool call.
A useful Code Mode runtime therefore needs nested traces:
codemode #42
├─ search #43
│ └─ 42 ms
├─ fetch #44
│ └─ 181 ms
├─ fetch #45
│ └─ 172 ms
└─ update #46
└─ approval required
Pi explicitly records nested calls and their parent IDs.
Cloudflare’s durable runtime records steps and replay information.
PydanticAI exposes nested Code Mode calls in its tracing and can integrate them with durable execution systems.
This is not optional infrastructure.
Once a generated program can call fifty tools, “the Code Mode tool failed” is not useful telemetry.
Code Mode introduces new failure modes
It removes some agent failures and adds others.
Normal tool calling can fail because the model:
forgets a step
calls a tool repeatedly
loads too much data
serializes work unnecessarily
loses intermediate state
Code Mode can reduce those.
But now the generated program can have:
syntax errors
type errors
infinite loops
unbounded fan-out
bad retry logic
stale state
race conditions
incorrect aggregation
side effects inside loops
A serious runtime needs limits.
At minimum:
execution timeout
memory limit
maximum tool calls
concurrency limit
output limit
nesting limit
cancellation
Pi already places deadlines and output limits around its sandbox, and large outputs can be spilled rather than pushed back into context.
PydanticAI’s Monty environment intentionally exposes a restricted subset of Python and limits host access.
This suggests another useful rule:
Code Mode should be powerful enough for orchestration, not powerful enough to become an accidental operating system.
JavaScript or Python?
The current implementations split here.
Cloudflare and Pi use JavaScript.
Anthropic, PydanticAI and smolagents use Python.
This is less important than it first appears.
The generated language needs:
variables
arrays/maps
loops
conditionals
functions
async calls
parallel calls
basic data transformation
Both languages satisfy that.
The more important properties are:
How cheap is sandbox startup?
What host capabilities exist?
Can execution be interrupted?
Can state persist?
Can nested calls be traced?
Can calls pause for approval?
Can programs be replayed?
Can tool types be generated cleanly?
Pydantic’s Monty is interesting precisely because it treats this as a runtime problem rather than “just execute Python.”
Pi does the same with QuickJS in WASM.
The language is the visible part.
The runtime is the product.
A useful comparison
| System | Generated language | Main sandbox | Tool discovery | Persistent state | Main emphasis |
|---|---|---|---|---|---|
| Cloudflare Code Mode | JavaScript / TypeScript-shaped API | Workers / isolated execution | search() + describe() | Durable runtime / snippets | Huge API surfaces, MCP, durable execution |
| Pi Codemode | JavaScript | QuickJS WASM in worker | Inline declarations + tool search | Session store | Coding-agent harness integration |
| Anthropic Programmatic Tool Calling | Python | Anthropic code execution container | Provider tool definitions / related advanced tool use | Container reuse | Managed programmatic calling |
| PydanticAI Code Mode | Python | Monty | Tool Search integration | Persistent REPL state | Safe embeddable runtime |
| smolagents CodeAgent | Python | Local or configured sandbox | Agent tool set | Agent execution state | Code as primary action representation |
They are different implementations of roughly the same abstraction:
LLM
↓
small program
↓
capability runtime
↓
tools
What I think Code Mode actually is
Calling this feature “Code Mode” makes it sound optional:
normal agent
+
sometimes let it execute code
I think the deeper interpretation is different.
Normal tool calling asks an LLM to act as both:
reasoning engine
and
workflow interpreter
Code Mode separates them.
The LLM produces a temporary program.
The runtime executes it.
Then the LLM returns when semantic reasoning is needed again.
That looks much closer to a compiler architecture:
intent
↓
LLM
↓
temporary program / IR
↓
runtime
↓
observations
The program does not need to survive forever.
It can exist for 200 milliseconds.
It can be generated specifically for one observation.
It can call five tools and disappear.
This makes it a useful intermediate representation between language models and tools.
The interesting future is not bigger tool calls
There is a natural progression here.
Generation 1
LLM → text
Generation 2
LLM → JSON tool call
Generation 3
LLM → code → many tool calls
The next step is probably not simply more code.
The runtime can start providing primitives that are cheaper or safer than asking the main model again:
classify(...)
validate(...)
wait(...)
watch(...)
parallel(...)
approve(...)
spawn(...)
store(...)
Pi already exposes model classifiers from its Code Mode environment. Its merged implementation allows generated programs to inspect the model catalog and invoke classification operations.
This is interesting because the program no longer orchestrates only tools.
It can orchestrate computation at different intelligence levels.
For example:
const issues = await tools.github.searchIssues({ repo })
const relevant = await Promise.all(
issues.map(issue =>
models.classify(
"Is this issue related to authentication?",
issue.body
)
)
)
return issues.filter((_, i) => relevant[i])
The expensive model writes the algorithm once.
A cheaper decision mechanism executes inside the loop.
Now the architecture becomes:
main LLM
│
▼
code
┌────────┼─────────┐
▼ ▼ ▼
tools small AI local compute
That starts to look like a real computational runtime for agents.
When I would use Code Mode
Code Mode is a good fit when the task contains:
fan-out
loops
large intermediate results
filtering
aggregation
parallel calls
large tool catalogs
reusable procedures
mostly deterministic control flow
Examples:
check 100 services and return the failed ones
search ten sources and deduplicate results
read every changed file and run a validator
query several APIs and join their results
inspect every failed CI run
find all matching records, update a subset
I would keep direct tool calls for:
open this file
run this command
ask the user
make one API request
perform one dangerous action requiring approval
Anthropic’s own measurements support this distinction: workloads with many calls and large intermediate results benefit; workloads with one or two short sequential calls may not.
So I would not make Code Mode a universal replacement for tools.
I would make it a first-class execution path beside them.
The design I would build
A minimal Code Mode implementation only needs:
one sandbox
one code tool
a bridge to existing tools
A production implementation needs more.
I would want:
1. Capability-based sandbox
2. Typed tool bindings
3. Progressive tool discovery
4. Direct and code-only tool exposure
5. Nested tool-call tracing
6. Per-call policy and approval
7. Tool-call concurrency limits
8. Runtime deadline
9. Output budget
10. Persistent session-local state
11. Cancellation
12. Optional durable execution
The central API can remain very small:
tools.search(...)
tools.describe(...)
await tools.foo(...)
store(...)
load(...)
text(...)
Everything else belongs in the harness.
The generated program should not know where credentials live, how permissions work, how traces are exported, or whether a tool is implemented through MCP, HTTP, a shell process or another agent.
That is the host’s job.
One consequence for MCP
Code Mode also changes how I think about MCP.
MCP is useful as a transport and capability description protocol.
It is much less convincing as the language an LLM should directly program against.
Exposing hundreds of MCP tools directly to the model couples:
capability transport
with:
model action representation
Those do not need to be the same thing.
Cloudflare makes this explicit: MCP can remain underneath Code Mode. An MCP server may itself expose a Code Mode interface, or existing MCP tools can become functions callable by generated code.
Pi has now made a similar separation. MCP tools can be available through Codemode or discovered and exposed directly depending on configuration.
A cleaner stack is:
LLM
↓
Code Mode
↓
tool abstraction
↓
MCP / HTTP / local functions / shell / another agent
MCP becomes infrastructure.
It stops consuming the entire model interface.
Conclusion
Code Mode looks like a token optimization at first.
It is more useful to see it as a change in the execution model.
Traditional agents repeatedly ask the LLM:
What should I do next?
even when “next” is determined by a loop, a filter or a boolean condition.
Code Mode lets the model answer a larger question:
What program should control the next section of execution?
Then the harness runs that program against a restricted set of capabilities.
This reduces model round trips.
It keeps intermediate data out of context.
It makes huge tool catalogs practical.
It also creates a new runtime layer that needs sandboxing, permissions, tracing, resource limits and state.
Cloudflare approaches the problem from large APIs.
Anthropic approaches it from model-native programmatic tool calling.
PydanticAI approaches it from a safe embedded Python runtime.
Pi has now integrated the pattern directly into a coding-agent harness.
The implementations differ.
The architecture is converging:
LLM for semantic decisions.
Code for control flow.
Tools for authority.
Harness for policy and state.
That separation is more important than the name “Code Mode.”