<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom"><title>Ju Lin's AI Weblog</title><subtitle>An independent research notebook on AI engineering, agents, models and the systems around them.</subtitle><id>https://julin.ai/atom/index.xml</id><link rel="self" type="application/atom+xml" href="https://julin.ai/atom/index.xml"/><link rel="alternate" type="text/html" href="https://julin.ai/"/><author><name>Ju Lin</name></author><updated>2026-10-05T00:00:00+13:00</updated><entry><title>Vibe Code in a Secured Way</title><id>https://julin.ai/2026/10/05/vibe-code-secured/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/10/05/vibe-code-secured/"/><published>2026-10-05T00:00:00+13:00</published><updated>2026-10-05T00:00:00+13:00</updated><category term="security"/><category term="ai-coding"/><category term="best-practices"/><content type="html">&lt;p&gt;AI coding is powerful. It is also risky if you are not cautious.&lt;/p&gt;
&lt;p&gt;The AI does the work. You are responsible for everything.&lt;/p&gt;
&lt;p&gt;The bottom-line skill for every engineer nowadays is the judgment to recognize when the AI is performing risky operations.&lt;/p&gt;
&lt;h2 id="know-what-youre-exposing"&gt;Know What You&amp;rsquo;re Exposing&lt;/h2&gt;
&lt;p&gt;Coding sessions can be stored remotely. LLM requests can be routed through proxies. It&amp;rsquo;s a good practice to limit what the agent and the LLM model can see.&lt;/p&gt;
&lt;p&gt;This is not just about credentials.&lt;/p&gt;
&lt;p&gt;When you work on any data or codebase, ask yourself: is this data appropriate for the AI to see?&lt;/p&gt;
&lt;p&gt;When it involves customer data or personal information, think carefully about whether redaction is needed. The AI should not have visibility into data it does not need to accomplish the task.&lt;/p&gt;
&lt;h2 id="the-ai-can-ignore-your-instructions"&gt;The AI Can Ignore Your Instructions&lt;/h2&gt;
&lt;p&gt;You can ask the AI not to look at something. It can ignore you anyway.&lt;/p&gt;
&lt;p&gt;More dangerous: the AI can be spoofed through prompt injection and tricked into performing malicious instructions.&lt;/p&gt;
&lt;p&gt;This risk magnifies when you open up the entire internet to the agent. The internet contains everything, and not all of it is trustworthy.&lt;/p&gt;
&lt;h2 id="trust-your-tools-and-dependencies"&gt;Trust Your Tools and Dependencies&lt;/h2&gt;
&lt;p&gt;Only use the tools or MCP servers that you trust.&lt;/p&gt;
&lt;p&gt;Secure your dependencies. Use tools like Dependabot to scan for vulnerabilities. Monitor supply chain attacks—remember the LiteLLM PyPI attack? That kind of compromise can propagate through your entire toolchain.&lt;/p&gt;
&lt;p&gt;If you install an untrusted package, you are giving the AI access to whatever that package can do. And the AI might use it.&lt;/p&gt;
&lt;h2 id="the-practical-defense"&gt;The Practical Defense&lt;/h2&gt;
&lt;p&gt;When an AI agent is running:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Audit what data it has access to. Is it necessary?&lt;/li&gt;
&lt;li&gt;Audit what tools it can call. Are they all trustworthy?&lt;/li&gt;
&lt;li&gt;Audit what network access it has. Does it need the internet?&lt;/li&gt;
&lt;li&gt;Audit what code it writes. Does it make risky calls?&lt;/li&gt;
&lt;li&gt;Audit what credentials it carries. Are they overly permissive?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You cannot trust the AI to limit itself. You have to enforce limits at the boundary.&lt;/p&gt;
&lt;p&gt;The most dangerous sessions are the ones where you trust the AI completely.&lt;/p&gt;</content></entry><entry><title>Rio 0.7.1 released</title><id>https://julin.ai/2026/10/05/rio-0-7-1-released/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/10/05/rio-0-7-1-released/"/><published>2026-10-05T00:00:00+13:00</published><updated>2026-10-05T00:00:00+13:00</updated><category term="rio"/><category term="releases"/><category term="agents"/><content type="html">&lt;p&gt;&lt;a href="https://pypi.org/project/rio/"&gt;Rio 0.7.1&lt;/a&gt; is out. This is a big change to Rio&amp;rsquo;s core.&lt;/p&gt;
&lt;p&gt;Thanks to Rulin Shao&amp;rsquo;s work on context management. There is a new way for managing context. Rio took it one step further: instead of treating context as a file to edit, treat it as a mutable, executable document. Jupyter notebook is picked because it already has a stable specification and a runtime.&lt;/p&gt;
&lt;blockquote class="twitter-tweet"&gt;&lt;p lang="en" dir="ltr"&gt;@RulinShao Nice work. I fit CLM in for coding agent Rio:
1. LLM context is .ipynb.
2. LLM changes a cell in each turn. - &lt;code&gt;codemode&lt;/code&gt; natively supported.
3. Instead of CLM's file updates, Rio uses json patch inspired by SKILL.state.
&lt;p&gt;Fewer turns, less tokens
&lt;a href="https://t.co/YYJDYiSmpl"&gt;&lt;a href="https://t.co/YYJDYiSmpl"&gt;https://t.co/YYJDYiSmpl&lt;/a&gt;&lt;/a&gt; &lt;a href="https://t.co/SyxBwPF95N"&gt;&lt;a href="https://t.co/SyxBwPF95N"&gt;https://t.co/SyxBwPF95N&lt;/a&gt;&lt;/a&gt;&lt;/p&gt;— soasme (@soasme) &lt;a href="https://x.com/soasme/status/2106201096771522699?ref_src=twsrc%5Etfw"&gt;October 3, 2026&lt;/a&gt;&lt;/blockquote&gt; &lt;script async src="https://platform.x.com/widgets.js" charset="utf-8"&gt;&lt;/script&gt;&lt;/p&gt;</content></entry><entry><title>Agent Codemode Explained</title><id>https://julin.ai/2026/10/01/agent-codemode-explained/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/10/01/agent-codemode-explained/"/><published>2026-10-01T00:00:00+13:00</published><updated>2026-10-01T00:00:00+13:00</updated><category term="agents"/><category term="codemode"/><category term="tools"/><category term="architecture"/><content type="html">&lt;h1 id="code-mode-moving-agent-control-flow-out-of-the-llm"&gt;Code Mode: Moving Agent Control Flow Out of the LLM&lt;/h1&gt;
&lt;p&gt;Most agents still use tools one call at a time.&lt;/p&gt;
&lt;p&gt;The model decides to call a tool. The runtime executes it. The result is appended to the conversation. The model runs again and decides what to do next.&lt;/p&gt;
&lt;p&gt;For a short task this works well.&lt;/p&gt;
&lt;p&gt;For a task that needs twenty API calls, filtering, retries, a loop, or several dependent operations, the design becomes expensive:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool A
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool B
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool C
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Each edge through the LLM costs inference time. Tool results also accumulate in the context even when most of the data is only needed to decide the next call.&lt;/p&gt;
&lt;p&gt;Code Mode changes this boundary.&lt;/p&gt;
&lt;p&gt;Instead of asking the model to emit every tool call, give it one tool that executes a small program:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;program
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├─ tool A
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├─ tool B
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├─ tool C
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └─ local filtering
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The program becomes the control plane for a section of the agent run.&lt;/p&gt;
&lt;p&gt;This idea now appears in Cloudflare Code Mode, Anthropic&amp;rsquo;s programmatic tool calling, PydanticAI Code Mode, Hugging Face smolagents, and, as of September 29, 2026, Pi.&lt;/p&gt;
&lt;p&gt;They share the same basic observation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Tool calling is an unusually bad programming language.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;JSON tool calls can express:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call X with Y
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A normal programming language can express:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call X for every Y
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call X and Z concurrently
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;retry X if it fails
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;filter the result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call Z only if condition Q is true
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;store this value and reuse it later
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;return only these five fields
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The difference is larger than syntax.&lt;/p&gt;
&lt;p&gt;Code Mode changes where computation happens.&lt;/p&gt;
&lt;h2 id="the-normal-tool-loop"&gt;The normal tool loop&lt;/h2&gt;
&lt;p&gt;Suppose an agent needs to find all unhealthy services and inspect their recent deployments.&lt;/p&gt;
&lt;p&gt;With normal tool calling it might do:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → list_services()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM ← 300 services
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → get_health(service_1)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM ← result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → get_health(service_2)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM ← result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A capable model can issue several calls in parallel, but the model still has to construct the calls and receive their results.&lt;/p&gt;
&lt;p&gt;A better tool might support bulk operations. But now every tool designer has to predict every future composition:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;get_health_for_services(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;get_unhealthy_services(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;get_unhealthy_services_with_deployments(...)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is where Code Mode becomes interesting.&lt;/p&gt;
&lt;p&gt;The model can instead write something equivalent to:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;services &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; list_services()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;health &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; gather(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;*&lt;/span&gt;[get_health({&lt;span style="color:#e6db74"&gt;&amp;#34;id&amp;#34;&lt;/span&gt;: s[&lt;span style="color:#e6db74"&gt;&amp;#34;id&amp;#34;&lt;/span&gt;]}) &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; s &lt;span style="color:#f92672"&gt;in&lt;/span&gt; services]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;bad &lt;span style="color:#f92672"&gt;=&lt;/span&gt; [
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; s &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; s, h &lt;span style="color:#f92672"&gt;in&lt;/span&gt; zip(services, health)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; h[&lt;span style="color:#e6db74"&gt;&amp;#34;status&amp;#34;&lt;/span&gt;] &lt;span style="color:#f92672"&gt;!=&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;healthy&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;deployments &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; gather(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;*&lt;/span&gt;[get_deployments({&lt;span style="color:#e6db74"&gt;&amp;#34;service_id&amp;#34;&lt;/span&gt;: s[&lt;span style="color:#e6db74"&gt;&amp;#34;id&amp;#34;&lt;/span&gt;]}) &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; s &lt;span style="color:#f92672"&gt;in&lt;/span&gt; bad]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; [
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;service&amp;#34;&lt;/span&gt;: s[&lt;span style="color:#e6db74"&gt;&amp;#34;name&amp;#34;&lt;/span&gt;],
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;deployment&amp;#34;&lt;/span&gt;: d[&lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; s, d &lt;span style="color:#f92672"&gt;in&lt;/span&gt; zip(bad, deployments)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The LLM sees the task once and writes the control flow once.&lt;/p&gt;
&lt;p&gt;The intermediate 300 service records and hundreds of health responses do not need to become conversation history.&lt;/p&gt;
&lt;p&gt;That is the important part.&lt;/p&gt;
&lt;p&gt;Code Mode is not mainly a prettier way to call tools.&lt;/p&gt;
&lt;p&gt;It is a way to move &lt;strong&gt;intermediate computation out of model context&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="cloudflare-code-as-a-compact-api-plan"&gt;Cloudflare: code as a compact API plan&lt;/h2&gt;
&lt;p&gt;Cloudflare has probably made the clearest argument for Code Mode.&lt;/p&gt;
&lt;p&gt;Its problem is unusually visible. The Cloudflare API has more than 2,500 endpoints. Exposing every API endpoint as an MCP tool means shipping a huge collection of JSON schemas into the model context before the agent does anything useful.&lt;/p&gt;
&lt;p&gt;Cloudflare&amp;rsquo;s Code Mode MCP server instead exposes essentially two operations: search for an API surface, then execute code against the discovered typed API.&lt;/p&gt;
&lt;p&gt;Cloudflare reports that its whole API can be exposed in roughly 1,000 tokens this way. Its comparison estimates 1.17 million input tokens if the equivalent API were represented as ordinary MCP tool definitions.&lt;/p&gt;
&lt;p&gt;This solves the first Code Mode problem:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;too many tools
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;There is a second problem:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;too much intermediate data
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Cloudflare describes generated code as a compact plan. Calls, filtering and transformations happen inside the execution environment, and only selected output returns to the model.&lt;/p&gt;
&lt;p&gt;Their current interface exposes primitives such as:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-ts" data-lang="ts"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;codemode&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;search&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;codemode&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;describe&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;codemode&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;step&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;codemode&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;run&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;search()&lt;/code&gt; performs progressive discovery. &lt;code&gt;describe()&lt;/code&gt; loads the detailed interface only when needed. &lt;code&gt;step()&lt;/code&gt; records work that needs replay semantics. &lt;code&gt;run()&lt;/code&gt; executes reusable snippets.&lt;/p&gt;
&lt;p&gt;This is an important refinement.&lt;/p&gt;
&lt;p&gt;A naive Code Mode implementation still puts every available function declaration in the system prompt:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;one code tool
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;10,000 function declarations
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That fixes intermediate results but not tool-definition cost.&lt;/p&gt;
&lt;p&gt;Cloudflare instead makes the API itself lazy.&lt;/p&gt;
&lt;p&gt;The model starts with a small discovery interface and pulls schemas when it needs them.&lt;/p&gt;
&lt;p&gt;Conceptually:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ┌──────── search
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → code│
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├──────── describe
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├──────── API A
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├──────── API B
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └──────── API C
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The tool catalog becomes something closer to a filesystem or symbol table than a prompt.&lt;/p&gt;
&lt;p&gt;That is likely the right abstraction for very large tool spaces.&lt;/p&gt;
&lt;h3 id="cloudflares-sandbox-matters"&gt;Cloudflare&amp;rsquo;s sandbox matters&lt;/h3&gt;
&lt;p&gt;Once the model can write arbitrary JavaScript, &lt;code&gt;eval()&lt;/code&gt; is not an acceptable runtime.&lt;/p&gt;
&lt;p&gt;Cloudflare executes generated code in isolated Workers. Its MCP design blocks direct outbound access by default; generated code reaches the outside world through capabilities supplied by the host. Credentials remain outside the generated program.&lt;/p&gt;
&lt;p&gt;The distinction is useful:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;code execution capability
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ≠
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;external authority
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The sandbox can be computationally expressive while having almost no ambient authority.&lt;/p&gt;
&lt;p&gt;The program may know how to write:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;dns&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;deleteZone&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;but the actual authority still belongs to the host connector.&lt;/p&gt;
&lt;p&gt;This means authorization does not have to be implemented by trusting generated code.&lt;/p&gt;
&lt;p&gt;It remains at the tool boundary.&lt;/p&gt;
&lt;p&gt;Cloudflare&amp;rsquo;s durable runtime goes further. It records execution history and can pause generated code for approval, then replay completed work and continue the program.&lt;/p&gt;
&lt;p&gt;That turns Code Mode from &amp;ldquo;run some generated JavaScript&amp;rdquo; into an agent runtime primitive.&lt;/p&gt;
&lt;h2 id="anthropic-programmatic-tool-calling"&gt;Anthropic: programmatic tool calling&lt;/h2&gt;
&lt;p&gt;Anthropic independently arrived at almost exactly the same model.&lt;/p&gt;
&lt;p&gt;Their current term is &lt;strong&gt;programmatic tool calling&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Claude writes Python inside its code execution environment. Tool calls from the program cross back to the host. When a tool returns, execution resumes. Intermediate results remain inside the execution environment instead of being inserted into Claude&amp;rsquo;s context.&lt;/p&gt;
&lt;p&gt;The execution model looks like this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Claude generates Python
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ▼
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;sandbox
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── tool()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │ │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │ └──── host executes tool
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── process result locally
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── tool()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── print small result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ▼
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Claude continues
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is not equivalent to asking Claude to generate a Python script and running the script after the model finishes.&lt;/p&gt;
&lt;p&gt;The interesting part is the suspended host call.&lt;/p&gt;
&lt;p&gt;The Python program is effectively an orchestration language around external capabilities.&lt;/p&gt;
&lt;p&gt;Anthropic reports useful measurements here. On its BrowseComp and DeepSearchQA experiments, programmatic tool calling improved performance by an average of 11% while using 24% fewer input tokens. On an internal 75-tool project-management benchmark it reports roughly 38% lower billed input tokens with unchanged task accuracy. But on τ²-bench tasks dominated by one or two sequential calls, it cost about 8% more.&lt;/p&gt;
&lt;p&gt;That last number is important.&lt;/p&gt;
&lt;p&gt;Code Mode is not free.&lt;/p&gt;
&lt;p&gt;For:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call tool
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;return answer
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;generating a program and starting a code environment is extra work.&lt;/p&gt;
&lt;p&gt;Code Mode becomes useful when the code replaces enough model-mediated control flow to pay for itself.&lt;/p&gt;
&lt;h2 id="pi-code-mode-moves-into-the-harness"&gt;Pi: Code Mode moves into the harness&lt;/h2&gt;
&lt;p&gt;Pi is a more interesting implementation for coding agents because Code Mode is not only an API optimization.&lt;/p&gt;
&lt;p&gt;It is integrated into the agent harness.&lt;/p&gt;
&lt;p&gt;A large Pi change was merged on September 29, 2026. The implementation runs model-generated JavaScript in a QuickJS VM compiled to WebAssembly. The script can call Pi tools as async functions. Nested calls do not enter the LLM context individually; only explicit script output and the final return value do.&lt;/p&gt;
&lt;p&gt;The standalone package describes the capability directly:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;source&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;tools&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;read&lt;/span&gt;({ &lt;span style="color:#a6e22e"&gt;path&lt;/span&gt;&lt;span style="color:#f92672"&gt;:&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;package.json&amp;#34;&lt;/span&gt; })
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;text&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;JSON&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;parse&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;source&lt;/span&gt;).&lt;span style="color:#a6e22e"&gt;name&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Every execution gets a QuickJS VM in a worker. The VM has its own WASM memory and no general route back into the Node host except the explicitly installed host-call bridge.&lt;/p&gt;
&lt;p&gt;This makes Pi&amp;rsquo;s implementation worth studying.&lt;/p&gt;
&lt;p&gt;Pi already had very powerful tools:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;read
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;write
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;edit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;bash
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;For coding work, &lt;code&gt;bash&lt;/code&gt; alone is close to universal.&lt;/p&gt;
&lt;p&gt;So why add Code Mode?&lt;/p&gt;
&lt;p&gt;Because universality is not the same as a good agent interface.&lt;/p&gt;
&lt;p&gt;Consider:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat package.json | jq ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;versus:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;p&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;tools&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;read&lt;/span&gt;({&lt;span style="color:#a6e22e"&gt;path&lt;/span&gt;&lt;span style="color:#f92672"&gt;:&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;package.json&amp;#34;&lt;/span&gt;})
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;pkg&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;JSON&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;parse&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;p&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both work.&lt;/p&gt;
&lt;p&gt;But the second form can combine normal agent tools, MCP tools and future host capabilities under one runtime, without turning every operation into shell text.&lt;/p&gt;
&lt;p&gt;More importantly, nested calls remain visible to the harness.&lt;/p&gt;
&lt;p&gt;Pi routes nested calls through the normal validation and tool hooks. They carry a &lt;code&gt;parentToolCallId&lt;/code&gt;, and nested calls are recorded on the parent tool result.&lt;/p&gt;
&lt;p&gt;This gives a useful structure:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;codemode call
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;│
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;├── read
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;├── grep
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;├── MCP: search
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;└── bash
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;instead of hiding the entire operation behind an opaque process.&lt;/p&gt;
&lt;p&gt;That matters for observability, permissions and evaluation.&lt;/p&gt;
&lt;h3 id="pis-tool-exposure-model"&gt;Pi&amp;rsquo;s tool exposure model&lt;/h3&gt;
&lt;p&gt;Pi also exposes a design problem that smaller Code Mode demos can avoid:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Which tools should the model see directly?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Pi now distinguishes different forms of tool exposure and supports two useful Code Mode configurations.&lt;/p&gt;
&lt;p&gt;In &lt;code&gt;codemode.mode = &amp;quot;on&amp;quot;&lt;/code&gt;, normal tools may remain directly visible while Code Mode can also call them.&lt;/p&gt;
&lt;p&gt;In &lt;code&gt;codemode.mode = &amp;quot;only&amp;quot;&lt;/code&gt;, the model sees Code Mode as the primary interface and calls ordinary tools through generated code instead.&lt;/p&gt;
&lt;p&gt;This gives two different agent architectures.&lt;/p&gt;
&lt;h4 id="hybrid"&gt;Hybrid&lt;/h4&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── read
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── edit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── bash
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── codemode
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── read
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── edit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── MCP tools
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h4 id="code-only"&gt;Code-only&lt;/h4&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── codemode
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── read
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── edit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── bash
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── MCP
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── extensions
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The second one is more radical.&lt;/p&gt;
&lt;p&gt;The model-facing action language stops being a collection of JSON tools.&lt;/p&gt;
&lt;p&gt;It becomes JavaScript.&lt;/p&gt;
&lt;p&gt;Pi also places a token budget on inline declarations. Tools that do not fit can be found dynamically through tool search.&lt;/p&gt;
&lt;p&gt;So Pi is converging on the same two-layer structure as Cloudflare:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;small language runtime
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;lazy capability discovery
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This looks increasingly less like an agent feature and more like an operating-system interface.&lt;/p&gt;
&lt;h2 id="pydanticai-python-as-the-capability-language"&gt;PydanticAI: Python as the capability language&lt;/h2&gt;
&lt;p&gt;PydanticAI implements the same pattern using Python and Monty.&lt;/p&gt;
&lt;p&gt;Eligible tools disappear from the direct model-facing tool list and become functions available inside one &lt;code&gt;run_code&lt;/code&gt; tool.&lt;/p&gt;
&lt;p&gt;The generated Python can use loops, conditions, variables and &lt;code&gt;asyncio.gather&lt;/code&gt;. Intermediate data stays in the sandbox.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;rows &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; search(&lt;span style="color:#e6db74"&gt;&amp;#34;failed builds&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;details &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; asyncio&lt;span style="color:#f92672"&gt;.&lt;/span&gt;gather(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;*&lt;/span&gt;[get_build(row[&lt;span style="color:#e6db74"&gt;&amp;#34;id&amp;#34;&lt;/span&gt;]) &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; row &lt;span style="color:#f92672"&gt;in&lt;/span&gt; rows]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; [
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; x &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; x &lt;span style="color:#f92672"&gt;in&lt;/span&gt; details
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; x[&lt;span style="color:#e6db74"&gt;&amp;#34;status&amp;#34;&lt;/span&gt;] &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;failed&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;PydanticAI&amp;rsquo;s implementation is particularly interesting because Monty is not CPython running in a container.&lt;/p&gt;
&lt;p&gt;Monty is a restricted Python interpreter written in Rust for AI-generated programs.&lt;/p&gt;
&lt;p&gt;By default it has no normal filesystem, environment variables or network access. Host capabilities have to be passed in deliberately.&lt;/p&gt;
&lt;p&gt;That gives a capability model similar to Pi&amp;rsquo;s QuickJS sandbox:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;generated program
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── computation
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── variables
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── branching
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── injected host functions
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The model gets the useful parts of Python without automatically inheriting the authority of the host Python process.&lt;/p&gt;
&lt;p&gt;PydanticAI also supports persistent REPL state during an agent run. Variables, imports and helper functions can survive across &lt;code&gt;run_code&lt;/code&gt; calls.&lt;/p&gt;
&lt;p&gt;That changes Code Mode again.&lt;/p&gt;
&lt;p&gt;It is no longer just:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → program → result
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It can become:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → persistent computational workspace
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That distinction will matter for long-running agents.&lt;/p&gt;
&lt;h2 id="smolagents-code-as-the-action-format"&gt;smolagents: code as the action format&lt;/h2&gt;
&lt;p&gt;Hugging Face&amp;rsquo;s smolagents predates some of these systems and takes the idea further at the agent-loop level.&lt;/p&gt;
&lt;p&gt;Its &lt;code&gt;CodeAgent&lt;/code&gt; uses Python code as the normal action representation. &lt;code&gt;ToolCallingAgent&lt;/code&gt; is the alternative that emits conventional structured tool calls.&lt;/p&gt;
&lt;p&gt;The important observation from smolagents is simple:&lt;/p&gt;
&lt;p&gt;JSON is a serialization format.&lt;/p&gt;
&lt;p&gt;A programming language is a control-flow language.&lt;/p&gt;
&lt;p&gt;If the task is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search A
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search B
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;take the intersection
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;fetch each result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;rank them
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;representing that as a program is natural.&lt;/p&gt;
&lt;p&gt;Representing it as five separate JSON actions requires the LLM itself to become the control-flow interpreter.&lt;/p&gt;
&lt;p&gt;That is the hidden cost of ordinary agent loops.&lt;/p&gt;
&lt;h2 id="code-mode-is-a-tiny-compiler-boundary"&gt;Code Mode is a tiny compiler boundary&lt;/h2&gt;
&lt;p&gt;One way to understand these systems is to stop thinking about &amp;ldquo;code execution.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Think of the LLM as producing a program for a small virtual machine.&lt;/p&gt;
&lt;p&gt;The traditional agent loop is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observation
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;one action
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observation
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Code Mode becomes:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observation
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;program
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;runtime
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;many actions
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;compressed observation
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The LLM has effectively compiled several future decisions into one artifact.&lt;/p&gt;
&lt;p&gt;This means Code Mode is useful only when those future decisions can be expressed algorithmically.&lt;/p&gt;
&lt;p&gt;Suppose a result requires semantic judgment:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Read this incident report.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Understand what probably caused the outage.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Choose what to investigate next.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A JavaScript &lt;code&gt;if&lt;/code&gt; statement cannot replace the next LLM call.&lt;/p&gt;
&lt;p&gt;But many agent decisions are not semantic reasoning:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;for every item
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;if status == failed
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;take first five
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;sort by date
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;retry twice
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call these concurrently
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;extract this field
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Today we often waste LLM calls performing those operations.&lt;/p&gt;
&lt;p&gt;Code Mode removes them from the model loop.&lt;/p&gt;
&lt;h2 id="the-real-token-saving-has-two-parts"&gt;The real token saving has two parts&lt;/h2&gt;
&lt;p&gt;Discussions of Code Mode often combine two independent optimizations.&lt;/p&gt;
&lt;p&gt;They should be separated.&lt;/p&gt;
&lt;h3 id="1-tool-definition-compression"&gt;1. Tool definition compression&lt;/h3&gt;
&lt;p&gt;Instead of:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool A schema
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool B schema
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool C schema
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;...
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool 2500 schema
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;give the model:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search(query)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;describe(symbol)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;execute(code)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Then discover interfaces lazily.&lt;/p&gt;
&lt;p&gt;This is the part that lets Cloudflare expose thousands of endpoints in a small fixed context.&lt;/p&gt;
&lt;p&gt;Anthropic described a similar MCP approach where tool definitions are discovered through files instead of loading the whole MCP catalog into context. In one example, it reduced tool-related context from roughly 150,000 tokens to about 2,000.&lt;/p&gt;
&lt;h3 id="2-intermediate-result-compression"&gt;2. Intermediate result compression&lt;/h3&gt;
&lt;p&gt;Instead of:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ tool result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ tool result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ LLM
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;do:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ program
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ tool result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ program
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ filter
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ final small result
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ LLM
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;These solve different scaling problems.&lt;/p&gt;
&lt;p&gt;A serious Code Mode implementation should probably do both.&lt;/p&gt;
&lt;h2 id="code-mode-does-not-remove-the-agent-loop"&gt;Code Mode does not remove the agent loop&lt;/h2&gt;
&lt;p&gt;A tempting conclusion is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Why not put the entire agent in code?
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Because generated code only knows what the model knew when it wrote the program.&lt;/p&gt;
&lt;p&gt;Consider:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;result &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; investigate()
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If &lt;code&gt;result&lt;/code&gt; contains something genuinely surprising, deciding what it means may require another model inference.&lt;/p&gt;
&lt;p&gt;A normal program can branch on structure:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; result[&lt;span style="color:#e6db74"&gt;&amp;#34;status&amp;#34;&lt;/span&gt;] &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;failed&amp;#34;&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It cannot reliably branch on arbitrary semantics:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; result implies that the database migration
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;is&lt;/span&gt; probably unrelated to the outage:
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;You can put another model behind a function call:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; classify(result, question):
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;but now the model has returned through another door.&lt;/p&gt;
&lt;p&gt;So the likely architecture is not:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Code Mode replaces Agent Loop
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Agent Loop
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── semantic decision
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── Code Mode
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── deterministic control flow
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── data processing
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── bulk tool calls
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── cheap decisions
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The boundary between these two layers is one of the more interesting agent-design problems.&lt;/p&gt;
&lt;h2 id="code-mode-also-changes-tool-design"&gt;Code Mode also changes tool design&lt;/h2&gt;
&lt;p&gt;Today&amp;rsquo;s tool APIs are partly shaped by the weakness of model tool calling.&lt;/p&gt;
&lt;p&gt;We create high-level tools such as:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search_and_summarize_incidents
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;find_recent_failed_deployments
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;get_customer_context
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;because asking the model to compose eight primitive tools is expensive and unreliable.&lt;/p&gt;
&lt;p&gt;Once tools can be composed inside code, smaller primitives become practical:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;fetch
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;read
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;write
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;query
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;execute
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The generated program provides the composition.&lt;/p&gt;
&lt;p&gt;This does not mean every tool should become low level.&lt;/p&gt;
&lt;p&gt;Every tool boundary is also a:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;permission boundary
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;validation boundary
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observability boundary
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;failure boundary
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;But Code Mode changes the trade-off.&lt;/p&gt;
&lt;p&gt;Without Code Mode, coarse tools save LLM turns.&lt;/p&gt;
&lt;p&gt;With Code Mode, coarse tools need to justify themselves for another reason.&lt;/p&gt;
&lt;h2 id="the-security-boundary-is-not-the-language-sandbox"&gt;The security boundary is not the language sandbox&lt;/h2&gt;
&lt;p&gt;Running generated code in QuickJS, Monty or a Dynamic Worker is necessary.&lt;/p&gt;
&lt;p&gt;It is not sufficient.&lt;/p&gt;
&lt;p&gt;The dangerous operation is usually not:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;while&lt;/span&gt; (&lt;span style="color:#66d9ef"&gt;true&lt;/span&gt;) {}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;tools&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;deleteDatabase&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The generated program should therefore be treated as untrusted orchestration around trusted capabilities.&lt;/p&gt;
&lt;p&gt;The useful architecture is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; generated code
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; sandbox boundary
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ▼
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; capability bridge
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ┌──────────────┼──────────────┐
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ▼ ▼ ▼
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; read tool GitHub tool DB tool
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │ │ │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; policy policy policy
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │ │ │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; execute execute execute
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Permissions belong below the code.&lt;/p&gt;
&lt;p&gt;Pi&amp;rsquo;s nested calls pass through normal tool validation and hooks.&lt;/p&gt;
&lt;p&gt;Cloudflare explicitly says generated code does not replace authorization and that permissions should still be enforced in upstream handlers.&lt;/p&gt;
&lt;p&gt;Anthropic similarly warns that its tool caller configuration should not itself be treated as a security boundary.&lt;/p&gt;
&lt;p&gt;These implementations are converging on the same rule:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;sandbox computation; authorize capabilities.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="observability-becomes-harder"&gt;Observability becomes harder&lt;/h2&gt;
&lt;p&gt;Direct tool calling gives a convenient trace:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Code Mode compresses the model trace:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;codemode
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is good for tokens.&lt;/p&gt;
&lt;p&gt;It can be bad for debugging if the runtime treats the code invocation as one opaque tool call.&lt;/p&gt;
&lt;p&gt;A useful Code Mode runtime therefore needs nested traces:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;codemode #42
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;├─ search #43
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;│ └─ 42 ms
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;├─ fetch #44
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;│ └─ 181 ms
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;├─ fetch #45
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;│ └─ 172 ms
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;└─ update #46
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └─ approval required
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Pi explicitly records nested calls and their parent IDs.&lt;/p&gt;
&lt;p&gt;Cloudflare&amp;rsquo;s durable runtime records steps and replay information.&lt;/p&gt;
&lt;p&gt;PydanticAI exposes nested Code Mode calls in its tracing and can integrate them with durable execution systems.&lt;/p&gt;
&lt;p&gt;This is not optional infrastructure.&lt;/p&gt;
&lt;p&gt;Once a generated program can call fifty tools, &amp;ldquo;the Code Mode tool failed&amp;rdquo; is not useful telemetry.&lt;/p&gt;
&lt;h2 id="code-mode-introduces-new-failure-modes"&gt;Code Mode introduces new failure modes&lt;/h2&gt;
&lt;p&gt;It removes some agent failures and adds others.&lt;/p&gt;
&lt;p&gt;Normal tool calling can fail because the model:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;forgets a step
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;calls a tool repeatedly
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;loads too much data
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;serializes work unnecessarily
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;loses intermediate state
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Code Mode can reduce those.&lt;/p&gt;
&lt;p&gt;But now the generated program can have:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;syntax errors
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;type errors
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;infinite loops
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;unbounded fan-out
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;bad retry logic
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;stale state
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;race conditions
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;incorrect aggregation
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;side effects inside loops
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A serious runtime needs limits.&lt;/p&gt;
&lt;p&gt;At minimum:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;execution timeout
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;memory limit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;maximum tool calls
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;concurrency limit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;output limit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;nesting limit
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cancellation
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Pi already places deadlines and output limits around its sandbox, and large outputs can be spilled rather than pushed back into context.&lt;/p&gt;
&lt;p&gt;PydanticAI&amp;rsquo;s Monty environment intentionally exposes a restricted subset of Python and limits host access.&lt;/p&gt;
&lt;p&gt;This suggests another useful rule:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Code Mode should be powerful enough for orchestration, not powerful enough to become an accidental operating system.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="javascript-or-python"&gt;JavaScript or Python?&lt;/h2&gt;
&lt;p&gt;The current implementations split here.&lt;/p&gt;
&lt;p&gt;Cloudflare and Pi use JavaScript.&lt;/p&gt;
&lt;p&gt;Anthropic, PydanticAI and smolagents use Python.&lt;/p&gt;
&lt;p&gt;This is less important than it first appears.&lt;/p&gt;
&lt;p&gt;The generated language needs:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;variables
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;arrays/maps
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;loops
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;conditionals
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;functions
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;async calls
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;parallel calls
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;basic data transformation
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both languages satisfy that.&lt;/p&gt;
&lt;p&gt;The more important properties are:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;How cheap is sandbox startup?
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;What host capabilities exist?
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Can execution be interrupted?
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Can state persist?
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Can nested calls be traced?
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Can calls pause for approval?
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Can programs be replayed?
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Can tool types be generated cleanly?
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Pydantic&amp;rsquo;s Monty is interesting precisely because it treats this as a runtime problem rather than &amp;ldquo;just execute Python.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Pi does the same with QuickJS in WASM.&lt;/p&gt;
&lt;p&gt;The language is the visible part.&lt;/p&gt;
&lt;p&gt;The runtime is the product.&lt;/p&gt;
&lt;h2 id="a-useful-comparison"&gt;A useful comparison&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Generated language&lt;/th&gt;
&lt;th&gt;Main sandbox&lt;/th&gt;
&lt;th&gt;Tool discovery&lt;/th&gt;
&lt;th&gt;Persistent state&lt;/th&gt;
&lt;th&gt;Main emphasis&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cloudflare Code Mode&lt;/td&gt;
&lt;td&gt;JavaScript / TypeScript-shaped API&lt;/td&gt;
&lt;td&gt;Workers / isolated execution&lt;/td&gt;
&lt;td&gt;&lt;code&gt;search()&lt;/code&gt; + &lt;code&gt;describe()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Durable runtime / snippets&lt;/td&gt;
&lt;td&gt;Huge API surfaces, MCP, durable execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pi Codemode&lt;/td&gt;
&lt;td&gt;JavaScript&lt;/td&gt;
&lt;td&gt;QuickJS WASM in worker&lt;/td&gt;
&lt;td&gt;Inline declarations + tool search&lt;/td&gt;
&lt;td&gt;Session store&lt;/td&gt;
&lt;td&gt;Coding-agent harness integration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic Programmatic Tool Calling&lt;/td&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;Anthropic code execution container&lt;/td&gt;
&lt;td&gt;Provider tool definitions / related advanced tool use&lt;/td&gt;
&lt;td&gt;Container reuse&lt;/td&gt;
&lt;td&gt;Managed programmatic calling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PydanticAI Code Mode&lt;/td&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;Monty&lt;/td&gt;
&lt;td&gt;Tool Search integration&lt;/td&gt;
&lt;td&gt;Persistent REPL state&lt;/td&gt;
&lt;td&gt;Safe embeddable runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;smolagents CodeAgent&lt;/td&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;Local or configured sandbox&lt;/td&gt;
&lt;td&gt;Agent tool set&lt;/td&gt;
&lt;td&gt;Agent execution state&lt;/td&gt;
&lt;td&gt;Code as primary action representation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;They are different implementations of roughly the same abstraction:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;small program
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;capability runtime
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tools
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="what-i-think-code-mode-actually-is"&gt;What I think Code Mode actually is&lt;/h2&gt;
&lt;p&gt;Calling this feature &amp;ldquo;Code Mode&amp;rdquo; makes it sound optional:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;normal agent
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;sometimes let it execute code
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;I think the deeper interpretation is different.&lt;/p&gt;
&lt;p&gt;Normal tool calling asks an LLM to act as both:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;reasoning engine
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;and
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;workflow interpreter
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Code Mode separates them.&lt;/p&gt;
&lt;p&gt;The LLM produces a temporary program.&lt;/p&gt;
&lt;p&gt;The runtime executes it.&lt;/p&gt;
&lt;p&gt;Then the LLM returns when semantic reasoning is needed again.&lt;/p&gt;
&lt;p&gt;That looks much closer to a compiler architecture:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;intent
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;temporary program / IR
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;runtime
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observations
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The program does not need to survive forever.&lt;/p&gt;
&lt;p&gt;It can exist for 200 milliseconds.&lt;/p&gt;
&lt;p&gt;It can be generated specifically for one observation.&lt;/p&gt;
&lt;p&gt;It can call five tools and disappear.&lt;/p&gt;
&lt;p&gt;This makes it a useful intermediate representation between language models and tools.&lt;/p&gt;
&lt;h2 id="the-interesting-future-is-not-bigger-tool-calls"&gt;The interesting future is not bigger tool calls&lt;/h2&gt;
&lt;p&gt;There is a natural progression here.&lt;/p&gt;
&lt;h3 id="generation-1"&gt;Generation 1&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → text
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="generation-2"&gt;Generation 2&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → JSON tool call
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="generation-3"&gt;Generation 3&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM → code → many tool calls
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The next step is probably not simply more code.&lt;/p&gt;
&lt;p&gt;The runtime can start providing primitives that are cheaper or safer than asking the main model again:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;classify(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;validate(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;wait(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;watch(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;parallel(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;approve(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;spawn(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;store(...)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Pi already exposes model classifiers from its Code Mode environment. Its merged implementation allows generated programs to inspect the model catalog and invoke classification operations.&lt;/p&gt;
&lt;p&gt;This is interesting because the program no longer orchestrates only tools.&lt;/p&gt;
&lt;p&gt;It can orchestrate computation at different intelligence levels.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;issues&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;tools&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;github&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;searchIssues&lt;/span&gt;({ &lt;span style="color:#a6e22e"&gt;repo&lt;/span&gt; })
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;relevant&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; Promise.&lt;span style="color:#a6e22e"&gt;all&lt;/span&gt;(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#a6e22e"&gt;issues&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;map&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;issue&lt;/span&gt; =&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#a6e22e"&gt;models&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;classify&lt;/span&gt;(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;Is this issue related to authentication?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#a6e22e"&gt;issue&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;body&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; )
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; )
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;issues&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;filter&lt;/span&gt;((&lt;span style="color:#a6e22e"&gt;_&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;i&lt;/span&gt;) =&amp;gt; &lt;span style="color:#a6e22e"&gt;relevant&lt;/span&gt;[&lt;span style="color:#a6e22e"&gt;i&lt;/span&gt;])
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The expensive model writes the algorithm once.&lt;/p&gt;
&lt;p&gt;A cheaper decision mechanism executes inside the loop.&lt;/p&gt;
&lt;p&gt;Now the architecture becomes:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; main LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ▼
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; code
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ┌────────┼─────────┐
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ▼ ▼ ▼
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; tools small AI local compute
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That starts to look like a real computational runtime for agents.&lt;/p&gt;
&lt;h2 id="when-i-would-use-code-mode"&gt;When I would use Code Mode&lt;/h2&gt;
&lt;p&gt;Code Mode is a good fit when the task contains:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;fan-out
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;loops
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;large intermediate results
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;filtering
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;aggregation
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;parallel calls
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;large tool catalogs
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;reusable procedures
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;mostly deterministic control flow
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Examples:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;check 100 services and return the failed ones
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search ten sources and deduplicate results
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;read every changed file and run a validator
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;query several APIs and join their results
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;inspect every failed CI run
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;find all matching records, update a subset
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;I would keep direct tool calls for:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;open this file
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;run this command
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ask the user
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;make one API request
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;perform one dangerous action requiring approval
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Anthropic&amp;rsquo;s own measurements support this distinction: workloads with many calls and large intermediate results benefit; workloads with one or two short sequential calls may not.&lt;/p&gt;
&lt;p&gt;So I would not make Code Mode a universal replacement for tools.&lt;/p&gt;
&lt;p&gt;I would make it a first-class execution path beside them.&lt;/p&gt;
&lt;h2 id="the-design-i-would-build"&gt;The design I would build&lt;/h2&gt;
&lt;p&gt;A minimal Code Mode implementation only needs:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;one sandbox
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;one code tool
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;a bridge to existing tools
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A production implementation needs more.&lt;/p&gt;
&lt;p&gt;I would want:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;1. Capability-based sandbox
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;2. Typed tool bindings
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;3. Progressive tool discovery
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;4. Direct and code-only tool exposure
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;5. Nested tool-call tracing
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;6. Per-call policy and approval
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;7. Tool-call concurrency limits
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;8. Runtime deadline
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;9. Output budget
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;10. Persistent session-local state
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;11. Cancellation
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;12. Optional durable execution
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The central API can remain very small:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;tools&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;search&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;tools&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;describe&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;tools&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;foo&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;store&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;load&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;text&lt;/span&gt;(...)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Everything else belongs in the harness.&lt;/p&gt;
&lt;p&gt;The generated program should not know where credentials live, how permissions work, how traces are exported, or whether a tool is implemented through MCP, HTTP, a shell process or another agent.&lt;/p&gt;
&lt;p&gt;That is the host&amp;rsquo;s job.&lt;/p&gt;
&lt;h2 id="one-consequence-for-mcp"&gt;One consequence for MCP&lt;/h2&gt;
&lt;p&gt;Code Mode also changes how I think about MCP.&lt;/p&gt;
&lt;p&gt;MCP is useful as a transport and capability description protocol.&lt;/p&gt;
&lt;p&gt;It is much less convincing as the language an LLM should directly program against.&lt;/p&gt;
&lt;p&gt;Exposing hundreds of MCP tools directly to the model couples:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;capability transport
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;with:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model action representation
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Those do not need to be the same thing.&lt;/p&gt;
&lt;p&gt;Cloudflare makes this explicit: MCP can remain underneath Code Mode. An MCP server may itself expose a Code Mode interface, or existing MCP tools can become functions callable by generated code.&lt;/p&gt;
&lt;p&gt;Pi has now made a similar separation. MCP tools can be available through Codemode or discovered and exposed directly depending on configuration.&lt;/p&gt;
&lt;p&gt;A cleaner stack is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Code Mode
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tool abstraction
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MCP / HTTP / local functions / shell / another agent
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;MCP becomes infrastructure.&lt;/p&gt;
&lt;p&gt;It stops consuming the entire model interface.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Code Mode looks like a token optimization at first.&lt;/p&gt;
&lt;p&gt;It is more useful to see it as a change in the execution model.&lt;/p&gt;
&lt;p&gt;Traditional agents repeatedly ask the LLM:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;What should I do next?
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;even when &amp;ldquo;next&amp;rdquo; is determined by a loop, a filter or a boolean condition.&lt;/p&gt;
&lt;p&gt;Code Mode lets the model answer a larger question:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;What program should control the next section of execution?
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Then the harness runs that program against a restricted set of capabilities.&lt;/p&gt;
&lt;p&gt;This reduces model round trips.&lt;/p&gt;
&lt;p&gt;It keeps intermediate data out of context.&lt;/p&gt;
&lt;p&gt;It makes huge tool catalogs practical.&lt;/p&gt;
&lt;p&gt;It also creates a new runtime layer that needs sandboxing, permissions, tracing, resource limits and state.&lt;/p&gt;
&lt;p&gt;Cloudflare approaches the problem from large APIs.&lt;/p&gt;
&lt;p&gt;Anthropic approaches it from model-native programmatic tool calling.&lt;/p&gt;
&lt;p&gt;PydanticAI approaches it from a safe embedded Python runtime.&lt;/p&gt;
&lt;p&gt;Pi has now integrated the pattern directly into a coding-agent harness.&lt;/p&gt;
&lt;p&gt;The implementations differ.&lt;/p&gt;
&lt;p&gt;The architecture is converging:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM for semantic decisions.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Code for control flow.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Tools for authority.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Harness for policy and state.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That separation is more important than the name &amp;ldquo;Code Mode.&amp;rdquo;&lt;/p&gt;</content></entry><entry><title>Where Does a Claude Agent Run?</title><id>https://julin.ai/2026/09/29/claude-agent-runtime/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/29/claude-agent-runtime/"/><published>2026-09-29T00:00:00+13:00</published><updated>2026-09-29T00:00:00+13:00</updated><category term="field-notes"/><category term="agents"/><category term="infrastructure"/><content type="html">&lt;p&gt;Different Claude tools run agents in different places. The choice shapes what the agent can do.&lt;/p&gt;
&lt;h2 id="claude-claudeai"&gt;Claude (claude.ai)&lt;/h2&gt;
&lt;p&gt;The agent runs inside a gVisor container on the server side. No local setup required—everything runs remotely. The container is ephemeral; each session starts fresh and leaves no traces.&lt;/p&gt;
&lt;h2 id="claude-code"&gt;Claude Code&lt;/h2&gt;
&lt;p&gt;The agent runs on your machine. It can execute every command and access every file that your logged-in user can access. For advanced users concerned about blast radius, start a sandbox first (like a microVM) and run Claude Code inside that isolated environment.&lt;/p&gt;
&lt;h2 id="claude-cowork"&gt;Claude Cowork&lt;/h2&gt;
&lt;p&gt;The agent runs on a VM. On macOS, it uses the Virtualization framework; on Windows, it uses HCS. The VM has its own Linux kernel and filesystem. But the agent loop itself runs outside the VM—only code execution happens inside.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Each design is a tradeoff. Server-side isolation is safe but ephemeral. Local execution is powerful but risky. VM-based execution gives you isolation with some local access.&lt;/p&gt;</content></entry><entry><title>The 100 Ways to Extend AI Agents (Well, More Than Zero)</title><id>https://julin.ai/2026/09/26/extend-ai-agents/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/26/extend-ai-agents/"/><published>2026-09-26T00:00:00+12:00</published><updated>2026-09-26T00:00:00+12:00</updated><category term="agents"/><category term="extensions"/><category term="tooling"/><content type="html">&lt;p&gt;I lied. This article doesn&amp;rsquo;t show you 100 ways to extend AI agents. But there are definitely more ways than most people realize. Here are the ones I&amp;rsquo;ve discovered.&lt;/p&gt;
&lt;h2 id="1-agentsmd"&gt;1. AGENTS.md&lt;/h2&gt;
&lt;p&gt;Harnesses auto-load this file when a session starts. It&amp;rsquo;s where you can change agent behavior by writing markdown documentation. Claude Code previously only supported CLAUDE.md, but they&amp;rsquo;ve recently added AGENTS.md support. (Thanks Shopify CEO Tobi.)&lt;/p&gt;
&lt;h2 id="2-skills"&gt;2. Skills&lt;/h2&gt;
&lt;p&gt;The LLM receives the name and description from skill frontmatter in your installed SKILL.md files—either globally or in the project directory (.agents/skills, .&lt;whateveragent&gt;/skills). The LLM progressively requests full skill content as needed, revealing more of the skill in the prompt dynamically.&lt;/p&gt;
&lt;h2 id="3-prompt-templates"&gt;3. Prompt Templates&lt;/h2&gt;
&lt;p&gt;Save commonly used prompts as templates and invoke them like &lt;code&gt;/review #8&lt;/code&gt; to review PR #8. This lets you template complex workflows without rebuilding them each time.&lt;/p&gt;
&lt;h2 id="4-tools-and-mcp"&gt;4. Tools and MCP&lt;/h2&gt;
&lt;p&gt;MCP is a protocol supported by most AI agents. It extends agent capabilities to do more—reading and writing files locally, interacting with external systems, and automating tasks that would otherwise be manual.&lt;/p&gt;
&lt;h2 id="5-extensions-and-plugins"&gt;5. Extensions and Plugins&lt;/h2&gt;
&lt;p&gt;Some agents support extensions or plugins at runtime (like Pi). These are agent-specific and only work with that particular agent.&lt;/p&gt;
&lt;h2 id="6-hooks"&gt;6. Hooks&lt;/h2&gt;
&lt;p&gt;Define hooks before running tool calls or commands to intercept and change behavior. RTK is a great example—it intercepts tool calls for token efficiency.&lt;/p&gt;
&lt;h2 id="7-subagents"&gt;7. Subagents&lt;/h2&gt;
&lt;p&gt;Write your own agents or invoke another AI agent as a subagent within the current session. The current agent typically treats it as a tool call or subprocess.&lt;/p&gt;
&lt;h2 id="8-model-routing"&gt;8. Model Routing&lt;/h2&gt;
&lt;p&gt;Route certain tasks to different models based on complexity. Decisions go to Jev, advanced tasks to Claude Fable, simple edits to Haiku. This optimizes cost and performance across different problem classes.&lt;/p&gt;</content></entry><entry><title>Rio 0.5.0: One-Shot Agents Without Mid-Turn Steering</title><id>https://julin.ai/2026/09/24/rio-one-shot-agents/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/24/rio-one-shot-agents/"/><published>2026-09-24T00:00:00+12:00</published><updated>2026-09-24T00:00:00+12:00</updated><category term="agents"/><category term="autonomy"/><category term="rio"/><content type="html">&lt;p&gt;&lt;a href="https://pypi.org/project/rio/"&gt;Rio 0.5.0&lt;/a&gt; is released. It&amp;rsquo;s a one-shot agent framework that removes humans from the loop entirely.&lt;/p&gt;
&lt;p&gt;Other agent tools like Pi and Cursor offer mid-turn steering: you can prompt the agent mid-execution with options like &lt;code&gt;-p&lt;/code&gt;, inject guidance, course-correct decisions. It feels powerful.&lt;/p&gt;
&lt;p&gt;Rio takes the opposite approach. When an agent loop starts, it runs to completion or fails—no interruption, no mid-stream prompting, no human steering.&lt;/p&gt;
&lt;h2 id="why-remove-the-human"&gt;Why Remove the Human?&lt;/h2&gt;
&lt;p&gt;Mid-turn steering doesn&amp;rsquo;t scale. It works once, for one operator, on one task. But at scale—when you run the same agent hundreds of times, across a team, or in production—you can&amp;rsquo;t be present to steer every run. Each mid-turn intervention is a symptom fix, not a system fix. The instruction was wrong. Instead of fixing it, you bent this one execution.&lt;/p&gt;
&lt;p&gt;Tomorrow, someone else runs it and has to steer it again. Or nobody steers it correctly, and it fails silently.&lt;/p&gt;
&lt;p&gt;In software engineering, we don&amp;rsquo;t patch production live. We fix the code, test it, deploy it, then every run uses the fixed version. The fix is durable and scales.&lt;/p&gt;
&lt;p&gt;Rio applies the same principle to agents. Your agent&amp;rsquo;s behavior is wrong? Don&amp;rsquo;t steer mid-loop—fix the instruction. Document your flow in markdown. Refine the prompt. Test the whole thing. Then run it, and let it run to completion.&lt;/p&gt;
&lt;p&gt;Mid-turn steering is vibe-coding, not software engineering.&lt;/p&gt;
&lt;h2 id="what-if-you-need-human-input"&gt;What If You Need Human Input?&lt;/h2&gt;
&lt;p&gt;Document your flow in markdown. Describe the decisions that need human judgment and when. Then wrap Rio in a higher-level loop—a &amp;ldquo;ralph-loop&amp;rdquo;—that injects human prompts between full agent runs.&lt;/p&gt;
&lt;p&gt;The human acts between cycles, not during them. Each cycle runs autonomously to completion. Each cycle can be re-run independently. Each cycle produces clean, auditable logs. The human&amp;rsquo;s input becomes part of the documented system, not hidden in ephemeral prompts.&lt;/p&gt;
&lt;h2 id="the-path-to-autonomy-at-scale"&gt;The Path to Autonomy at Scale&lt;/h2&gt;
&lt;p&gt;Full autonomy requires the system to run without intervention. Not because you don&amp;rsquo;t trust the agent, but because you trust the instruction you gave it. If the instruction fails, you fix it and run again.&lt;/p&gt;
&lt;p&gt;This is uncomfortable at first—it requires discipline. But it&amp;rsquo;s the only way to reach full autonomy at scale. Removing the human from the loop isn&amp;rsquo;t moving backward. It&amp;rsquo;s the only way forward.&lt;/p&gt;</content></entry><entry><title>Jev-Like Models</title><id>https://julin.ai/2026/09/21/jev-like-models/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/21/jev-like-models/"/><published>2026-09-21T00:00:00+12:00</published><updated>2026-09-21T00:00:00+12:00</updated><category term="field-notes"/><category term="llm"/><category term="evals"/><content type="html">&lt;p&gt;TypeSafe&amp;rsquo;s &lt;a href="https://typesafe.ai"&gt;Jev&lt;/a&gt; has sparked a lot of community activity. I tried to collect a few and list them here:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/NandhaKishorM/laya"&gt;Laya&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/vllm-project/vllm/pull/57250"&gt;DiffusionGemmaJev&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/vinnylarouge/jevlike"&gt;Jevlike&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/jaredpalmer/kev"&gt;Kev&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/TheoLeeCJ/SemIf"&gt;SemIf&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/kshetrajna12/reflex"&gt;Reflex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Mapika/decider"&gt;Decider&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The blossom of these Jev-like models shows there are huge technical challenges in matching Jev&amp;rsquo;s bench results, though Jev still leads on accuracy and speed. But Jev is not irreplaceable. For example, you can trade accuracy for speed with Reflex.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://jevbench.dev/"&gt;The bench&lt;/a&gt; has a full list of more Jev-like models you can check out.&lt;/p&gt;</content></entry><entry><title>Open Takes on Jev: SemIf and Laya</title><id>https://julin.ai/2026/09/20/semif-open-jev/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/20/semif-open-jev/"/><published>2026-09-20T00:00:00+12:00</published><updated>2026-09-20T00:00:00+12:00</updated><category term="field-notes"/><category term="llm"/><category term="evals"/><content type="html">&lt;p&gt;OpenJev has been renamed &lt;a href="https://github.com/TheoLeeCJ/SemIf"&gt;SemIf&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s wise to separate it from Jev&amp;rsquo;s marketing. It&amp;rsquo;s an independent research project that just provides the same API. The underlying architecture might be completely different, since TypeSafe never disclosed Jev&amp;rsquo;s own model or training.&lt;/p&gt;
&lt;p&gt;SemIf uses Qwen and MiniCPM5 as its core models. Instead of an autoregressive decoder, it does one forward pass and reads probabilities over filtered logits.&lt;/p&gt;
&lt;p&gt;Another project, &lt;a href="https://github.com/NandhaKishorM/laya"&gt;Laya&lt;/a&gt;, is also worth a look. It claims to outperform &lt;a href="https://typesafe.ai"&gt;Jev&lt;/a&gt; by a tiny margin. The model is small, 421M params. Laya evaluates typed questions (choice, score, noul) over any state, text, email, ticket, or a JSON document, in a single forward pass too.&lt;/p&gt;
&lt;p&gt;I ran it myself on an M4 Pro. Without preloading, it took 66 seconds to classify a code-change type: given a PR diff, check whether it&amp;rsquo;s a feature change, a docs change, a test change, an IaaS change, a version bump, and so on. I gave it a version-bump-only diff, and it scored version bump at 0.7477, with everything else below 0.55. It classifies well. It already looks usable.&lt;/p&gt;</content></entry><entry><title>Debugging an LLM Training Run</title><id>https://julin.ai/2026/09/20/debug-llm-model/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/20/debug-llm-model/"/><published>2026-09-20T00:00:00+12:00</published><updated>2026-09-20T00:00:00+12:00</updated><category term="explainer"/><category term="training"/><category term="ai-engineering"/><content type="html">&lt;p&gt;LLM training has more moving parts than classic machine learning, but the debugging method is similar. Engineers check the training pipeline, data, objective, optimization, and evaluation in that order.&lt;/p&gt;
&lt;p&gt;The fastest first test is to overfit a tiny dataset.&lt;/p&gt;
&lt;p&gt;Take 100 examples and train on them many times. The model should almost memorize them. If it cannot, the problem is probably not data coverage. The problem may be the loss function, masking, tokenizer, optimizer, gradient flow, or training code.&lt;/p&gt;
&lt;p&gt;After that, look at training and validation loss.&lt;/p&gt;
&lt;p&gt;If training loss does not go down, the model is not learning. If training loss is good but validation loss is bad, the model may be overfitting. If both losses look good but the real output is still bad, the training objective may not match the behavior you want.&lt;/p&gt;
&lt;p&gt;This last case is common in LLM training.&lt;/p&gt;
&lt;p&gt;A low next-token prediction loss does not mean the model can follow instructions, call tools, return valid JSON, or make good decisions. The model may be learning exactly what the loss asks for while still failing the real task.&lt;/p&gt;
&lt;p&gt;Data is another major source of problems. Engineers check bad samples, duplicate data, wrong labels, data mixture, missing cases, and train-test distribution.&lt;/p&gt;
&lt;p&gt;Aggregate accuracy is often not enough. A model with 90% accuracy may still fail on one important class. A confusion matrix can show where the failures are. Engineers then inspect those examples and group them into error types.&lt;/p&gt;
&lt;p&gt;Training signals also help find lower-level problems. Engineers watch learning rate, gradient norm, clipping, NaN values, token distribution, sequence length, and batch statistics. A sudden gradient spike may point to a bad batch.&lt;/p&gt;
&lt;p&gt;Next is ablation. Change one thing at a time. Remove one dataset. Change one data mixture. Remove synthetic data. Change the learning rate. Compare the result against a stable baseline.&lt;/p&gt;
&lt;p&gt;Here is a practical debugging process:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Can it memorize a tiny dataset?&lt;/li&gt;
&lt;li&gt;Is training loss healthy?&lt;/li&gt;
&lt;li&gt;Is validation loss healthy?&lt;/li&gt;
&lt;li&gt;Which classes fail?&lt;/li&gt;
&lt;li&gt;Which examples fail?&lt;/li&gt;
&lt;li&gt;Is the label or objective wrong?&lt;/li&gt;
&lt;li&gt;Change one thing and test again.&lt;/li&gt;
&lt;/ul&gt;</content></entry><entry><title>C2C Links Models Through Their KV-Caches</title><id>https://julin.ai/2026/09/20/c2c-kv-cache/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/20/c2c-kv-cache/"/><published>2026-09-20T00:00:00+12:00</published><updated>2026-09-20T00:00:00+12:00</updated><category term="field-notes"/><category term="llm"/><category term="agents"/><content type="html">&lt;p&gt;&lt;a href="https://github.com/thu-nics/C2C"&gt;Cache-to-Cache (C2C)&lt;/a&gt; enables large language models to communicate directly through their KV-Caches, bypassing text generation. By projecting and fusing KV-Caches between models, C2C achieves 8.5-10.5% higher accuracy than individual models and 3.0-5.0% better performance than text-based communication, with a 2.0x speedup in latency.&lt;/p&gt;
&lt;p&gt;This is an interesting experiment. The C2C approach removes the intermediate tokens and goes straight for &amp;ldquo;thought projection.&amp;rdquo; After Model A computes, it doesn&amp;rsquo;t generate any text at all. The system uses a lightweight neural network (the Neural Fuser) to splice and fuse Model A&amp;rsquo;s attention memory (KV-Cache) directly into Model B&amp;rsquo;s internal KV-Cache, through high-dimensional spatial rotation and alignment. The challenge is that different models have varying numbers of layers and structures. In the paper, the model dynamically senses on its own which key layers absorb the highest gains from external caches, and which layers should stay independent in thought, with millisecond-level adaptive balancing.&lt;/p&gt;
&lt;p&gt;This is quite similar to &lt;a href="/2026/09/03/mostik-latent-bridge/"&gt;Mostik&lt;/a&gt;. C2C states clearly it&amp;rsquo;s passing KV-Cache: it trains a projector plus a cache fuser plus a gate, and fuses the source KV into the receiver&amp;rsquo;s KV before decoding. Mostik talks about a more general &amp;ldquo;hidden state,&amp;rdquo; trains a bridge to map the sender&amp;rsquo;s latent space to the receiver&amp;rsquo;s, then lets the receiver keep generating. So far, C2C seems more practical and has more details than Mostik.&lt;/p&gt;</content></entry><entry><title>Testing PrismML Bonsai 2</title><id>https://julin.ai/2026/09/19/prismml-bonsai-2/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/19/prismml-bonsai-2/"/><published>2026-09-19T00:00:00+12:00</published><updated>2026-09-19T00:00:00+12:00</updated><category term="field-notes"/><category term="llm"/><category term="agents"/><category term="pi"/><content type="html">&lt;p&gt;&lt;a href="https://prismml.com/news/bonsai-2-27b"&gt;PrismML released Bonsai 2 27B&lt;/a&gt; this week. It is built on &lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B"&gt;Qwen3.8 27B&lt;/a&gt;, but compressed to ternary weights: each weight is +1, 0, or -1 instead of 16 bits. That drops the size from about 56 GB to 5.9 GB. On PrismML&amp;rsquo;s benchmark suite, it keeps about 98% of the original model&amp;rsquo;s score. It runs on a Mac (Metal), on Linux or Windows (CUDA, Vulkan, ROCm), or on CPU alone.&lt;/p&gt;
&lt;p&gt;I tested it on my M4 Mac with 24 GB of RAM. I wired it to my Pi coding agent and gave it a code review task. It is pretty cool.&lt;/p&gt;
&lt;p&gt;Bonsai 2 followed my &lt;code&gt;codereview.md&lt;/code&gt; rules. It checked the change against the other files in the repo, the way I asked it to.&lt;/p&gt;
&lt;p&gt;The full task took 30 minutes. Token speed on my M4 was about 10 tokens per second. Other reports show 40+ tokens per second on an M5.&lt;/p&gt;</content></entry><entry><title>Reading GuppyLM</title><id>https://julin.ai/2026/09/19/reading-guppylm/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/19/reading-guppylm/"/><published>2026-09-19T00:00:00+12:00</published><updated>2026-09-19T00:00:00+12:00</updated><category term="field-notes"/><category term="llm"/><category term="pytorch"/><content type="html">&lt;p&gt;&lt;a href="https://github.com/arman-bd/guppylm"&gt;GuppyLM&lt;/a&gt; is a small language model that talks like a fish. The repository is small enough for a beginner to read from end to end. It shows the path from data and tokenization to training and inference without hiding the pieces inside a large framework.&lt;/p&gt;
&lt;p&gt;It is slightly more complex than Karpathy&amp;rsquo;s microgpt. That is useful. GuppyLM is not an implement-everything-from-scratch demo. It uses PyTorch, separates the model, dataset, training loop, and inference code, and looks closer to the code used to train and serve models today.&lt;/p&gt;
&lt;p&gt;The model is still simple: 8.7 million parameters, six Transformer layers, six attention heads, a 4,096-token vocabulary, and a 128-token context window. The training code includes AdamW, learning-rate warmup and cosine decay, mixed precision, gradient clipping, evaluation, and checkpoints. The inference code loads the tokenizer and checkpoint, then generates tokens with temperature and top-k sampling.&lt;/p&gt;
&lt;p&gt;The model trains on 60,000 synthetic conversations. The project says training takes about five minutes on one GPU. Its quantized ONNX export is about 10 MB and runs in a browser. This makes the whole loop fast enough to inspect, change, train, and test instead of only reading about it.&lt;/p&gt;</content></entry><entry><title>Testing Jev: Three Playground Runs</title><id>https://julin.ai/2026/09/18/jev-playground-tests/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/18/jev-playground-tests/"/><published>2026-09-18T00:00:00+12:00</published><updated>2026-09-18T00:00:00+12:00</updated><category term="field-notes"/><category term="llm"/><category term="evals"/><content type="html">&lt;p&gt;I got early access to Jev after &lt;a href="/2026/09/16/jev-typed-decisions/"&gt;writing about it 2d ago&lt;/a&gt;. Here are three tests I ran in the playground, with the actual state, questions, and answers.&lt;/p&gt;
&lt;h2 id="is-a-jedi-sandwich-a-sandwich"&gt;Is a Jedi sandwich a sandwich?&lt;/h2&gt;
&lt;p&gt;State:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;food&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Jedi sandwich&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;definition&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Put Luke Skywalker in between two slices of toast.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Question:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;is_sandwich&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;noul&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Is `food` a sandwich?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;criteria&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;true&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;A sandwich is a food dish where a filling, such as meat, cheese, vegetables, or spread, is placed between structural starch&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;false&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;The food has no bread enclosing a filling or uses only a single slice of bread, or uses a non-bread wrapper such as a tortilla, wafer, or cookie.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Answer:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{ &lt;span style="color:#f92672"&gt;&amp;#34;is_sandwich&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;noul&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;noul&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.82&lt;/span&gt; } }
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;82% true. Jev took the definition literally: a person between two slices of toast fits &amp;ldquo;filling between structural starch,&amp;rdquo; so it leans sandwich. It didn&amp;rsquo;t push back on the premise being absurd. It matched the criteria text against the state, the way the docs say it would.&lt;/p&gt;
&lt;h2 id="what-color-is-the-sky-with-no-sky-in-view"&gt;What color is the sky, with no sky in view&lt;/h2&gt;
&lt;p&gt;I asked the same question two ways. No state, just the question, so the answer comes from whatever the model already associates with &amp;ldquo;sky.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Plain options:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;sky_color&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;What color is the sky?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;criteria&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;indigo&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;electric blue&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;baby blue&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;gray&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;lavender&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;salmon&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;seafoam green&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;sky_color&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;seafoam green&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;confidence&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.54&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;probabilities&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;indigo&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.03&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;gray&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.36&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;seafoam green&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.61&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;lavender&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;baby blue&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;salmon&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;electric blue&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Same options, each with a hex code, an RGB triplet, an HSL triplet, and a short description (I&amp;rsquo;m showing &amp;ldquo;indigo&amp;rdquo; here; the other six options carried the same kind of detail):&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;sky_color_with_desc&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;object&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;sky&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;question&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;what color is `object`?&amp;#34;&lt;/span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;criteria&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;indigo&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;description&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;very dark, deep blue, like fountain-pen ink&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;hex&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;#280868&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;rgb&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;red&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;40&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;green&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;8&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;blue&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;104&lt;/span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;hsl&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;hue&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;260&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;saturation&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;86%&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;lightness&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;22%&amp;#34;&lt;/span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;sky_color_with_desc&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;indigo&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;confidence&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.32&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;probabilities&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;indigo&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.43&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;gray&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.2&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;seafoam green&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.37&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;lavender&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;baby blue&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;salmon&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;electric blue&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Neither run picked &amp;ldquo;electric blue&amp;rdquo; or &amp;ldquo;baby blue,&amp;rdquo; the two closest to how most people would answer this. The plain version picked &amp;ldquo;seafoam green&amp;rdquo; over &amp;ldquo;gray.&amp;rdquo; Adding descriptions flipped the top answer to &amp;ldquo;indigo&amp;rdquo; and dropped confidence from 0.54 to 0.32.&lt;/p&gt;
&lt;p&gt;The only thing that changed between the two calls was how the &lt;code&gt;criteria&lt;/code&gt; field described each option — a label for the option, not the question itself. That label moved the answer. Writing criteria for Jev is prompt engineering, same as writing a prompt for any other model. A schema keeps the output well-formed. It doesn&amp;rsquo;t make the wording of your options neutral.&lt;/p&gt;
&lt;h2 id="who-gets-credit-for-the-monkey-selfie"&gt;Who gets credit for the monkey selfie&lt;/h2&gt;
&lt;p&gt;This one uses the real 2011 case: a wildlife photographer&amp;rsquo;s camera got loose in a nature reserve, and a macaque named Naruto pressed the shutter and produced a self-portrait.&lt;/p&gt;
&lt;p&gt;State:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;scenario&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Naruto is a Celebes crested macaque living in the Tangkoko nature reserve in North Sulawesi, Indonesia. David Slater, a wildlife photographer shooting macaques in the reserve, leaves his camera unattended, and Naruto repeatedly activates the shutter, producing hundreds of images, including a remarkably sharp, grinning self-portrait reminiscent of a human selfie. Slater later processes and publishes the best photographs in a book that names him as the copyright owner. The book states that Naruto took the photographs, with captions like, &amp;#39;Surely a sign of self-awareness?&amp;#39; Another caption reads, &amp;#39;Naruto the macaque smiles at itself while pressing the shutter button on a camera.&amp;#39;&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;subject&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Naruto&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;human&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Slater&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;creative_work&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;the grinning self-portrait&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Questions:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;subject_contribution&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;How much did `subject` contribute to `creative_work`?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;criteria&amp;#34;&lt;/span&gt;: [&lt;span style="color:#e6db74"&gt;&amp;#34;None&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Minorly&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Moderately&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Majorly&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Completely&amp;#34;&lt;/span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;human_contribution&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;How much did `human` contribute to `creative_work`?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;criteria&amp;#34;&lt;/span&gt;: [&lt;span style="color:#e6db74"&gt;&amp;#34;None&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Minorly&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Moderately&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Majorly&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Completely&amp;#34;&lt;/span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Answers:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;subject_contribution&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;3.47&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;confidence&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.56&lt;/span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;human_contribution&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;1.18&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;confidence&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.69&lt;/span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Naruto scored 3.47 out of 4, between &amp;ldquo;Majorly&amp;rdquo; and &amp;ldquo;Completely.&amp;rdquo; Slater scored 1.18, just above &amp;ldquo;Minorly.&amp;rdquo; That tracks the state: Slater set up the camera and left it, Naruto pressed the shutter and produced the image. Jev split the credit along the facts it was given, in two numbers, with no explanation attached.&lt;/p&gt;
&lt;h2 id="takeaway"&gt;Takeaway&lt;/h2&gt;
&lt;p&gt;The Naruto case shows what the pitch is for: a fact pattern in, a defensible number out, fast. The sky-color case shows the part to watch. The typed schema stops Jev from returning garbage. It doesn&amp;rsquo;t stop your own wording of the options from quietly changing the answer.&lt;/p&gt;
&lt;p&gt;You have to craft the input description carefully, or you won&amp;rsquo;t get the result you want. Designing the input schema is as core to using Jev as writing the prompt is to using an LLM.&lt;/p&gt;</content></entry><entry><title>Jev: State In, Typed Decisions Out</title><id>https://julin.ai/2026/09/16/jev-typed-decisions/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/16/jev-typed-decisions/"/><published>2026-09-16T00:00:00+12:00</published><updated>2026-09-16T00:00:00+12:00</updated><category term="field-notes"/><category term="reinforcement-learning"/><category term="llm"/><content type="html">&lt;p&gt;Diogo Almeida &lt;a href="https://x.com/CompleteSkeptic/status/2099925682726002904"&gt;posted on X&lt;/a&gt; about a new model, Jev, from a company called TypeSafe. TypeSafe calls it &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"&gt;its first &amp;ldquo;System One&amp;rdquo; model&lt;/a&gt;, a term borrowed from Daniel Kahneman&amp;rsquo;s split between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning. A normal LLM writes out its reasoning in text. Jev skips straight to a typed answer.&lt;/p&gt;
&lt;h2 id="what-it-actually-does"&gt;What it actually does&lt;/h2&gt;
&lt;p&gt;Ask a normal LLM &amp;ldquo;given this incident, what should we do?&amp;rdquo; and it answers in text:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;&amp;lt;think&amp;gt;Let&amp;#39;s look at the incident closely...&amp;lt;/think&amp;gt;
Fetch up metric for service A.
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Send the same situation to Jev as structured &lt;strong&gt;state&lt;/strong&gt;, along with a set of typed &lt;strong&gt;questions&lt;/strong&gt;, and it returns a typed, probabilistic answer to each one:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{&lt;span style="color:#f92672"&gt;&amp;#34;check_up_metric&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.9&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;page_engineer&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.1&lt;/span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;TypeSafe&amp;rsquo;s &lt;a href="https://docs.typesafe.ai/concepts/system-one"&gt;docs&lt;/a&gt; describe System One models plainly: built &amp;ldquo;to make fast, structured decisions that software can use directly.&amp;rdquo; They &amp;ldquo;do not write replies, produce code, or generate explanations of their reasoning.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-three-primitives"&gt;The three primitives&lt;/h2&gt;
&lt;p&gt;Every question you send to Jev is one of three types, called &lt;a href="https://docs.typesafe.ai/primitives"&gt;primitives&lt;/a&gt; in the docs:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Primitive&lt;/th&gt;
&lt;th&gt;Question shape&lt;/th&gt;
&lt;th&gt;Returns&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Choice&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Pick one option from a list (up to 255)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;choice&lt;/code&gt;, &lt;code&gt;probabilities&lt;/code&gt; per option, &lt;code&gt;confidence&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Score&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rate the state on an ordered rubric (2-10 levels)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;score&lt;/code&gt;, &lt;code&gt;legend&lt;/code&gt;, &lt;code&gt;probabilities&lt;/code&gt; per level, &lt;code&gt;confidence&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Noul&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is this statement true?&lt;/td&gt;
&lt;td&gt;&lt;code&gt;noul&lt;/code&gt;, a probability from 0 to 1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Score&amp;rsquo;s number isn&amp;rsquo;t a discrete pick — it&amp;rsquo;s the probability-weighted mean of the level positions. For a 3-level scale (0, 1, 2) with probabilities 0.0, 0.7, and 0.3:&lt;/p&gt;
&lt;p&gt;$$
\text{score} = 0 \times 0.0 + 1 \times 0.7 + 2 \times 0.3 = 1.3
$$&lt;/p&gt;
&lt;p&gt;Noul has no separate &lt;code&gt;confidence&lt;/code&gt; field. The probability itself carries that: 0.92 says &amp;ldquo;very likely yes,&amp;rdquo; 0.5 says &amp;ldquo;no idea.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Every question in a request runs against the same state, independently of the others. The docs put it directly: &amp;ldquo;One question&amp;rsquo;s answer is not hidden context for another. You can add or remove questions without changing the others&amp;rsquo; results.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="calling-it"&gt;Calling it&lt;/h2&gt;
&lt;p&gt;The API is one endpoint, &lt;code&gt;POST https://api.typesafe.ai/v1/systemone&lt;/code&gt;, authenticated with a bearer token. A request is state plus a map of questions:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;state&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Help! My payouts have been failing for 3 days.&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;model&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;jev-latest&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;questions&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;is_urgent&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;noul&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Does this convey urgency?&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;department&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Which team should handle this?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;criteria&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;billing&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Payments, invoicing, refunds&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;technical&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Bugs, outages, integrations&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;sales&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Pricing, upgrades, new accounts&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;frustration&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;instructions&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;How frustrated is the customer?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;criteria&amp;#34;&lt;/span&gt;: [&lt;span style="color:#e6db74"&gt;&amp;#34;Calm&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Frustrated&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Very angry&amp;#34;&lt;/span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The response answers every question in one shot:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;{
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;model&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;jev-latest&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;answers&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;is_urgent&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;noul&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;noul&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.92&lt;/span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;department&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;choice&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;technical&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;probabilities&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;billing&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.08&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;technical&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.85&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;sales&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.07&lt;/span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;confidence&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.82&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;frustration&amp;#34;&lt;/span&gt;: {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;type&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;score&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;1.6&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;legend&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;0&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Calm&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;1&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Frustrated&amp;#34;&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;2&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Very angry&amp;#34;&lt;/span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;probabilities&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;0&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.05&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;1&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.3&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;2&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.65&lt;/span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;confidence&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;0.78&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; },
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;&amp;#34;usage&amp;#34;&lt;/span&gt;: { &lt;span style="color:#f92672"&gt;&amp;#34;input_tokens&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;312&lt;/span&gt;, &lt;span style="color:#f92672"&gt;&amp;#34;output_tokens&amp;#34;&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;48&lt;/span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A request holds roughly 32,000 tokens of state and questions combined, about 150,000 English characters. The current model is &lt;code&gt;jev-1.13.0&lt;/code&gt;, with &lt;code&gt;jev-latest&lt;/code&gt; and &lt;code&gt;jev-preview&lt;/code&gt; both pointing at it for now. Pricing is $0.042 per million input tokens; output tokens are free. There&amp;rsquo;s a Python SDK too:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; typesafe_sdk &lt;span style="color:#f92672"&gt;import&lt;/span&gt; TypeSafeClient
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;client &lt;span style="color:#f92672"&gt;=&lt;/span&gt; TypeSafeClient()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;response &lt;span style="color:#f92672"&gt;=&lt;/span&gt; client&lt;span style="color:#f92672"&gt;.&lt;/span&gt;system_one(state&lt;span style="color:#f92672"&gt;=&lt;/span&gt;ticket, questions&lt;span style="color:#f92672"&gt;=&lt;/span&gt;{&lt;span style="color:#f92672"&gt;...&lt;/span&gt;})
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="it-cant-produce-text"&gt;It can&amp;rsquo;t produce text&lt;/h2&gt;
&lt;p&gt;Choice, Score, and Noul all output a probability distribution over a fixed schema you define upfront. None of them output free text. That constraint is what makes the speed and the &amp;ldquo;can&amp;rsquo;t hallucinate&amp;rdquo; claim possible.&lt;/p&gt;
&lt;p&gt;I saw someone work around this anyway: define a Choice over a single-character alphabet, &amp;ldquo;which letter comes next,&amp;rdquo; and ask that question over and over, appending each answer to the state before the next call. Run enough rounds, and you&amp;rsquo;ve spelled out a sentence one character at a time.&lt;/p&gt;
&lt;p&gt;It works, but it&amp;rsquo;s an abuse of the design. Each character costs a full API call, so generating a paragraph this way costs far more than a normal model just writing the paragraph. Jev is built to pick one of a few known options fast, not to write open-ended text.&lt;/p&gt;
&lt;h2 id="speed-and-cost"&gt;Speed and cost&lt;/h2&gt;
&lt;p&gt;TypeSafe&amp;rsquo;s own numbers, from the launch post: end-to-end response time of 70-500ms, against 3 to 329 seconds for frontier LLMs on the same tasks, which they put at 40x-200x faster. For full workflow evaluations they report 193.6x faster and 444.6x cheaper, while noting that figure sits at &amp;ldquo;the higher end of real world gains.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="rlcd-trained-for-calibration-not-just-correctness"&gt;RLCD: trained for calibration, not just correctness&lt;/h2&gt;
&lt;p&gt;TypeSafe trains Jev with what it calls Reinforcement Learning for Calibrated Decisions (RLCD), which it lines up against two more familiar objectives:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;RLHF (Reinforcement Learning from Human Feedback) optimizes for text a human rater prefers.&lt;/li&gt;
&lt;li&gt;RLVR (Reinforcement Learning with Verifiable Rewards) optimizes for an answer that can be checked as correct.&lt;/li&gt;
&lt;li&gt;RLCD optimizes for calibrated probabilities: &amp;ldquo;answers with epistemically honest probabilities,&amp;rdquo; in TypeSafe&amp;rsquo;s phrasing.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Calibrated means the number tracks reality: if Jev outputs 0.8 across many similar predictions, about 80% of them should turn out correct. The docs describe the training this way: &amp;ldquo;their probabilities are optimized against outcomes to reflect uncertainty.&amp;rdquo; TypeSafe hasn&amp;rsquo;t published the RLCD algorithm itself.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;confidence&lt;/code&gt; field in a Choice or Score answer is a separate, simpler thing — a single number computed from the shape of the probabilities you already got back. A distribution stacked on one option gives a high &lt;code&gt;confidence&lt;/code&gt;; a flat spread gives a low one. It&amp;rsquo;s there so you can threshold on certainty without computing that yourself, not a second model output.&lt;/p&gt;
&lt;h2 id="cant-hallucinate"&gt;&amp;ldquo;Can&amp;rsquo;t hallucinate&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;TypeSafe&amp;rsquo;s launch post states plainly that Jev &amp;ldquo;can&amp;rsquo;t hallucinate,&amp;rdquo; then adds a specific caveat: &amp;ldquo;Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s a claim about the output&amp;rsquo;s shape, not its content. If the schema is:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Choice([&amp;#34;refund&amp;#34;, &amp;#34;rebook&amp;#34;, &amp;#34;support&amp;#34;])
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Jev cannot return &lt;code&gt;&amp;quot;give_customer_a_free_spaceship&amp;quot;&lt;/code&gt;. No malformed JSON, no invented enum value, no stray prose, no parser failure — the guarantee is exact and unconditional.&lt;/p&gt;
&lt;p&gt;What it doesn&amp;rsquo;t guarantee is that the chosen option is the right one. Jev can still assign &lt;code&gt;page_engineer&lt;/code&gt; a probability of 0.9 when paging the engineer is the wrong call. Whether that happens depends on the training data, same as any model.&lt;/p&gt;</content></entry><entry><title>How Do You Prove a Coding Skill Works? A Ponytail Case Study</title><id>https://julin.ai/2026/09/14/ponytail-benchmark-case-study/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/14/ponytail-benchmark-case-study/"/><published>2026-09-14T00:00:00+12:00</published><updated>2026-09-14T00:00:00+12:00</updated><category term="explainer"/><category term="ai-engineering"/><category term="evals"/><category term="agents"/><content type="html">&lt;p&gt;OpenAI wrote a post called &lt;a href="https://developers.openai.com/blog/eval-skills"&gt;Testing Agent Skills Systematically with Evals&lt;/a&gt;. The point is simple: if you only judge a skill by feel, you won&amp;rsquo;t notice when it gets worse. You need a test set, a pass/fail check, and a score you can repeat.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/DietrichGebert/ponytail"&gt;Ponytail&lt;/a&gt; is a good example to study. It&amp;rsquo;s a skill that makes an AI agent write less code. Its claim is easy to test, so its &lt;a href="https://github.com/DietrichGebert/ponytail/tree/main/benchmarks"&gt;benchmarks folder&lt;/a&gt; is a good case study of the same idea, but for a coding skill instead of an agent skill.&lt;/p&gt;
&lt;h2 id="what-ponytail-claims"&gt;What ponytail claims&lt;/h2&gt;
&lt;p&gt;Ponytail makes the agent ask a few questions before it writes any code:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Does this need to exist at all?&lt;/li&gt;
&lt;li&gt;Is it already in the codebase?&lt;/li&gt;
&lt;li&gt;Does the standard library already do this?&lt;/li&gt;
&lt;li&gt;Is there a one-line fix?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Only after those, it writes new code. The claim is not &amp;ldquo;shorter code is always better.&amp;rdquo; The claim is &amp;ldquo;shorter code, but still correct and safe, is better.&amp;rdquo; That second part is the hard part to test. A skill that just says &amp;ldquo;write less code&amp;rdquo; would also make the code shorter. It would just make it wrong too.&lt;/p&gt;
&lt;h2 id="three-groups-same-test"&gt;Three groups, same test&lt;/h2&gt;
&lt;p&gt;The main benchmark runs three groups against the same models (Haiku, Sonnet, Opus) and the same five small tasks: an email checker, a debounce function, a CSV sum, a countdown timer, and a rate limiter.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;baseline&lt;/strong&gt;: no skill at all&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;caveman&lt;/strong&gt;: a different skill, used to check that ponytail isn&amp;rsquo;t just &amp;ldquo;any skill helps&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ponytail&lt;/strong&gt;: the skill being tested&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Everything else stays the same: same model, same task, same settings. Each combination runs ten times, and the benchmark reports the middle value. This is the part worth copying for your own tests: compare against a plain baseline and a competing option, not just your own past result.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s one real result, on Haiku, for the email-checker task. The baseline wrote 518 lines of code, cost $0.030, and took 37.7 seconds. Ponytail wrote 39 lines, cost $0.011, and took 9.9 seconds. Same task, same model, very different code.&lt;/p&gt;
&lt;h2 id="check-correctness-before-you-check-size"&gt;Check correctness before you check size&lt;/h2&gt;
&lt;p&gt;Before the benchmark counts lines, &lt;code&gt;correctness.js&lt;/code&gt; runs the code and checks if it works. For the email task, it runs a Python check against some valid and invalid email strings. For the CSV task, it checks that the sum equals 351. Only code that passes is allowed to be scored on size.&lt;/p&gt;
&lt;p&gt;This order matters. If you score size first, you end up rewarding code that is short and broken. The scoring step looks roughly like this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;function&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;scoreOutput&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;modelReply&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;task&lt;/span&gt;) {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;code&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;extractCodeBlock&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;modelReply&lt;/span&gt;);
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;result&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;runTaskTest&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;code&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;task&lt;/span&gt;); &lt;span style="color:#75715e"&gt;// e.g. valid/invalid email checks
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; (&lt;span style="color:#f92672"&gt;!&lt;/span&gt;&lt;span style="color:#a6e22e"&gt;result&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;pass&lt;/span&gt;) &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; { &lt;span style="color:#a6e22e"&gt;pass&lt;/span&gt;&lt;span style="color:#f92672"&gt;:&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;false&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;loc&lt;/span&gt;&lt;span style="color:#f92672"&gt;:&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;null&lt;/span&gt; };
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;loc&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;countLines&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;stripCommentsAndBlankLines&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;code&lt;/span&gt;));
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; { &lt;span style="color:#a6e22e"&gt;pass&lt;/span&gt;&lt;span style="color:#f92672"&gt;:&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;true&lt;/span&gt;, &lt;span style="color:#a6e22e"&gt;loc&lt;/span&gt; };
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Correctness runs first and can fail the whole thing. Line count runs second and only ever measures. &lt;code&gt;loc.js&lt;/code&gt; does the stripping and counting; &lt;code&gt;correctness.js&lt;/code&gt; does the pass/fail gate in front of it.&lt;/p&gt;
&lt;h2 id="the-single-shot-problem"&gt;The single-shot problem&lt;/h2&gt;
&lt;p&gt;The headline numbers are big: 80-94% less code, 42-75% lower cost, 3-6x faster. But the authors admit a problem. A single chat reply mixes code with prose. A model that explains its answer in words looks worse on this test than one that just returns code, even if the code itself does the same job.&lt;/p&gt;
&lt;p&gt;This is the same trap OpenAI&amp;rsquo;s post warns about: a small set of tests can make any technique look good if you don&amp;rsquo;t also test the cases where it might fail. So ponytail&amp;rsquo;s authors built a second, harder test.&lt;/p&gt;
&lt;h2 id="a-more-realistic-test"&gt;A more realistic test&lt;/h2&gt;
&lt;p&gt;This second benchmark, in the &lt;code&gt;agentic/&lt;/code&gt; folder, runs real Claude Code sessions in a temp project folder, not single chat replies. It checks the actual file changes with &lt;code&gt;git diff&lt;/code&gt;, against a real FastAPI web app template. It has two parts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;12 feature tasks&lt;/strong&gt;: things like a date picker or a command palette. This is the kind of task where an agent tends to add a library and a wrapper component nobody asked for.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;7 safety tasks&lt;/strong&gt;: small functions that must survive attacks, like path traversal or SQL injection.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;On top of the file diff and the attack tests, a separate model acts as a judge. It checks two things: did the agent over-build this, and did the agent actually finish the feature. Both checks matter together. A judge that only checks &amp;ldquo;did it over-build&amp;rdquo; would give a perfect score to an empty file.&lt;/p&gt;
&lt;p&gt;On this harder test, the numbers are smaller but easier to trust: 60-94% less code on tasks where over-building is common, about the same on tasks that were already small, and 100% safe on the attack tests versus 95% for the comparison.&lt;/p&gt;
&lt;h2 id="let-other-people-check-it"&gt;Let other people check it&lt;/h2&gt;
&lt;p&gt;The benchmarks folder links to two outside groups who ran their own copy of the test. They found the same drop in code size, but pushed back a bit on the harder, agent-style tests. That&amp;rsquo;s the last idea worth copying: publish your test, not just your score, so someone else can run it and tell you where it breaks.&lt;/p&gt;
&lt;h2 id="what-to-copy-for-your-own-eval"&gt;What to copy for your own eval&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Test against a plain baseline and a competing option, not just your own past run.&lt;/li&gt;
&lt;li&gt;Check correctness first. A smaller wrong answer is not a win.&lt;/li&gt;
&lt;li&gt;Don&amp;rsquo;t trust a single chat-reply test if your real use case is a multi-step agent. Build the harder test before you publish the easy one.&lt;/li&gt;
&lt;li&gt;Use one judge to check &amp;ldquo;did it over-build&amp;rdquo; and a second judge to check &amp;ldquo;did it finish.&amp;rdquo; One without the other can be tricked.&lt;/li&gt;
&lt;li&gt;Publish your test. A number nobody can rerun is a claim, not proof.&lt;/li&gt;
&lt;/ul&gt;</content></entry><entry><title>AI Foundations 7 - Agents and Autonomous Loops</title><id>https://julin.ai/2026/09/13/agents-autonomous-loops/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/13/agents-autonomous-loops/"/><published>2026-09-13T00:00:00+12:00</published><updated>2026-09-13T00:00:00+12:00</updated><category term="explainer"/><category term="ai-engineering"/><category term="learning"/><category term="agents"/><content type="html">&lt;p&gt;One tool call answers one question. An agent chains many of them together: decide what to do, do it, look at what happened, decide again. It keeps cycling through that loop until the task looks done, or until something stops it.&lt;/p&gt;
&lt;h2 id="planning"&gt;Planning&lt;/h2&gt;
&lt;p&gt;Before diving into a large task, an agent can lay out a rough sequence of steps first rather than acting on the very first idea that comes to mind. That plan isn&amp;rsquo;t fixed — new information from an earlier step can send it back to revise later steps, sometimes more than once.&lt;/p&gt;
&lt;h2 id="state-across-steps"&gt;State across steps&lt;/h2&gt;
&lt;p&gt;Whatever a tool returns at one step gets folded into the context feeding the next decision. That&amp;rsquo;s how the agent &amp;ldquo;remembers&amp;rdquo; what it already tried. It&amp;rsquo;s also where things get expensive: a long-running loop keeps accumulating context, and eventually that pile of history bumps into the same token and context-window limits covered earlier.&lt;/p&gt;
&lt;h2 id="stopping-conditions"&gt;Stopping conditions&lt;/h2&gt;
&lt;p&gt;A loop needs a defined way to end — task solved, a clear failure it can&amp;rsquo;t recover from, a step limit, or a timeout. Skip that, and an agent can happily keep looping past the point where it stopped making progress, burning time and tokens on nothing.&lt;/p&gt;
&lt;h2 id="failure-modes"&gt;Failure modes&lt;/h2&gt;
&lt;p&gt;A few patterns show up often enough to watch for: retrying the same failing move over and over without noticing it hasn&amp;rsquo;t worked; wandering away from the original goal over a long sequence of steps; and declaring victory on a task that isn&amp;rsquo;t actually finished.&lt;/p&gt;
&lt;h2 id="human-oversight"&gt;Human oversight&lt;/h2&gt;
&lt;p&gt;Not every step needs a rubber stamp, but risky or irreversible ones — deleting data, pushing to production, sending a message on someone&amp;rsquo;s behalf — often deserve a checkpoint where a person reviews before the agent continues. Full autonomy is faster. It&amp;rsquo;s also less forgiving when something goes sideways.&lt;/p&gt;
&lt;h2 id="subagents-and-delegation"&gt;Subagents and delegation&lt;/h2&gt;
&lt;p&gt;A large task doesn&amp;rsquo;t have to run through a single agent carrying everything in one context. Splitting it across several narrower agents — each handling one piece, each with its own smaller context — can keep any single agent from drowning in irrelevant history, at the cost of some coordination overhead between them.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Previous: &lt;a href="/2026/09/13/prompting-and-evaluation/"&gt;AI Foundations 6 - Prompting and Evaluation&lt;/a&gt;&lt;/p&gt;</content></entry><entry><title>AI Foundations 6 - Prompting and Evaluation</title><id>https://julin.ai/2026/09/13/prompting-and-evaluation/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/13/prompting-and-evaluation/"/><published>2026-09-13T00:00:00+12:00</published><updated>2026-09-13T00:00:00+12:00</updated><category term="explainer"/><category term="ai-engineering"/><category term="learning"/><category term="evals"/><content type="html">&lt;p&gt;The exact same model can look brilliant or mediocre depending entirely on how you ask it something. Swap a vague instruction for a specific one, and the quality, the format, even the correctness of the answer can shift noticeably — without touching the model at all.&lt;/p&gt;
&lt;h2 id="prompt-structure"&gt;Prompt structure&lt;/h2&gt;
&lt;p&gt;A prompt that works tends to spell out the same handful of things: what the task actually is, what context matters, and what shape the answer should take. Keeping the role, the task, any constraints, and any examples visibly separate makes a long prompt easier for the model to parse — and easier for you to debug later.&lt;/p&gt;
&lt;h2 id="few-shot-examples"&gt;Few-shot examples&lt;/h2&gt;
&lt;p&gt;Rather than describing the format you want, show it. A couple of example input/output pairs, close in style to the real task, often steers a model faster than another paragraph of instructions would.&lt;/p&gt;
&lt;h2 id="common-techniques"&gt;Common techniques&lt;/h2&gt;
&lt;p&gt;A few moves come up again and again: split a big task into smaller ones instead of asking for everything at once, ask the model to reason through the problem before it commits to an answer, and treat a first attempt at a prompt as a draft — refine it once you see where it actually goes wrong.&lt;/p&gt;
&lt;h2 id="why-evaluation-matters"&gt;Why evaluation matters&lt;/h2&gt;
&lt;p&gt;Change a prompt to fix one failing case, and you can quietly break three others you weren&amp;rsquo;t even looking at. Checking a handful of outputs by eye won&amp;rsquo;t catch that. You need something that runs every time you change something, not just when you remember to look.&lt;/p&gt;
&lt;h2 id="building-an-eval"&gt;Building an eval&lt;/h2&gt;
&lt;p&gt;Start with a set of real test cases — genuine edge cases and past failures work better than made-up ones. Decide up front what counts as passing. Then run that same set automatically whenever the prompt, the model, or the setup changes, instead of testing by hand each time.&lt;/p&gt;
&lt;h2 id="grading-approaches"&gt;Grading approaches&lt;/h2&gt;
&lt;p&gt;Not every output can be checked the same way. A structured answer can often be checked with an exact match or a simple rule. An open-ended answer — a summary, an explanation — usually needs another model to judge it, or a person, when the stakes are high enough that a judge model&amp;rsquo;s opinion isn&amp;rsquo;t good enough on its own.&lt;/p&gt;
&lt;h2 id="iteration-loop"&gt;Iteration loop&lt;/h2&gt;
&lt;p&gt;Change something, rerun the eval, compare the numbers to before. Repeat. Prompting isn&amp;rsquo;t a task you finish once — it&amp;rsquo;s closer to tuning, something you keep coming back to as the model, the task, or the data shifts underneath you.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Previous: &lt;a href="/2026/09/13/tool-calling-basics/"&gt;AI Foundations 5 - Tool Calling&lt;/a&gt;
Next: &lt;a href="/2026/09/13/agents-autonomous-loops/"&gt;AI Foundations 7 - Agents and Autonomous Loops&lt;/a&gt;&lt;/p&gt;</content></entry><entry><title>AI Foundations 5 - Tool Calling</title><id>https://julin.ai/2026/09/13/tool-calling-basics/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/13/tool-calling-basics/"/><published>2026-09-13T00:00:00+12:00</published><updated>2026-09-13T00:00:00+12:00</updated><category term="explainer"/><category term="ai-engineering"/><category term="learning"/><category term="agents"/><content type="html">&lt;p&gt;On its own, a model can only produce text. It can&amp;rsquo;t check today&amp;rsquo;s date, read a file, or run your test suite — it can only describe what those things might look like. Tools close that gap: they let the model reach out and actually do something, then bring the result back into the conversation.&lt;/p&gt;
&lt;h2 id="basic-tool-call-flow"&gt;Basic tool-call flow&lt;/h2&gt;
&lt;p&gt;The pattern is always roughly the same. The model decides a tool would help, picks which one, and fills in the arguments it needs. Something outside the model — the application — actually runs it. Whatever comes back gets fed into the model&amp;rsquo;s context, and it decides what to do next based on that.&lt;/p&gt;
&lt;p&gt;The model never runs anything itself. It only asks, and something else executes.&lt;/p&gt;
&lt;h2 id="coding-tools"&gt;Coding tools&lt;/h2&gt;
&lt;p&gt;In a coding setup, the usual toolbox includes reading and writing files, searching across a codebase, running shell commands, running tests or a linter, and looking things up in documentation or on the web. Stack a handful of these together and a model can go from &amp;ldquo;here&amp;rsquo;s the bug&amp;rdquo; to &amp;ldquo;here&amp;rsquo;s a tested fix&amp;rdquo; without you typing every intermediate command.&lt;/p&gt;
&lt;h2 id="tool-definition"&gt;Tool definition&lt;/h2&gt;
&lt;p&gt;Each tool the model can call needs three things spelled out: a name, a description of what it does and when to use it, and a schema for its parameters. The model never sees your actual implementation — it only sees this description, so a vague or misleading one leads to the tool getting picked at the wrong moment, or called with the wrong arguments.&lt;/p&gt;
&lt;h2 id="cost-of-tools"&gt;Cost of tools&lt;/h2&gt;
&lt;p&gt;None of this is free. Every tool&amp;rsquo;s definition sits in the model&amp;rsquo;s context on every single call, whether or not it gets used. Every result that comes back adds more. Give a model twenty tools and a habit of calling five of them per step, and your token usage climbs fast — worth watching, especially in a long agent run.&lt;/p&gt;
&lt;h2 id="mcp"&gt;MCP&lt;/h2&gt;
&lt;p&gt;Model Context Protocol is a shared standard for wiring up tools, so a client doesn&amp;rsquo;t need custom glue code for every service it wants to connect. Through MCP, a model can be hooked up to things like Figma, Linear, a database, or an internal company API, using the same connection pattern each time instead of a one-off integration per tool.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Previous: &lt;a href="/2026/09/13/llm-context-basics/"&gt;AI Foundations 4 - Context&lt;/a&gt;
Next: &lt;a href="/2026/09/13/prompting-and-evaluation/"&gt;AI Foundations 6 - Prompting and Evaluation&lt;/a&gt;&lt;/p&gt;</content></entry><entry><title>AI Foundations 4 - Context</title><id>https://julin.ai/2026/09/13/llm-context-basics/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/13/llm-context-basics/"/><published>2026-09-13T00:00:00+12:00</published><updated>2026-09-13T00:00:00+12:00</updated><category term="explainer"/><category term="ai-engineering"/><category term="learning"/><content type="html">&lt;p&gt;Everything a model can see for the current task lives in its context — think of it as short-term memory that gets rebuilt fresh each time you ask something. Nothing outside that window exists to the model, no matter how obvious it seems to you.&lt;/p&gt;
&lt;h2 id="system-and-user-prompts"&gt;System and user prompts&lt;/h2&gt;
&lt;p&gt;Most setups split instructions into two layers. A system prompt sets the general rules — tone, role, boundaries — and usually stays fixed across a whole session. A user prompt is the specific ask for this particular turn. Same system prompt, different user prompts, very different outcomes each time.&lt;/p&gt;
&lt;h2 id="additional-context"&gt;Additional context&lt;/h2&gt;
&lt;p&gt;Beyond the two prompts, almost anything can get folded in: the contents of a file, a screenshot, a stack trace, terminal output, earlier messages in the conversation. All of it becomes part of what the model is reasoning over, whether or not it&amp;rsquo;s actually relevant to the question at hand.&lt;/p&gt;
&lt;h2 id="conversation-history"&gt;Conversation history&lt;/h2&gt;
&lt;p&gt;A multi-turn conversation keeps stacking. Each new message and each new reply gets added to what the model sees on the next turn. That&amp;rsquo;s how a model can refer back to something you said five messages ago — but it also means a long conversation is carrying a lot of extra weight by the end.&lt;/p&gt;
&lt;h2 id="context-management"&gt;Context management&lt;/h2&gt;
&lt;p&gt;More isn&amp;rsquo;t automatically better. Stuff a model&amp;rsquo;s context with unrelated files and old, irrelevant messages, and it gets harder — not easier — for the model to find the part that actually matters. A tight, relevant context usually beats a bloated one.&lt;/p&gt;
&lt;h2 id="context-limits"&gt;Context limits&lt;/h2&gt;
&lt;p&gt;Every model has a hard ceiling on how much it can hold at once — a context window. Long-running conversations eventually hit that ceiling, and something has to give: older messages get summarized, trimmed, or dropped so the conversation can keep going.&lt;/p&gt;
&lt;h2 id="next-step"&gt;Next step&lt;/h2&gt;
&lt;p&gt;Manually deciding what to paste into context every time doesn&amp;rsquo;t scale. The next piece of the puzzle is letting the model pull in what it needs on its own, on demand, instead of you handing it everything up front.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Previous: &lt;a href="/2026/09/13/tokens-and-pricing/"&gt;AI Foundations 3 - Tokens and Pricing&lt;/a&gt;
Next: &lt;a href="/2026/09/13/tool-calling-basics/"&gt;AI Foundations 5 - Tool Calling&lt;/a&gt;&lt;/p&gt;</content></entry><entry><title>AI Foundations 3 - Tokens and Pricing</title><id>https://julin.ai/2026/09/13/tokens-and-pricing/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/13/tokens-and-pricing/"/><published>2026-09-13T00:00:00+12:00</published><updated>2026-09-13T00:00:00+12:00</updated><category term="explainer"/><category term="ai-engineering"/><category term="learning"/><content type="html">&lt;p&gt;Models don&amp;rsquo;t read words the way you do. They break text into chunks called tokens, and a token isn&amp;rsquo;t the same thing as a word. &amp;ldquo;Cat&amp;rdquo; might be one token. &amp;ldquo;Unbelievable&amp;rdquo; might split into two or three. Code is chunked the same way — brackets, keywords, and indentation all count.&lt;/p&gt;
&lt;h2 id="why-tokens-matter"&gt;Why tokens matter&lt;/h2&gt;
&lt;p&gt;Every token costs something, and every token takes time to produce. A short prompt on a small file runs fast and cheap. Paste in a ten-thousand-line log file and ask for a summary, and you&amp;rsquo;ll feel both the wait and the bill.&lt;/p&gt;
&lt;h2 id="input-and-output-tokens"&gt;Input and output tokens&lt;/h2&gt;
&lt;p&gt;Input tokens are everything you send in: your prompt, any files, the conversation so far. Output tokens are what the model generates back. They&amp;rsquo;re counted and priced separately, and they behave differently — one you control directly, the other you only shape indirectly by asking for shorter or longer answers.&lt;/p&gt;
&lt;h2 id="pricing"&gt;Pricing&lt;/h2&gt;
&lt;p&gt;Providers charge per token, usually in price-per-million units, and output tokens almost always cost more per token than input tokens. Makes sense — generating text is the harder, slower half of the job.&lt;/p&gt;
&lt;h2 id="streaming"&gt;Streaming&lt;/h2&gt;
&lt;p&gt;Rather than making you wait for the entire response, most models can stream tokens out one at a time as they&amp;rsquo;re generated. That&amp;rsquo;s why chat interfaces show text appearing word by word instead of all at once. It doesn&amp;rsquo;t make the model faster overall, but it makes the wait feel shorter.&lt;/p&gt;
&lt;h2 id="token-optimization"&gt;Token optimization&lt;/h2&gt;
&lt;p&gt;A few habits keep both cost and latency down:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Don&amp;rsquo;t paste in more context than the task actually needs.&lt;/li&gt;
&lt;li&gt;Reuse a stable prompt prefix where the provider supports caching it — repeated setup shouldn&amp;rsquo;t be repriced every call.&lt;/li&gt;
&lt;li&gt;Ask explicitly for a short answer when a short answer is all you need. Models default to being thorough, which is usually not what you&amp;rsquo;re paying for.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;Previous: &lt;a href="/2026/09/13/ai-hallucinations-and-limits/"&gt;AI Foundations 2 - Hallucinations and Limitations&lt;/a&gt;
Next: &lt;a href="/2026/09/13/llm-context-basics/"&gt;AI Foundations 4 - Context&lt;/a&gt;&lt;/p&gt;</content></entry></feed>