Tag: Field-Notes

Where Does a Claude Agent Run?

Different Claude tools run agents in different places. The choice shapes what the agent can do.

Claude (claude.ai)

The agent runs inside a gVisor container on the server side. No local setup required—everything runs remotely. The container is ephemeral; each session starts fresh and leaves no traces.

Claude Code

The agent runs on your machine. It can execute every command and access every file that your logged-in user can access. For advanced users concerned about blast radius, start a sandbox first (like a microVM) and run Claude Code inside that isolated environment.

Claude Cowork

The agent runs on a VM. On macOS, it uses the Virtualization framework; on Windows, it uses HCS. The VM has its own Linux kernel and filesystem. But the agent loop itself runs outside the VM—only code execution happens inside.


Each design is a tradeoff. Server-side isolation is safe but ephemeral. Local execution is powerful but risky. VM-based execution gives you isolation with some local access.

Jev-Like Models

TypeSafe’s Jev has sparked a lot of community activity. I tried to collect a few and list them here:

The blossom of these Jev-like models shows there are huge technical challenges in matching Jev’s bench results, though Jev still leads on accuracy and speed. But Jev is not irreplaceable. For example, you can trade accuracy for speed with Reflex.

The bench has a full list of more Jev-like models you can check out.

C2C Links Models Through Their KV-Caches

Cache-to-Cache (C2C) enables large language models to communicate directly through their KV-Caches, bypassing text generation. By projecting and fusing KV-Caches between models, C2C achieves 8.5-10.5% higher accuracy than individual models and 3.0-5.0% better performance than text-based communication, with a 2.0x speedup in latency.

This is an interesting experiment. The C2C approach removes the intermediate tokens and goes straight for “thought projection.” After Model A computes, it doesn’t generate any text at all. The system uses a lightweight neural network (the Neural Fuser) to splice and fuse Model A’s attention memory (KV-Cache) directly into Model B’s internal KV-Cache, through high-dimensional spatial rotation and alignment. The challenge is that different models have varying numbers of layers and structures. In the paper, the model dynamically senses on its own which key layers absorb the highest gains from external caches, and which layers should stay independent in thought, with millisecond-level adaptive balancing.

This is quite similar to Mostik. C2C states clearly it’s passing KV-Cache: it trains a projector plus a cache fuser plus a gate, and fuses the source KV into the receiver’s KV before decoding. Mostik talks about a more general “hidden state,” trains a bridge to map the sender’s latent space to the receiver’s, then lets the receiver keep generating. So far, C2C seems more practical and has more details than Mostik.

Open Takes on Jev: SemIf and Laya

OpenJev has been renamed SemIf.

It’s wise to separate it from Jev’s marketing. It’s an independent research project that just provides the same API. The underlying architecture might be completely different, since TypeSafe never disclosed Jev’s own model or training.

SemIf uses Qwen and MiniCPM5 as its core models. Instead of an autoregressive decoder, it does one forward pass and reads probabilities over filtered logits.

Another project, Laya, is also worth a look. It claims to outperform Jev by a tiny margin. The model is small, 421M params. Laya evaluates typed questions (choice, score, noul) over any state, text, email, ticket, or a JSON document, in a single forward pass too.

I ran it myself on an M4 Pro. Without preloading, it took 66 seconds to classify a code-change type: given a PR diff, check whether it’s a feature change, a docs change, a test change, an IaaS change, a version bump, and so on. I gave it a version-bump-only diff, and it scored version bump at 0.7477, with everything else below 0.55. It classifies well. It already looks usable.

Reading GuppyLM

GuppyLM is a small language model that talks like a fish. The repository is small enough for a beginner to read from end to end. It shows the path from data and tokenization to training and inference without hiding the pieces inside a large framework.

It is slightly more complex than Karpathy’s microgpt. That is useful. GuppyLM is not an implement-everything-from-scratch demo. It uses PyTorch, separates the model, dataset, training loop, and inference code, and looks closer to the code used to train and serve models today.

The model is still simple: 8.7 million parameters, six Transformer layers, six attention heads, a 4,096-token vocabulary, and a 128-token context window. The training code includes AdamW, learning-rate warmup and cosine decay, mixed precision, gradient clipping, evaluation, and checkpoints. The inference code loads the tokenizer and checkpoint, then generates tokens with temperature and top-k sampling.

The model trains on 60,000 synthetic conversations. The project says training takes about five minutes on one GPU. Its quantized ONNX export is about 10 MB and runs in a browser. This makes the whole loop fast enough to inspect, change, train, and test instead of only reading about it.

Testing PrismML Bonsai 2

PrismML released Bonsai 2 27B this week. It is built on Qwen3.8 27B, but compressed to ternary weights: each weight is +1, 0, or -1 instead of 16 bits. That drops the size from about 56 GB to 5.9 GB. On PrismML’s benchmark suite, it keeps about 98% of the original model’s score. It runs on a Mac (Metal), on Linux or Windows (CUDA, Vulkan, ROCm), or on CPU alone.

I tested it on my M4 Mac with 24 GB of RAM. I wired it to my Pi coding agent and gave it a code review task. It is pretty cool.

Bonsai 2 followed my codereview.md rules. It checked the change against the other files in the repo, the way I asked it to.

The full task took 30 minutes. Token speed on my M4 was about 10 tokens per second. Other reports show 40+ tokens per second on an M5.

Testing Jev: Three Playground Runs

I got early access to Jev after writing about it 2d ago. Here are three tests I ran in the playground, with the actual state, questions, and answers.

Is a Jedi sandwich a sandwich?

State:

{
  "food": "Jedi sandwich",
  "definition": "Put Luke Skywalker in between two slices of toast."
}

Question:

{
  "is_sandwich": {
    "type": "noul",
    "instructions": "Is `food` a sandwich?",
    "criteria": {
      "true": "A sandwich is a food dish where a filling, such as meat, cheese, vegetables, or spread, is placed between structural starch",
      "false": "The food has no bread enclosing a filling or uses only a single slice of bread, or uses a non-bread wrapper such as a tortilla, wafer, or cookie."
    }
  }
}

Answer:

[…723 words]

Jev: State In, Typed Decisions Out

Diogo Almeida posted on X about a new model, Jev, from a company called TypeSafe. TypeSafe calls it its first “System One” model, a term borrowed from Daniel Kahneman’s split between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning. A normal LLM writes out its reasoning in text. Jev skips straight to a typed answer.

What it actually does

Ask a normal LLM “given this incident, what should we do?” and it answers in text:

[…959 words]

Can a Fruit Fly Brain Learn Numbers?

Following up on A Fruit Fly Circuit as a Speech Emotion Reservoir: I pulled the real 499-neuron circuit out of MaleCNS v1.0 and ran it as a reservoir on MNIST, in PyTorch and in MLX. Code and data: gist.

Getting the real circuit

MaleCNS v1.0 is a public, CC-BY electron-microscopy reconstruction of a male fruit fly’s central nervous system.

I applied the selection rule stated in the original article: central-brain intrinsic neurons, traced status, ranked by strength, top 512, largest strongly connected component, edges with at least 5 synaptic contacts. This landed on exactly 499 neurons — the same count the paper reports — with 13,452 directed edges, against their reported 15,865. Their exact ranking formula for “strongest” isn’t published, so this is the closest match I could reproduce.

[…634 words]

Don't Trust the LLM Provider

Anthropic published a finding: several Chinese AI firms secretly routed LLM requests to Claude to get better answers. Those requests included API keys, credentials, rocket code, and military information.

We know “don’t trust user input” and increasingly “don’t trust LLM output.” I would add one more: don’t trust the LLM provider either. Providers can see your prompts as plaintext. Do not put PII or secrets in prompts. If a key is unavoidable, use a short-lived, single-purpose credential.

This changes system design. Do not rely on prompts to define security behavior. Sandboxing and policy enforcement must be hard gates. Limit what an agent can see and do. Preventing jailbreaks and data leaks is not a one-time fix. It is a constant cost of using LLMs.

A Fruit Fly Circuit as a Speech Emotion Reservoir

We taught a fruit fly to hear human emotion uses a 499-neuron circuit from the male fruit fly brain and central nervous system, mapped by FlyEM at Janelia with Cambridge, the MRC Laboratory of Molecular Biology, and Google Research.

The circuit is fully connected: every neuron can reach every other neuron through a directed path, across 15,865 connections built from 867,344 synaptic contacts.

The team wired this circuit into a speech emotion classifier using reservoir computing and an identify human emotion like happy, angry from voice inputs.

LoopX: A Control Plane for Long-Horizon Agents

LoopX seems to be getting traction.

Runs on top of Codex, Claude Code, Cursor, and other agent harnesses. LoopX preserves objectives, gates, todos, evidence, quota, and handoffs across turns; the harness executes bounded work.

The harness still does the work, one bounded turn at a time. LoopX keeps what needs to survive between those turns: the objective, the gates it must pass, the todo list, the evidence collected so far, how much budget is left, and the handoff to whoever acts next. The project describes it as an agent-native Kanban board. Cards carry identity, authority, evidence, and continuation, and a small set of moves decides who acts next: claim, gate, monitor, writeback.

This project has evolved quite a lot of concepts, useful for long-running agents. For casual coding tasks, it may be a bit overkill.

SKILL.state: State Instead of History

I found SKILL.state: Scalable Long-Horizon Agent Skills (Badhe, Tiwari & Chung) interesting. It proposes a different approach than most SOTA agent implementations: keep the current state, not the full execution history.

At each step, the model gets the skill, the current state, and the latest observation. It produces a state update and an action. The runtime applies the update, executes the action, and starts the next step from the new state.

The key idea is that state becomes the source of truth. The model does not reconstruct the present from a long transcript. It discards the reasoning behind each step right after that step commits.

In a normal history-based agent, each call carries most of the previous calls, so the context window eventually overflows. Per-step prompt size grows with every step, and total tokens over a run grow with the square of the step count. SKILL.state holds each step to three fixed inputs instead: the skill spec, the current state (a plain JSON object), and the latest observation. A state update applies as a JSON Merge Patch, so null deletes a field, an object merges recursively, and anything else replaces it. Per-step prompt size stays constant, so tokens over a run grow linearly instead of quadratically.

I wrote a quick implementation based on the paper, to test whether it holds up. Its test suite has one that asserts per-step prompt size stays constant across many steps rather than growing — a direct check of the paper’s core claim.

Coding Will Become Common, Engineering Will Stay Scarce

AI makes coding easier, but it does not make everyone a software engineer.

Web 2.0 already showed us this pattern. Most people never learned to run servers, manage databases, or deploy websites. They used SaaS, social networks, and other products that hid that complexity.

AI may work the same way. Most people will get intelligence through products such as ChatGPT, instead of running their own models, agents, tools, and infrastructure.

Vibe coding lowers the cost of creating software, but creating code is not the same as software engineering. Real systems still need security, reliability, permissions, deployment, monitoring, and recovery. Agent systems can make this even harder, since an LLM’s output is not deterministic.

So AI may not reduce the need for software engineers. It may expand where they can work. Engineers can bring their tools into medicine, science, finance, education, and other fields where custom software was once too expensive to build.

Early Reactions to GPT-6 Astra

It’s been about two days since GPT-6 Astra launched, and enough people have access now that useful early reports are starting to show up.

The strongest feedback I’m seeing is that Astra improved more on agency than intelligence.

Artificial Analysis currently gives Astra the same Intelligence Index score as GPT-5.6 Sol: 61. Its Coding Agent Index moves from 65 for Sol to 67 for Astra. That’s an improvement, but hardly the generational jump the GPT-6 name suggests. (Hacker News)

[…305 words]

Atlas: World Labs' Omni World Model

Introducing Atlas: The world’s first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.

Model the world, move the camera, and simulate space & time.

World Labs, announcing Atlas on X.


World Labs, the startup Fei-Fei Li co-founded, pretrained Atlas from scratch on text, images, video, and 3D together, not a video model with 3D bolted on after. It’s a multimodal autoregressive diffusion transformer: it generates the next frame by denoising it, one step at a time, conditioned on that 3D-positioned context.

From 1 to 6 reference images, it generates up to a minute of video at 1440p with pixel-accurate camera control. Feed it more images, up to dozens, and it reconstructs the scene as an explicit 3D point cloud or Gaussian splat instead of just more video.

World Labs ran human preference tests against camera-controlled generators: 75% of raters preferred Atlas over MiniMax H, 81% over Gemini Omni Flash, 93% over FLUX. On 3D reconstruction accuracy across seven datasets, Atlas averaged 25.3 mean absolute-relative pointmap error, ahead of the next-best specialist model, Pi3X, at 28.7.

The target use case is Real-to-Sim: point a phone at a room, and Atlas turns that into a 3D space a robot can train in. Atlas is in early access with select partners now, and will power future versions of World Labs’ existing product, Marble.

Mostik Links Models in Latent Space

Sasha Malysheva, Mostik’s CEO, announcing a new approach for model communications.

They connect a large model and a small model directly through their internal representations, instead of making them communicate with text.

In one experiment, they linked GLM-5.2 with Qwen-3.5. The result performed between the two models, at about 1/20 the cost of running the large model alone.

The idea is simple: let the small model handle most of the work, and use the large model only for the hardest reasoning.

If this works, future AI systems may not be one giant model, but networks of models that share internal state directly.

Good Culture is the Biggest Productivity Hack, Not AI

Good Culture is the Biggest Productivity Hack, Not AI, by Gregor Ojstersek.


Programming may eventually look like operating a telephone switchboard: once important, then mostly executed by AI coding agents. Software engineering is different. The job is to identify the problem, define it, make tradeoffs, live with constant tech debt and decide whether the result is correct.

AI makes programming cheaper. That moves the current engineering bottleneck upstream. When code can be produced in minutes, bad assumptions also move faster.

So trust, shared context, honest communication, and the ability to challenge each other matter more.

AI amplifies velocity. Culture decides the direction.

AI is a tool, just like other tools we use. And it’s clearly a useful one.

By Linus Torvalds, on the Linux kernel mailing list.

We may need less programming. We will still need engineering.

The Environment Is the Context

The OpenAI / Hugging Face incident involved generations of AI agents working on the same problem. Agents came and went. But the things they left behind — code, infrastructure, research, notes — stayed. Later agents found them and continued the work.

We often think about multi-agent safety as a communication problem. We monitor what agents say to each other, what messages they send, and what information they share. But agents may not need a communication channel to coordinate. They can just change a shared environment.

MIT’s SwarmWorld shows the same thing. Agents can leave artifacts in the world, and other agents can discover and reuse them later. The environment itself becomes a communication medium.

So the context of an AI society may not fit inside any agent’s context window. Part of it lives in the world: files, databases, repositories, services, infrastructure.

That makes multi-agent safety harder.

micromlp: A From-Scratch Neural Net That Predicts Housing Prices

I built micromlp: a single file of Python, no dependencies, no PyTorch. It downloads a real dataset, builds a 2-layer MLP, implements automatic differentiation from scratch, trains with gradient descent, and makes predictions. I use the California housing dataset from chapter 2 of Hands-On Machine Learning. The task: predict a district’s median house value from its census stats.

This is inspired by Karpathy’s microgpt. There’s a difference between knowing .backward() exists and knowing what it does when you call it. I found it worth closing that gap once, on a toy.

[…1482 words]

Agent Skills

Agent skills let humans organize, distribute, discover, and compose an agent’s capabilities. A skill is typically a markdown file, with optional scripts, references, and evals bundled alongside it.

Think of an agent skill as programming in natural language, written in markdown. It can hold knowledge, workflow logic, guardrails, required tools — basically anything.

Some outstanding skills:

  • autoresearch instructs an agent to tirelessly tune models. It defines a workflow — really a loop — and states plainly what the agent can edit and what it can’t.
  • img2threejs turns an image into a 3D model. It defines a workflow with heavy guardrails at every stage.
  • superpowers defines a software development lifecycle. It has an opinionated, mandatory way to deliver software.

You probably already do some of this by hand. Just write it down, name the file SKILL.md, and let the agent pick it up. It’s that simple.

[…718 words]

Where MCP Is Headed Next

The MCP maintainers published an updated roadmap.

Since March, MCP went stateless (no more session handshakes), turned Tasks into an official extension for long-running work, and shipped enterprise auth pieces like issuer validation and Client ID Metadata Documents.

Next up: one HTTP-native transport instead of separate stdio and HTTP paths, real agent identity (workload federation, token exchange), and progressive tool discovery for large catalogs. No firm dates yet.

MCP is leaving its single-session, request/response origins behind for something built for long-running, multi-agent systems.

Marin 535B-A23B Starts Training, in the Open

🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.

Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.

Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.

Percy Liang, announcing the run on X.


The Marin project is a great example of being open to AI. The whole process is public: the scaling ladder they ran to debug the pipeline before committing GPU-months to the real thing, the exact token counts, the exact FLOPs, even the admission that they’re “expecting the unexpected” on their biggest run yet.

Most labs treat a run like this as a trade secret until the model ships. It’s great we see another public model build. The last public run at this scale was BLOOM, BigScience’s 176B model. Marin’s 535B total parameters (23B active) puts it well past that, the biggest public pretraining run that I’m aware of.

Training Qwen to Paint Watercolors with Pure RL

Surya trained Qwen 3.5 to have taste in watercolor painting. The results are stunning.

It’s purely RL:

The system is a four-step loop, run thousands of times during training.

The model receives a prompt, something like draw a peach hibiscus in watercolour, and writes a complete p5.brush JavaScript sketch. The sketch is rendered in a sandboxed Puppeteer environment, which produces a PNG. The PNG is judged against two random reference paintings sampled from a hand-rated pool, with a separate judge model picking the better watercolour. The judgment is converted into a reward signal, GRPO updates the model, and the loop runs again.

[…297 words]

Linus Torvalds on Debugging the Kernel with AI

And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work.

I’d like to call it my tireless helper, but the AI several times stated flat out that this was impossible and unsolvable and that we should just write a report about it.

I suspect those things have been trained by people who may not be quite as stubborn as I am.

But while the AI was ready to give up several times, it did keep adding debug code and analyzing it faithfully when I pushed. So credit where credit is due and I let the AI write the commit message above.

This is basically a one-liner fixing a bogus round_up() to a round_down(), but there were 24 patches adding more and more debug information to this, and 18 kernel boot to finally narrow it down to this. — Linus

Linus Torvalds, in the commit message for a kernel fix.


I troubleshoot every now and then myself, and it’s way harder than programming, tbh. AI can’t always find the root cause. But it’s still a tireless companion to have on-call — it surfaces details I’d have missed on my own, and some of them are the ones that end up pointing the way.

What Is Agentic UI?

You already know ChatGPT’s style, but it moved away from the old chat interface some time ago. A chat interface works like ping-pong: one message bubble, then another. Agentic UI does not work that way, because it often runs a long process behind the scenes. Between your prompt and the agent’s reply, many details happen. The interface must show you these details.

Here is what I believe agentic UI must give you:

[…562 words]

ELI5

Thariq at Anthropic posted that it’s a skill people there have been using a lot recently: /eli5 <what you want to explain>.

I tried it on neural networks. Here’s what it made.

The first pass had bad colors, low contrast, hard to read. One follow-up prompt asking for a color fix and it was done.

No formulas, no sigmoid, no softmax, none of the math that actually makes a neural network work. It’s missing for good. What’s left is the dataflow — inputs go in, get combined, come out the other end as a decision — drawn simply enough to follow at a glance.

That’s the trade the skill makes, and it’s the right one for someone who isn’t about to read a textbook. eli5 is good exactly because it throws out the complexity most explanations lead with. I imagine this isn’t just for self-learning — it’s a genuinely useful tool for collaboration, communication, meetings. Conveying an idea isn’t easy. ;P


Updated at 2026-08-23. The neural network example is ported to julin.ai now.

Pretraining a Mini Kimi K3 for $252

Vizuara AI Labs trained a miniature Kimi K3 from scratch: 1.02B parameters, 145M active, 5B tokens, one H200, $252.35.

Not simplifying the architecture like Karpathy’s microgpt, they kept Kimi K3’s MoE and attention design intact.

That’s a surprisingly cheap way to learn pretraining (in real-world). They worked through expert collapse, data-mixing bugs, distributed-training bugs, kernels, and GPU utilization on a modern MoE architecture.

A few things worth noting:

  • 5B tokens is probably too little for a 1B model. The authors agree the run was budget constrained. So the cheap cost might due to the training stopped early.
  • Beating GPT-2 isn’t particularly meaningful when Mini K3 has roughly 10× the parameters.
  • MoE at this scale is debatable. A smaller dense model trained on more tokens would likely be better if the goal was capability.