Rio Durable
I’ve been working on Rio, a coding agent designed to run without a human continuously steering it. A task may take minutes, hours, or days. The agent needs to keep making progress even when nobody is watching.
The first versions of Rio used Jupyter notebooks as both context and execution environment. The model could write and edit cells, execute them, and use their results to decide what to do next.
This worked surprisingly well. A notebook is already a structured, executable document. It has cells, outputs, and a format that can be saved to disk. There is no need to invent a new programming language or an elaborate workflow graph.
But a notebook is not a durable computation runtime.
A running kernel holds variables, open files, subprocesses, and other state in memory. When the kernel dies, that state disappears. The notebook may contain the code and its previous outputs, but it cannot tell us whether an interrupted operation completed.
Rio Durable is an attempt to solve this problem with fewer abstractions.
The design document is currently a draft. The behavior described here is the proposed design, not an implemented and tested durability guarantee.
What should survive?
Consider an agent fixing a Flask application. It reads the source, edits a route, and starts the tests. The machine restarts before the tests finish.
What should happen next?
Restarting the entire task wastes completed work. Replaying the interrupted script may repeat changes that already happened. Restoring the Python kernel is not necessarily possible, and restoring its memory would not restore the filesystem or external services.
Rio makes a distinction between computation state and the working environment.
Computation state describes what the agent knows, what it intended to run, what results it received, and what remains uncertain. The working environment consists of files, processes, installed packages, and external services.
Rio persists the former. It does not promise to restore the latter.
After a restart, the agent receives its previous computation state. It inspects the current environment and decides how to continue.
This is the central durability contract.
State is context
Rio uses one SKILL.state object as the model’s context. It contains the goal, an ordered collection of cells, and runtime information.
A simplified state looks like this:
{
"schema": "rio.state/1",
"id": "s1",
"revision": 12,
"goal": "Fix the Flask route and pass tests",
"cells": [
{
"id": 1,
"previous_id": null,
"kind": "note",
"text": "Inspect app/routes.py"
},
{
"id": 2,
"previous_id": null,
"kind": "cmd",
"argv": ["pytest", "-q"]
}
],
"runtime": {
"next_cell_id": 3,
"cells": {},
"inbox": [],
"usage": {}
}
}
A cell is either a note, a Python script, or a command. Notes hold context. Scripts and commands request execution.
The model manages this context itself. It can add cells, replace them, summarize observations into notes, and remove cells it no longer needs.
There is no separate conversation transcript that must be compacted into another representation. The selected cells are the working context. The complete history stays in persistent storage.
The distinction matters for long-running agents. An agent working for a month should not need to carry a month’s worth of tool calls into every model request. It needs the current state and enough evidence to make its next decision.
One tool: step
Most coding agents expose tools such as read, write, edit, and bash. Rio Durable exposes a single model tool called step.
The model submits a JSON Patch to update State.
For example, it can replace a note:
{
"patch": [
{
"op": "test",
"path": "/revision",
"value": 12
},
{
"op": "test",
"path": "/cells/0/id",
"value": 1
},
{
"op": "replace",
"path": "/cells/0",
"value": {
"id": 3,
"previous_id": 1,
"kind": "note",
"text": "Check the route and its tests"
}
}
]
}
The Harness validates the patch and commits it atomically. It then evaluates the new State and schedules any new executable cells.
The model does not directly mutate the persistent database. It proposes a State transition. The Harness decides whether that transition is valid.
Each model round has a durable identity. A repeated submission cannot apply the same patch twice.
This gives Rio a simple boundary between model decisions and computation.
Immutable cells
Every cell is immutable. Replacing a cell creates a new version with a new ID.
Suppose State contains cells [1, 2]. Updating Cell 1 creates Cell 3, leaving the current array as [3, 2]. Cell 3 records Cell 1 as its predecessor.
Cell 1 remains in Session history, but it is no longer part of the current context.
For executable cells, each new version requests exactly one execution. Rebuilding State does not execute code. Repeatedly evaluating State does not execute it either.
To execute the same code again, the model creates another version.
This removes the need for an execution flag, a retry counter on each cell, or a system that compares source code to decide whether it has changed.
It also gives results a precise identity. An output belongs to the cell version that produced it, not to whatever code happens to occupy that position later.
Scripts instead of kernels
Rio’s notebook prototype used a shared Python kernel. That made it convenient to pass variables between cells, but it also made execution depend on hidden state.
Rio Durable removes the shared kernel.
Every Python cell is an independent script. It starts in a fresh process and declares its dependencies using PEP 723.
# /// script
# requires-python = ">=3.12"
# dependencies = ["httpx==0.28.1"]
# ///
import httpx
response = httpx.get("https://example.com")
print(response.status_code)
The Harness materializes the script and runs it with uv run. It records the output and exit status.
A command cell runs an argv array directly, without an implicit shell. For example, ["pytest", "-q"] runs the tests.
Cells do not share Python variables or import definitions from other cells. They communicate through ordinary files, services, and recorded observations.
This is less convenient than a notebook kernel for interactive exploration. It is easier to reason about after a restart.
Each execution has an explicit beginning, a durable identity, and a recorded outcome. Package environments and caches can be recreated when dependencies are available. There is no interpreter stack to restore.
Surviving crashes
An executable cell has three phases:
| Phase | Meaning | After restart |
|---|---|---|
queued | Execution accepted but not started | Can be dispatched |
started | Launch authorized; code may have run | Mark as uncertain |
finished | Final outcome committed | Preserve result |
Rio commits the started record before launching the process.
If Rio crashes after that commit, it cannot reliably know whether the process started, how far it ran, or whether it completed an external operation.
It therefore records UnknownExecution.
{
"type": "UnknownExecution",
"cell_id": 2,
"reason": "worker_lost_before_result_commit",
"worker_stopped": true,
"evidence_refs": ["session:event:104"]
}
Rio does not automatically run that cell again.
Instead, it gives the uncertainty to the model. The model can create new cells to inspect the filesystem, query an API, examine logs, or run a test.
If a script edited a Flask route before the crash, the agent can read the file and check the tests. It does not need to repeat the original edit blindly.
An inspection may establish that the desired final state exists. It may not prove that every operation in the interrupted script completed.
Rio preserves that distinction. A resolution note records the evidence and the model’s decision, while the original uncertainty remains in history.
This does not provide exactly-once execution of external effects. It provides durable knowledge of what was accepted, what finished, and what remains unknown.
SQLite as the source of truth
Rio uses SQLite for persistent Sessions.
The database records State changes, immutable cells, execution authorization, outputs, model rounds, observations, timers, and run endings.
The current State is reconstructed from committed records. Rebuilding it is a pure operation: it does not execute scripts or contact the model.
A single transaction accepts a valid patch, allocates new cell IDs, creates execution records, advances the State revision, and consumes observations delivered to that model round.
External work happens after the commit.
This ordering is important. A crash before a patch commits leaves no accepted work. A crash after the commit leaves a durable record of work that still needs attention.
Rio uses SQLite WAL mode with synchronous=FULL for its intended local durability guarantees. It also needs exclusive Session ownership and reliable worker cleanup before recovery can admit new execution.
SQLite protects committed computation records under the supported storage assumptions. It does not protect against lost disks, corrupted storage, or effects in an external service.
Events wake the agent
An agent does not need a running worker while it is waiting.
Rio uses a persistent Inbox for observations. A finished execution, an expired timer, or an external event can add an observation and wake the model.
The model sees a fixed snapshot of State for each round. Events arriving while it generates a response go into Inbox and become visible in a subsequent round.
This avoids rewriting the model’s context in the middle of generation.
A code cell can register a timer. The timer survives a restart, but it does not suspend and resume the original Python stack. When it expires, Rio creates an observation that can trigger another model decision.
If there is no executable work, no timer, and no external wake-up source, Rio does not keep calling the model in the hope that something changes.
For an agent intended to run for months, an idle Session should consume almost no resources.
Cancellation and completion are State changes
Rio also avoids separate model tools for cancellation and completion.
A note can request cancellation by naming a target cell. Another note can request a successful or failed conclusion.
These notes are committed through the same step interface. The Harness recognizes their roles and processes the requests.
Removing a cell from context does not cancel its execution. Context management and execution control are separate operations.
A successful conclusion requires no outstanding work, no unresolved execution uncertainty, and passing results from any required validators. The model decides whether the natural-language goal has been achieved. The Harness checks the mechanical conditions.
Rio has no human approval state in this core protocol. Permissions and execution limits must be enforced outside the model’s decisions. When the agent cannot proceed within those limits, it records failure rather than waiting indefinitely for approval.
Why not tasks?
Pi Durable takes a broader approach. It provides durable tasks, checkpoints, tool execution policies, ownership, extension hooks, and application documents. Its task model can represent work that waits for other work, resumes from checkpoints, or runs compensating actions.
Rio Durable deliberately does less.
An executable cell is the unit of work. A Session is the durable history. State is the model’s current context. Inbox carries observations.
There is no general Task class, workflow graph, dependency scheduler, persistent scratchpad, or operation-level effect ledger.
This choice has a cost. Rio cannot resume a Python function from its last internal operation. It cannot automatically replay the safe parts of a script or compensate for completed effects. An interrupted cell may require the model to investigate and repeat work.
For Rio’s intended workload, that is an acceptable tradeoff to test.
A coding agent can inspect the working environment. It can check whether a file exists, whether a test passes, or whether a commit has been created. These observations can inform recovery without requiring every filesystem or network operation to implement a separate durability protocol.
That does not make arbitrary side effects safe to repeat. Deployments, payments, and other externally consequential operations still need permissions and application-level safeguards.
The question is whether cell-level durability is sufficient for useful autonomous coding work, not whether it can reproduce every guarantee of a general-purpose durable workflow engine.
What remains to be tested
The most important part of this design is not the happy path. It is what happens at the boundaries between database commits, process execution, and external effects.
Rio needs fault-injection tests for crashes before dispatch, after launch authorization, during execution, and before result commits. Recovery must also handle surviving child processes, uncertain database commits, lost output buffers, and model requests that complete without an accepted patch.
Process termination tests are not enough to establish power-loss durability. Both need to be tested.
I also want to measure whether the simpler model improves actual coding performance: completion rate, token usage, execution overhead, recovery quality, and the number of model rounds spent resolving uncertainty.
The design assumes that an LLM can manage its context and investigate interrupted work effectively. That assumption needs evidence.
Less runtime, more computation
Rio started with the idea that executable cells could serve as an agent’s context. The notebook prototype showed that the approach is practical, but it also exposed the problems of tying computation to a shared kernel.
Rio Durable keeps the cells and removes the kernel.
The runtime persists decisions, observations, and execution boundaries. It does not try to preserve every detail of a running program. The model manages the context and decides how to continue when the environment changes.
The goal is not to build another workflow engine. It is to find the smallest runtime that allows an autonomous coding agent to make progress across process failures, restarts, and long periods of inactivity.
The design is available in Rio’s repository. The next step is implementation, fault injection, and evaluation.