Earendil, Armin Ronacher's shop, shipped Pi 1.0 with MCP support plus a long-run layer called Pi Durable.

One line sums it up. Treat agents as long-lived processes that pause, resume, and keep state, not one-shot prompt-response calls. Laptop asleep, container redeployed, memory exhausted: open the same storage and continue where it stopped.

Why durable now: the bottleneck moved from models to runs

As agent jobs got longer, failure changed shape.

Old bottleneck: is the model smart (the prompt and model-pick era)
New bottleneck: does it survive and continue (the state, recovery, and retry era)

The METR time-horizon post showed agents working solo for longer, and the Muse Spark post called long tasks a state-management problem. Pi Durable answers at the harness level.

Inside Pi Durable: every step is a checkpoint

One rule to remember.

Store before you show. Nothing displays before it commits.

Run flow:
submit(input) → generation(model call) → n tool calls → answer
                ↓ checkpoint at every step (SQLite or JSONL)

When the process dies:
a new process opens the same storage → finds unfinished work
→ continues from the last checkpoint (resume)

Cut-off handling:
- Cut model request: sent again (partial answer stays, marked aborted)
- Cut tool call: rerun if safe, else the model is told it was interrupted
- Effects never run twice (replay: never is recorded)

Concretely, a Task is a durable state machine checkpointing each step, with conversations, model turns, tool calls, and your own state committed as units. Storage is SQLite or JSONL, and reopening reads Harness.open(storage) followed by one resume() line. The idea is simple. Assume death, and design the afterlife first.

Cloudflare posted on October 2 about running the Pi Durable harness on its Agents (Durable Objects). This is not a local-only story; cloud runtimes point the same way.

Pi 1.0 itself: the minimal-harness philosophy

Pair it with the Pi 1.0 philosophy. "Many agent harnesses, but this one is yours." Minimal core plus extensions.

Built in: terminal TUI, built-in MCP, Codemode (compose tool calls in a JS sandbox),
          auto-saved sessions, AGENTS.md loader, mid-session model switching
Left out: subagents, plan mode (build with extensions or install packages)
Four modes: interactive / print and JSON / RPC / TypeScript SDK

Sessions auto-save and continue with pi --continue or /resume. That is layer 2 (session memory) from the three-layer memory post shipping as product. Durable adds "even if the process dies" on top.

CodeBridge Mini Lab: design resume points first

Before handing over a long coding task, check this list.

1. Fix checkpoint spots before starting:
   [ ] Snapshots around file edits (align with git commit units)
   [ ] Record test results (keep pass and fail logs)
   [ ] Record intent before external calls (payment, deploy, send)

2. Mark rerun safety:
   [ ] Reads and searches: safe to rerun
   [ ] Writes and edits: checkpoint required
   [ ] Deletes, deploys, payments: never rerun plus human approval

3. Rehearse one kill:
   - Kill the process mid-run on a long task
   - Check resume continues with no duplicated effects
   - Pair with the [WORKLOG pattern](/en/blog/meta-muse-spark-1-3-agentic-coding/)
     so humans can also see where to continue

It is the same table as the approval rules in the four-layer security post. Splitting the safe from the dangerous turns into checkpoint design.

Conclusion: after prompts, design for death

One line to close.

Checkpoints, not prompts, decide long-task quality.

Pi Durable shows the direction clearly. Agent engineering is shifting from model calls to state, recovery, retry, tool execution, and durability. One small task for today: before writing the prompt for your next long job, fix 3 resume points. Where to store, where to continue, what must never run twice. Those three lines separate agents that run all night from ones found broken in the morning.

Further reading

References

Go deeper with a course

If you want to design resumable run structures with verification loops, this course builds harness, loop, and graph layers exactly like the checkpoint design here.