AWS Pizza Bot is back in developer conversation this weekend. An open-source app handling agents through All, Unread, and Action inbox shapes instead of a watched chat. Scheduled and webhook tasks, background workers, specialist delegation, human approval, and persistent state all ship. Early versions served 2,000-plus people inside Amazon.

The core UX idea reads well.

Person: assigns the job
Agent: runs in background
Done: Unread
Needs human call: Action

Why an inbox: chats demand presence

The AWS blog sentence nails it. "You don't send an email and then sit watching the outbox until the reply lands." A thread is work you return to, not a session you must attend.

Live-chat premise: both sides present (fits quick exchanges)
Long jobs: minutes long, mid approvals, scheduled runs while away
→ the chat premise breaks
Inbox premise: assign, look away, read finished work and pending calls

Borrowing The Register's phrasing, it works like email. Finished work lands in Unread, calls land in Action, All holds history. A side Activity panel shows delegated specialist work with transcripts.

Structure: the server holds the job

4 layers:
Clients (Electron, browser, terminal)
  → Server (HTTP, run manager)
    → Runtime (DeepAgents and LangGraph, stateful)
      → Storage (SQLite plus files: checkpoints, threads, attachments)
        plus MCP servers, skill workers, your chosen model provider

Traits:
- threads as durable homes for talk plus agent state
- runs survive closes, reloads, device moves (checkpoints underneath)
- one folder holds everything (clear backup unit)
- credentials in the OS secret store, remote needs tokens plus allowed origins

No model lock-in. Anthropic, Bedrock, Gemini, OpenAI, OpenRouter, even local Ollama. It is community open source with no SLA or support, stated plainly. Analysts note its distance from AWS marketing exactly because it avoids Bedrock lock-in.

The skill shape deserves attention. One generalist plus SKILL.md specialists. A meeting-prep skill gets calendar, CRM, and docs tools only. The browser skill gets a browser and nothing else. Short tool lists are the point. Skills can demand approval before tools run, and that pause stays durable, answerable an hour later from another device in the Action queue.

The counterargument, fairly stated

InfoWorld and HFS objections go on record too. Inboxes can hide bad work. Watchers in chat catch derailment live. Inboxes learn late.

Chat: derailment visible at once, but demands watching (babysitting)
Inbox: judge-only attention, but bad work surfaces late
→ no winner, a trade-off; run both by job shape

Execution bundling from the observability post belongs here. The less you watch, the denser records must be.

CodeBridge Mini Lab: design 4 states first

Take the brief's action as is. States before streaming text.

4 states for long-agent UI:
[ ] Task: the assigned job (who, when, what)
[ ] Status: the running job (how far along)
[ ] Action Required: the stopped spot (approvals, answers pending)
[ ] Result: the finished job (read versus unread)

Apply in order:
1. Split 1 of your jobs into 4 states (example: weekly research)
2. Fix 1 approval point (where it stops)
3. Attach Unread and Action pings (when to look)
4. Leave chat for questions only (long jobs go inbox)

The approval card from the OpenDots post and the Queue from the gh-aw post share the shape. Stopping-and-continuing points become UI in this era.

Conclusion: from talk to work

One line to close.

Hour-long jobs outgrow chat windows. Conversation yields to work queues.

Pizza Bot asks plainly. Watch agents or assign and read only what matters. The answer is both. Short jobs chat, long jobs inbox. One task for today: split your longest agent job into the 4 states. Those 4 boxes become the next UI.

Further reading

References

Go deeper with a course

To practice building job-shaped screens and linking them with AI, the course that builds and ships websites through conversation continues exactly into the 4-state UI here.