Million-token context windows are now common. Yet something feels off. With all that memory, why does your agent forget yesterday's work when you ask today?
Because context and memory are different things. The million-token post and the RAG cost debate covered context. This post covers memory. Three layers make it clear.
See the three-layer structure first
┌ Layer 3 persistent — survives after sessions end
│ e.g. MEMORY.md lessons, project rules, user prefs
│ stored in: files, DB, vector store / lifetime: near-permanent
├ Layer 2 session — lasts while one work unit continues
│ e.g. WORKLOG (goal, done, next), chat summary, drafts
│ stored in: session state, summary notes / lifetime: per task
└ Layer 1 context — text fed into this inference
e.g. 1M-token window, files read this time
stored in: model input / lifetime: one call
Here is the confusion point. A bigger layer 1 does not solve layer 3. Context is what you see "this time." Memory is what you can recall "next time." The Muse Spark post called long work a state-management problem for the same reason. State never persists by itself.
Layer 1 context: wide but volatile
Opus and Sonnet 5.5, Astra, and Argon all treat 1M as the default. That helps when you load long docs and reason over them at once.
But limits remain.
- It vanishes when the session ends (yesterday's chat is not in today's input)
- More input means more cost and time (tokens = money = time)
- You barely notice when it quietly loses the thread
So use context as a "wide desk." Store what matters in the layers below. Think of desk notes as cleared at closing time.
Layer 2 session: your WORKLOG is session memory
Multi-turn work needs session memory. Goals, constraints, completed steps, next steps, unknowns. The WORKLOG format from the Muse Spark post is exactly this.
## Example WORKLOG
- Goal: fix login form validation bug
- Constraints: do not touch auth module, keep existing tests
- Done: found cause (periods.py boundary bug), added 1 test
- Next: run regression tests, then request review
- Unknowns: deploy schedule (ask a human)
- Verify: full tests pass + read the diff
The enemy of session memory is a reset. When the session breaks, layer 2 disappears. You have two defenses.
Defense 1: save WORKLOG to a file at each big step (keep it human-readable)
Defense 2: read the file first on resume (standardize the continue command)
Example resume prompt:
"Read WORKLOG.md first, verify everything up to Done,
then advance only one Next step.
If anything is in Unknowns, stop and ask."
The delegation pattern in the Qwen subagent post follows the same rule. You must pass state along when you hand work off.
Layer 3 persistent: AGENTS.md, SKILLS.md, MEMORY.md
Long projects need something that outlives sessions. Three docs are becoming standard.
| Doc | Role | What you write |
|---|---|---|
| AGENTS.md | Org chart and behavior policy | Owners, no-touch zones, approval rules |
| SKILLS.md | Technical taste | Preferred stack, banned patterns, style, UI direction |
| MEMORY.md | Long-term lessons | Past bugs and fixes, design intent, never-repeat items |
The permission-file discussion in the MCP comparison and the security four-layer post connects here. Put permissions in AGENTS.md and layer-3 memory becomes a safety device.
Good MEMORY.md entries (examples):
- 2026-09: never merge payments module without tests (outage history)
- Handle period boundaries only as half-open intervals (bug history)
- When adding an external API, write cache + retry policy together (repeated review comment)
Do NOT write:
- One-off task details (that belongs in layer-2 WORKLOG)
- Secrets or keys (memory is not a vault)
CodeBridge Mini Lab: install all three layers in one project
1. AGENTS.md, one page (30 min):
- Define 3 roles (explore, implement, review)
- List 5 banned paths + commands that need approval
2. SKILLS.md, half a page (20 min):
- Stack, style, banned patterns in under 10 lines
3. Start MEMORY.md (10 min):
- Write only your last 3 mistakes
- Rule: repeat a mistake twice, add one line
4. Fix the session routine:
- Start: read WORKLOG → verify → proceed
- End: update WORKLOG + add one MEMORY line if you learned something
This adds a memory layer to the verification loop from real-project practice and the harness post. Do not start big. Three files and two routines are enough.
Conclusion: memory never builds itself
One line to close.
Context is the desk, session is the progress board, persistent memory is the lessons book. Design all three separately.
Before you blame an agent for forgetting yesterday, ask yourself. Did you give it a place to write yesterday down? Three-layer memory starts with three files and two routines. Writing one WORKLOG today is already layer two of layer three.
Further reading
- What Muse Spark 1.3 means for long coding work
- Claude Opus 5.5 and the 1M-token context
- Do you still need RAG with 1M tokens?
References
- Anthropic: Introducing Claude Opus 5.5
- OpenAI Agents SDK: Sessions
- Model Context Protocol: Introduction
Go deeper with a course
If you want to design agent runtimes with memory built in, this course stacks harness, loop, and graph patterns exactly like the three layers here.