One number made the rounds on October 3. OpenRouter-based a16z data. August agent traffic at 7.3 trillion tokens against 1.4 trillion human. About a 5x gap. Agents first passed humans in February 2026 and the gap keeps widening, with 10x forecasts attached.
Scope first, precisely. This compares traffic inside OpenRouter, not the whole AI market. The direction still stands. And a more interesting figure hides beside it. Over 85% of agent tokens are cached prompts.
What 85% means: agents are rereading machines
Agent token mix (estimated):
- 85%+: cached prompts (past turns, tool results, context reinjected)
- rest: fresh input plus generation
Agents reread what they have seen
about as much as they create answers
Borrowing OpenRouter's own phrasing, agents resend the same system prompts, tool definitions, and schemas every turn. Cache reads cost 0.1 to 0.5x, which looks cheap, but they occupy KV cache and HBM. Chipmakers expecting HBM shortages deepening through 2028 count this among the causes.
The speed guide said token counts equal time. Now it is a space problem too. As long-running agents multiply, context engineering itself turns into infrastructure trouble.
Why model prices alone miss it
Old optimization: pick the cheap model (compare price sheets)
New optimization: cut rereading (compare structures)
Same job, different shapes:
A: full turns plus long tool logs reinjected → 90% cached, still slow and heavy
B: summary plus deltas only → 60% cached, fast and light
→ B can lose the unit-price game and win total cost and time
One line joins the prompt caching math. Watch the cache "share," not just the cache discount. 85% cached means caching works and simultaneously that "85% gets reread."
That is why OpenRouter pushes sticky routing, landing repeat requests on the provider holding the warm cache. Cache without sticky routing never stays warm. Caching, routing, and observability ship as one set.
The 4 to track together: start at tokens per task
Take the brief's action as is.
Weekly card for 1 agent:
[ ] tokens/task: total tokens per job (counted on success)
[ ] cache ratio: cached-read share (85% is average, above means suspect reinjection)
[ ] retries: retry counts (1 retry equals 1 full reread)
[ ] success rate: meaningless unless read with the 3 above
It is cost per success plus cache ratio and retries. Retries scare most. One retry equals rereading "everything so far" once.
CodeBridge Mini Lab: 1 week of reinjection diet
1. Find the 1 longest tool log (the token hog)
2. Change only this for 1 week:
- full-turn reinjection → summary plus deltas
- long tool output → trimmed to needed parts
- full tool definitions each turn → load only what runs
3. Compare:
[ ] tokens/task delta
[ ] cache ratio delta
[ ] success rate delta (roll back on drops — cheap must not mean broken)
4. Call it: lower tokens with held success means keep it, else revert
The 193k-token story from the Sonnet high-effort post lives here. Token-heavy models gain most from reinjection diets.
Conclusion: design the reading
One line to close.
Count the reading before optimizing the writing.
Agent optimization in the 5x era is context design, not model shopping. Track tokens per task, cache ratio, retries, and success rate together, and cut long reinjection first. Cache is infrastructure, not a discount trick. Teams that treat it like infrastructure ride out the next cost wave.
Further reading
- Cut AI API cost with prompt caching
- AI model speed and latency guide
- Judge AI model cost per successful task, not token price
References
- Gate News: AI Agents Use 5x More Tokens Than Humans (Oct 3, 2026)
- OpenRouter: The Cheapest Token Is a Cached One
- OpenRouter Docs: Prompt Caching
Go deeper with a course
To practice designing tokens, verification, and loops as a structure, this course builds harness, loop, and graph layers exactly like the 4 metrics here.