Some teams send every request to the most expensive model. Money drains away.
Other teams send everything to the cheapest model. They lose time to failures and retries.
The answer sits in the middle. Easy work goes to cheap models, hard work goes to strong models. Splitting them automatically is model routing. This post turns the theory from the Luna, Sol, and Astra guide into code.
Why routing matters now: the middle seat just got filled
Routing got much more practical in late September.
Top tier: Opus 5.5 max 58 pts, Astra max 53 pts, Argon 53 pts
Middle tier: GPT-6.1 Sol 52 pts (under a quarter of Astra's cost, shipped 9/29)
Sonnet 5.5 max 56 pts (tied with Opus xhigh)
Light tier: Luna-class small models + cache + fast responses
-> Now that a "middle to delegate to" exists, a router pays off
GPT-6.1 Sol replacing GPT-6 Sol in 7 days is the headline example. It delivers near-Astra intelligence at a far lower price. That is exactly the middle step a router needs.
Router anatomy: four steps are enough
Request ──> [1. classify] ──> try cheap model ──> [2. confidence + verify]
├─ pass → return + log
└─ fail → [3. escalate] stronger model
└─ [4. log] cost + outcomes saved
Each step in plain terms:
| Step | Job | On failure |
|---|---|---|
| 1. Classify | Estimate difficulty and type (length, expertise, tool needs) | Default to the middle tier |
| 2. Try + verify | Run the cheap model, check tests, format, and evidence | Escalate below the bar |
| 3. Escalate + fall back | Retry with a stronger model, then hand to a human | No infinite retries (cap at 2) |
| 4. Log | Record model, effort, cost, and success | Review weekly, tune escalation rate |
A minimal implementation
This small Python sketch shows the concept. Treat it as a structural reference, not production code.
# router.py — concept sketch
MAX_ESCALATIONS = 2
def classify(task):
"""Estimate difficulty: long, specialized, tool-using tasks count as hard."""
score = 0
if len(task.instructions) > 800:
score += 1
if task.needs_tools or task.needs_long_context:
score += 1
if task.domain in ("legal", "finance", "medical", "security"):
score += 1
if score >= 2:
return "strong" # Opus-class / Astra-class
if score == 1:
return "middle" # GPT-6.1 Sol-class / Sonnet-class
return "light" # Luna-class small
def verify(result, task):
"""Validation gate: only the checks that fit this task."""
checks = []
if task.tests:
checks.append(run_tests(result.patch))
if task.schema:
checks.append(matches_schema(result.output, task.schema))
if task.requires_citation:
checks.append(has_citation(result.output))
return all(checks)
def route(task, call_model, log):
tier = classify(task)
attempts = 0
for candidate in expand(tier): # e.g. light → middle → strong
attempts += 1
result = call_model(candidate, task)
log.record(model=candidate, cost=result.cost,
ok=verify(result, task))
if verify(result, task):
return result
if attempts > MAX_ESCALATIONS:
break
return escalate_to_human(task)
Three takeaways. Keep classification simple, vary validation per task, and cap retries. The router doesn't need to be smart. Strict verification lets a dumb router work.
Caching: the router's best friend
Prompt caching pairs beautifully with routing. Put repeated system instructions and documents in the cache to save input processing.
Cache-friendly prompt layout:
[fixed prefix: role, rules, frequently used docs] → cache hit
[variable suffix: this request, this input] → fresh each time
Effect (vendor announcements as of October):
- Opus 5.5 cache reads $0.20 (roughly 95% off)
- Gemini 4 Argon cache 95% off
- GPT-6 Astra cache reads 90% off
Attach the math from the prompt caching guide to your router logs. Tier costs plus cache hit rates show you exactly where money leaks.
CodeBridge Mini Lab: aim for a 20% escalation rate
1. Start with 2 tiers (light + strong):
- Classify with 2 rules only: length + tool use
- Verify with tests only: 1 pass/fail check
2. Judge after 2 weeks of logs:
- Escalation above 50% → classification is too optimistic, tighten it
- Escalation below 5% → strong tier is idle, sample the hard tasks
- Target: 10–20% escalation with success rate intact
3. Then add a 3rd tier (middle):
- Place a GPT-6.1 Sol-class model as escalation step 1
- Check whether cost per success dropped
As the GPT-5.6 Sol post and the GPT-6 Sol vs Luna post argue, using the strongest model for everything is wasteful. An escalation ladder fixes exactly that.
Related posts
- From Luna to Sol to Astra: stop using the top model for everything
- GPT-6 Sol vs Luna: why you don't always need the pricey model
- Cut AI API costs with prompt caching
References
- Artificial Analysis: GPT-6.1 Sol replaces GPT-6 Sol
- Artificial Analysis: Benchmarking GPT-6 Astra
- Artificial Analysis: Gemini 4 Argon
- OpenAI Agents SDK Documentation
Conclusion: a router is evaluation plus budget plus fallback
One last myth to clear. A router doesn't save money through brilliant classification. Verification catches failures, caps stop runaway loops, and logs change next week's decisions. That structure saves the money.
Start small today. Two tiers, one check, a retry cap of 2. Two weeks of logs will show you which middle model to buy next.
Go deeper with a course
If you want guided practice designing classify-delegate-verify systems, a structured course builds the habit faster than blog posts alone.