Some teams send every request to the most expensive model. Money drains away.

Other teams send everything to the cheapest model. They lose time to failures and retries.

The answer sits in the middle. Easy work goes to cheap models, hard work goes to strong models. Splitting them automatically is model routing. This post turns the theory from the Luna, Sol, and Astra guide into code.

Why routing matters now: the middle seat just got filled

Routing got much more practical in late September.

Top tier:    Opus 5.5 max 58 pts, Astra max 53 pts, Argon 53 pts
Middle tier: GPT-6.1 Sol 52 pts (under a quarter of Astra's cost, shipped 9/29)
             Sonnet 5.5 max 56 pts (tied with Opus xhigh)
Light tier:  Luna-class small models + cache + fast responses

-> Now that a "middle to delegate to" exists, a router pays off

GPT-6.1 Sol replacing GPT-6 Sol in 7 days is the headline example. It delivers near-Astra intelligence at a far lower price. That is exactly the middle step a router needs.

Router anatomy: four steps are enough

Request ──> [1. classify] ──> try cheap model ──> [2. confidence + verify]
                                                  ├─ pass → return + log
                                                  └─ fail → [3. escalate] stronger model
                                                              └─ [4. log] cost + outcomes saved

Each step in plain terms:

Step Job On failure
1. Classify Estimate difficulty and type (length, expertise, tool needs) Default to the middle tier
2. Try + verify Run the cheap model, check tests, format, and evidence Escalate below the bar
3. Escalate + fall back Retry with a stronger model, then hand to a human No infinite retries (cap at 2)
4. Log Record model, effort, cost, and success Review weekly, tune escalation rate

A minimal implementation

This small Python sketch shows the concept. Treat it as a structural reference, not production code.

 # router.py — concept sketch
MAX_ESCALATIONS = 2

def classify(task):
    """Estimate difficulty: long, specialized, tool-using tasks count as hard."""
    score = 0
    if len(task.instructions) > 800:
        score += 1
    if task.needs_tools or task.needs_long_context:
        score += 1
    if task.domain in ("legal", "finance", "medical", "security"):
        score += 1
    if score >= 2:
        return "strong"     # Opus-class / Astra-class
    if score == 1:
        return "middle"     # GPT-6.1 Sol-class / Sonnet-class
    return "light"          # Luna-class small

def verify(result, task):
    """Validation gate: only the checks that fit this task."""
    checks = []
    if task.tests:
        checks.append(run_tests(result.patch))
    if task.schema:
        checks.append(matches_schema(result.output, task.schema))
    if task.requires_citation:
        checks.append(has_citation(result.output))
    return all(checks)

def route(task, call_model, log):
    tier = classify(task)
    attempts = 0
    for candidate in expand(tier):  # e.g. light → middle → strong
        attempts += 1
        result = call_model(candidate, task)
        log.record(model=candidate, cost=result.cost,
                   ok=verify(result, task))
        if verify(result, task):
            return result
        if attempts > MAX_ESCALATIONS:
            break
    return escalate_to_human(task)

Three takeaways. Keep classification simple, vary validation per task, and cap retries. The router doesn't need to be smart. Strict verification lets a dumb router work.

Caching: the router's best friend

Prompt caching pairs beautifully with routing. Put repeated system instructions and documents in the cache to save input processing.

Cache-friendly prompt layout:
[fixed prefix: role, rules, frequently used docs] → cache hit
[variable suffix: this request, this input] → fresh each time

Effect (vendor announcements as of October):
- Opus 5.5 cache reads $0.20 (roughly 95% off)
- Gemini 4 Argon cache 95% off
- GPT-6 Astra cache reads 90% off

Attach the math from the prompt caching guide to your router logs. Tier costs plus cache hit rates show you exactly where money leaks.

CodeBridge Mini Lab: aim for a 20% escalation rate

1. Start with 2 tiers (light + strong):
   - Classify with 2 rules only: length + tool use
   - Verify with tests only: 1 pass/fail check

2. Judge after 2 weeks of logs:
   - Escalation above 50% → classification is too optimistic, tighten it
   - Escalation below 5% → strong tier is idle, sample the hard tasks
   - Target: 10–20% escalation with success rate intact

3. Then add a 3rd tier (middle):
   - Place a GPT-6.1 Sol-class model as escalation step 1
   - Check whether cost per success dropped

As the GPT-5.6 Sol post and the GPT-6 Sol vs Luna post argue, using the strongest model for everything is wasteful. An escalation ladder fixes exactly that.

Related posts

References

Conclusion: a router is evaluation plus budget plus fallback

One last myth to clear. A router doesn't save money through brilliant classification. Verification catches failures, caps stop runaway loops, and logs change next week's decisions. That structure saves the money.

Start small today. Two tiers, one check, a retry cap of 2. Two weeks of logs will show you which middle model to buy next.

Go deeper with a course

If you want guided practice designing classify-delegate-verify systems, a structured course builds the habit faster than blog posts alone.