AI products mix easy and hard requests.

"Convert this sentence to JSON"
"Find the cause of this repo outage and fix it"

Using the strongest model for both is simple. It is often not cost-efficient.

That is why model routing matters.

You can split three models by role

The GPT-6 family offers layers like Luna, Sol, and Astra with different cost and strength.

As a starting concept, split them like this:

Luna  → simple classification, extraction, short rewrites
Sol   → complex coding, agent workflows
Astra → hardest end-to-end work

This is not an absolute rule. It is a starting point for routing design.

The price gaps are very large

Under OpenAI standard API pricing from September 2026, short-context input and output prices differ sharply by model.

So sending everything to Astra versus starting in Luna and escalating some calls can create totally different cost structures.

But sending everything to the cheapest model is not the answer either. Failures and retries can raise costs instead.

A cheap first try still pays the input tokens, the wait time, and the retry. If Luna fails half your code tasks and each failure triggers a Sol rerun, you pay for two runs plus user delay. That total can pass a single careful Sol run. So always compare full paths: Luna alone, Sol alone, and Luna with escalation. Track success rate, retries, latency, and spend together before you lock a default.

Start with static routing

You do not need a complex AI router at first.

if task_type in {"extract", "classify", "rewrite"}:
    model = "luna"
elif task_type in {"code", "analysis"}:
    model = "sol"
else:
    model = "astra"

If you already know the task type, this is easiest to explain and operate.

A better way: failure-based escalation

One step further, let verification results decide routing.

Luna
  ↓ schema validation fails
Sol
  ↓ tests / rubric fail
Astra

This approach uses real success criteria instead of guessing "this request looks hard."

CodeBridge Mini Lab: simulate routing with 100 requests

Pull 100 past requests and record:

task,type,luna_ok,sol_ok,astra_ok,luna_cost,sol_cost,astra_cost
1,extract,1,1,1,0.001,0.01,0.05
2,code,0,1,1,0.003,0.08,0.31

Then compare three strategies:

A. Always Astra
B. Always Sol
C. Luna → Sol → Astra escalation

Compare on:

  • Total success rate
  • Total API cost
  • Average latency
  • Escalation rate

Even this much math shows whether routing truly pays.

Add one more column: cost per success. Divide each strategy's total spend by its successful tasks. Strategy C often wins here even when its raw success rate sits between A and B, because it spends top-model money on a small slice of traffic. Also note which task types escalate most. If code tasks escalate 60 percent of the time while extraction escalates 5 percent, move code to Sol by default and keep Luna for extraction. Your routing table should learn from that split.

A complex router adds its own cost

If another LLM call classifies each request, routing cost and failure risk grow.

So start in this order:

Rules
→ validation-based escalation
→ learned / LLM router only if needed

Routing also applies to effort, not only models

You can stage effort inside the same model:

Sol medium
→ Sol high
→ Sol max
→ Astra high

Models plus reasoning effort give you a finer cost-quality frontier.

Start cheap on both axes and climb one step at a time. A routine extract runs on Luna low. A draft comparison moves to Luna medium. Code review steps up to Sol medium, then Sol high only when tests fail. Reserve Astra high and max for the few tasks where a mistake costs real money or ships to customers. Each step should have a named check: schema valid, tests green, rubric score above your bar. No check, no promotion.

Conclusion: a good escalation policy can beat one best model

Your goal is not to use the leaderboard winner. It is to build a cost structure that meets your quality bar.

So change the question.

Which model should we use for every request?

Ask this more operational question instead:

Which failure signal should promote us to the next model?

Further reading

References

Go deeper with a course

If you want to turn routing rules and escalation checks into a repeatable team habit, practice with multi-model workflows.