AI products mix easy and hard requests.
"Convert this sentence to JSON"
"Find the cause of this repo outage and fix it"
Using the strongest model for both is simple. It is often not cost-efficient.
That is why model routing matters.
You can split three models by role
The GPT-6 family offers layers like Luna, Sol, and Astra with different cost and strength.
As a starting concept, split them like this:
Luna → simple classification, extraction, short rewrites
Sol → complex coding, agent workflows
Astra → hardest end-to-end work
This is not an absolute rule. It is a starting point for routing design.
The price gaps are very large
Under OpenAI standard API pricing from September 2026, short-context input and output prices differ sharply by model.
So sending everything to Astra versus starting in Luna and escalating some calls can create totally different cost structures.
But sending everything to the cheapest model is not the answer either. Failures and retries can raise costs instead.
A cheap first try still pays the input tokens, the wait time, and the retry. If Luna fails half your code tasks and each failure triggers a Sol rerun, you pay for two runs plus user delay. That total can pass a single careful Sol run. So always compare full paths: Luna alone, Sol alone, and Luna with escalation. Track success rate, retries, latency, and spend together before you lock a default.
Start with static routing
You do not need a complex AI router at first.
if task_type in {"extract", "classify", "rewrite"}:
model = "luna"
elif task_type in {"code", "analysis"}:
model = "sol"
else:
model = "astra"
If you already know the task type, this is easiest to explain and operate.
A better way: failure-based escalation
One step further, let verification results decide routing.
Luna
↓ schema validation fails
Sol
↓ tests / rubric fail
Astra
This approach uses real success criteria instead of guessing "this request looks hard."
CodeBridge Mini Lab: simulate routing with 100 requests
Pull 100 past requests and record:
task,type,luna_ok,sol_ok,astra_ok,luna_cost,sol_cost,astra_cost
1,extract,1,1,1,0.001,0.01,0.05
2,code,0,1,1,0.003,0.08,0.31
Then compare three strategies:
A. Always Astra
B. Always Sol
C. Luna → Sol → Astra escalation
Compare on:
- Total success rate
- Total API cost
- Average latency
- Escalation rate
Even this much math shows whether routing truly pays.
Add one more column: cost per success. Divide each strategy's total spend by its successful tasks. Strategy C often wins here even when its raw success rate sits between A and B, because it spends top-model money on a small slice of traffic. Also note which task types escalate most. If code tasks escalate 60 percent of the time while extraction escalates 5 percent, move code to Sol by default and keep Luna for extraction. Your routing table should learn from that split.
A complex router adds its own cost
If another LLM call classifies each request, routing cost and failure risk grow.
So start in this order:
Rules
→ validation-based escalation
→ learned / LLM router only if needed
Routing also applies to effort, not only models
You can stage effort inside the same model:
Sol medium
→ Sol high
→ Sol max
→ Astra high
Models plus reasoning effort give you a finer cost-quality frontier.
Start cheap on both axes and climb one step at a time. A routine extract runs on Luna low. A draft comparison moves to Luna medium. Code review steps up to Sol medium, then Sol high only when tests fail. Reserve Astra high and max for the few tasks where a mistake costs real money or ships to customers. Each step should have a named check: schema valid, tests green, rubric score above your bar. No check, no promotion.
Conclusion: a good escalation policy can beat one best model
Your goal is not to use the leaderboard winner. It is to build a cost structure that meets your quality bar.
So change the question.
Which model should we use for every request?
Ask this more operational question instead:
Which failure signal should promote us to the next model?
Further reading
References
Go deeper with a course
If you want to turn routing rules and escalation checks into a repeatable team habit, practice with multi-model workflows.