API price tables invite a simple comparison.
Model A: input $0.10 / output $0.50
Model B: input $2 / output $10
Model A looks 20x cheaper.
Real work misses one thing.
Did the task succeed?
A cheap model that fails often and needs two or three reruns can end up far less cheap.
Token price and task cost are different
Artificial Analysis uses Cost per Task in model comparisons.
It does not quote unit API rates. It computes average cost for one benchmark workload from actual input, cache, and output tokens.
That matters because models differ in answer length and reasoning-token appetite.
Same $1 / 1M tokens
Model A → uses 5,000 tokens
Model B → uses 30,000 tokens
Real costs differ
For production, go one step further down.
The CodeBridge version: one more calculation
The simplest useful metric is this.
Effective Cost per Success
=
Total run cost / Successful tasks
Say you ran models A and B ten times each.
| Model | Total cost | Wins | Cost per win |
|---|---|---|---|
| A | $0.80 | 5 | $0.16 |
| B | $1.60 | 9 | about $0.18 |
Model A costs half overall. Per success, the gap nearly vanishes.
For important work, add failure costs too.
Total Effective Cost
=
API cost
+
Retry cost
+
Human review time
+
Recovery cost from failures
You do not need exact dollars for everything. Direction alone helps.
GPT-6 Sol and Luna make a great example
As of September 2026, GPT-6 Luna costs far less than Sol per API token.
At max effort, Artificial Analysis Cost per Intelligence task looks like:
GPT-6 Luna about $0.07
GPT-6 Sol about $1.06
A huge gap.
But Sol stays stronger on overall quality and complex agentic work.
So both extremes waste money:
Every request → Sol
is inefficient, and:
Every request → Luna
can also waste money through failures.
The better question is:
Up to which task difficulty can Luna hold a good enough success rate?
CodeBridge Mini Lab: compute real cost with one CSV
Run a small experiment.
Prepare 20 tasks and log this data.
task_id,model,cost,success,retries,time_sec
1,luna,0.01,1,0,8
2,luna,0.02,0,2,31
3,sol,0.18,1,0,22
Python makes the math trivial.
import pandas as pd
df = pd.read_csv("runs.csv")
summary = (
df.groupby("model")
.agg(
total_cost=("cost", "sum"),
successes=("success", "sum"),
avg_time=("time_sec", "mean"),
avg_retries=("retries", "mean"),
)
)
summary["cost_per_success"] = (
summary["total_cost"] / summary["successes"]
)
print(summary)
The column that matters most is cost_per_success.
But define success first.
For coding:
All tests pass
Lint passes
No unexpected file changes
For document generation:
Required sections present
Evidence links present
No banned phrases
Auto-checkable conditions work best.
Why you should not optimize cost alone
Cost per success is still imperfect.
Two models can share a success rate while feeling totally different:
- One takes 5 seconds
- The other takes 3 minutes
User experience splits hard there.
So track at least these four together.
Success rate
Cost per success
Latency
Human review time
Those four beat "this token rate is cheaper" by a wide margin.
Why routing becomes the core cost lever
Not all work shares one difficulty.
Finding one model matters less than splitting traffic like this.
Simple tasks
→ low-cost model
Failed verification
→ stronger model
Still failing
→ human review
Here the price of one model matters less than the expected cost of the whole path.
Conclusion: find the cheapest success path, not the cheapest model
The easiest way to cut AI spend is not always picking the cheapest model.
Measure this instead.
How much does it cost on average to finish my task end to end?
One small CSV can change your selection logic.
Once those numbers accumulate, routing, effort tuning, and fallback become natural next steps.
Further reading
- GPT-6 Sol vs Luna: why you do not always need the expensive model
- Claude Opus 5.5 vs GPT-6 Astra
- How to split work across multiple AI tools
References
- Artificial Analysis: Benchmarking Methodology
- Artificial Analysis: GPT-6 Sol and Luna
- OpenAI API Pricing
Go deeper with a course
If you want a repeatable workflow for matching model strength to task difficulty, this course teaches practical tool selection by situation.