API price tables invite a simple comparison.

Model A: input $0.10 / output $0.50
Model B: input $2 / output $10

Model A looks 20x cheaper.

Real work misses one thing.

Did the task succeed?

A cheap model that fails often and needs two or three reruns can end up far less cheap.

Token price and task cost are different

Artificial Analysis uses Cost per Task in model comparisons.

It does not quote unit API rates. It computes average cost for one benchmark workload from actual input, cache, and output tokens.

That matters because models differ in answer length and reasoning-token appetite.

Same $1 / 1M tokens

Model A → uses 5,000 tokens
Model B → uses 30,000 tokens

Real costs differ

For production, go one step further down.

The CodeBridge version: one more calculation

The simplest useful metric is this.

Effective Cost per Success
=
Total run cost / Successful tasks

Say you ran models A and B ten times each.

Model Total cost Wins Cost per win
A $0.80 5 $0.16
B $1.60 9 about $0.18

Model A costs half overall. Per success, the gap nearly vanishes.

For important work, add failure costs too.

Total Effective Cost
=
API cost
+
Retry cost
+
Human review time
+
Recovery cost from failures

You do not need exact dollars for everything. Direction alone helps.

GPT-6 Sol and Luna make a great example

As of September 2026, GPT-6 Luna costs far less than Sol per API token.

At max effort, Artificial Analysis Cost per Intelligence task looks like:

GPT-6 Luna   about $0.07
GPT-6 Sol    about $1.06

A huge gap.

But Sol stays stronger on overall quality and complex agentic work.

So both extremes waste money:

Every request → Sol

is inefficient, and:

Every request → Luna

can also waste money through failures.

The better question is:

Up to which task difficulty can Luna hold a good enough success rate?

CodeBridge Mini Lab: compute real cost with one CSV

Run a small experiment.

Prepare 20 tasks and log this data.

task_id,model,cost,success,retries,time_sec
1,luna,0.01,1,0,8
2,luna,0.02,0,2,31
3,sol,0.18,1,0,22

Python makes the math trivial.

import pandas as pd

df = pd.read_csv("runs.csv")

summary = (
    df.groupby("model")
      .agg(
          total_cost=("cost", "sum"),
          successes=("success", "sum"),
          avg_time=("time_sec", "mean"),
          avg_retries=("retries", "mean"),
      )
)

summary["cost_per_success"] = (
    summary["total_cost"] / summary["successes"]
)

print(summary)

The column that matters most is cost_per_success.

But define success first.

For coding:

All tests pass
Lint passes
No unexpected file changes

For document generation:

Required sections present
Evidence links present
No banned phrases

Auto-checkable conditions work best.

Why you should not optimize cost alone

Cost per success is still imperfect.

Two models can share a success rate while feeling totally different:

  • One takes 5 seconds
  • The other takes 3 minutes

User experience splits hard there.

So track at least these four together.

Success rate
Cost per success
Latency
Human review time

Those four beat "this token rate is cheaper" by a wide margin.

Why routing becomes the core cost lever

Not all work shares one difficulty.

Finding one model matters less than splitting traffic like this.

Simple tasks
→ low-cost model

Failed verification
→ stronger model

Still failing
→ human review

Here the price of one model matters less than the expected cost of the whole path.

Conclusion: find the cheapest success path, not the cheapest model

The easiest way to cut AI spend is not always picking the cheapest model.

Measure this instead.

How much does it cost on average to finish my task end to end?

One small CSV can change your selection logic.

Once those numbers accumulate, routing, effort tuning, and fallback become natural next steps.

Further reading

References

Go deeper with a course

If you want a repeatable workflow for matching model strength to task difficulty, this course teaches practical tool selection by situation.