GPT-6.1 Sol was released on September 29, 2026.

At first, the price table tells a simple story.

GPT-6.1 Sol
Input  $2 / 1M
Output $10 / 1M

GPT-6 Astra
Input  $10 / 1M
Output $50 / 1M

At standard prices, that is exactly a 5x gap.

So is the conclusion simple too?

"Sol is 5x cheaper than Astra, so just use Sol."

Real agent costs do not work that way.

If a model fails and you run it twice more, or you raise reasoning effort and burn far more tokens, or tool calls grow, the answer changes.

So the more interesting question for GPT-6.1 Sol is this:

Not what one token costs, but what one successful task costs.

The specs are closer than you expect

Using the OpenAI API docs, the two models look like this.

Item GPT-6.1 Sol GPT-6 Astra
Input / 1M $2 $10
Cached input / 1M $0.10 $1
Output / 1M $10 $50
Context 1.05M 1.05M
Max output 128K 128K
Knowledge cutoff 2026-04-30 2026-04-30

Both support long context and tool use.

So Astra is not expensive just because it is "a longer model."

OpenAI's current model selection docs split the roles roughly like this.

Astra
Hardest, quality-first work

GPT-6.1 Sol
Complex work where you must also manage cost and time

Luna
Narrow, high-frequency repeat work

Model selection itself is becoming a routing problem.

One small GPT-6.1 Sol change matters more than it looks

Compared with GPT-6 Sol, base input and output prices are the same.

But cached input changed:

GPT-6 Sol      $0.20 / 1M
GPT-6.1 Sol    $0.10 / 1M

It dropped by half.

In agentic coding, the same repo description, system instructions, and tool schemas repeat often.

So with a high cache hit rate, a small price gap compounds.

Long shared context
+ Repeated runs
+ High cache hits
=
Bigger real cost gap

Independent tests also show "just below Astra"

In Artificial Analysis measurements from September 29, GPT-6.1 Sol Max scored 52 on the Intelligence Index, 1 point below GPT-6 Astra.

Cost per Intelligence Index task in the same analysis was:

GPT-6.1 Sol Max   about $0.72
GPT-6 Astra       about $3.26

So independent tests also show a zone where the task-cost gap is far bigger than the quality gap.

But do not simplify this to "Sol beat Astra."

Results change with reasoning effort.

For example, in the Artificial Analysis comparison:

Sol Medium
Intelligence Index 48
Cost per task about $0.21

Astra Xhigh
Intelligence Index 52
Cost per task about $2.31

You see a zone where buying a little more quality costs a lot more.

And on the hardest tasks, if a 4-point gap prevents real failures, Astra can still be cheaper.

Separate cost per task from cost per success

Split model cost into three stages. It becomes clearer.

Token price
↓
Cost of one run

Cost per Task
↓
Average cost when you evaluate one task

Cost per Success
↓
Cost of getting one real success

For example:

Sol
10 runs
8 successes
Total cost $2.40

Cost per Success = $0.30


Astra
10 runs
10 successes
Total cost $4.00

Cost per Success = $0.40

Here Sol is more economical.

But if Sol succeeds only 4 times:

$2.40 / 4
= $0.60 per success

Then Astra becomes cheaper.

That is why you cannot pick a model from the public price table alone.

You can calculate it with simple Python

You only need 10 to 20 real tasks.

task_id,model,cost,success,time_sec
1,gpt-6.1-sol,0.12,1,42
2,gpt-6.1-sol,0.09,1,31
3,gpt-6.1-sol,0.18,0,73
4,gpt-6-astra,0.41,1,55
5,gpt-6-astra,0.38,1,48
import pandas as pd

df = pd.read_csv("runs.csv")

summary = df.groupby("model").agg(
    total_cost=("cost", "sum"),
    successes=("success", "sum"),
    avg_time=("time_sec", "mean"),
)

summary["cost_per_success"] = (
    summary["total_cost"] / summary["successes"]
)

print(summary)

The most important column here is cost_per_success.

You get the price of a finished result, not the token price.

As a diagram, model routing feels natural

Instead of sending everything to Astra from the start, you can do this:

Task input
   ↓
GPT-6.1 Sol
   ↓
Pass completion checks?
  ├─ Yes → done
  └─ No
      ↓
   Escalate to Astra
      ↓
   Re-run / review

This is not a "always use the cheap model" strategy.

It is a strategy to finish cheap-enough work in Sol and pay Astra prices only when you truly need stronger capability.

Reasoning effort matters as much as model choice

GPT-6.1 Sol supports:

low
medium
high
xhigh
max

Even with the same model name, cost, time, and quality can shift a lot when effort changes.

So do not compare like this:

Sol vs Astra

In practice, it is more accurate to see this as one set:

Sol Medium
Sol High
Sol Max
Astra Medium
Astra Xhigh

That is why effort labels matter when you read benchmark scores.

CodeBridge Mini Lab: find the tasks that need Astra

Pick about 15 tasks from your project.

Mix difficulty levels if you can.

5 — small fixes
5 — multi-file changes
5 — tasks needing analysis + implementation + tests

Then set the same completion rules.

- All tests pass
- Lint passes
- No public API changes
- No unexpected file changes

Run Sol Medium first.

Promote only failures to Sol High, and only remaining failures to Astra.

Sol Medium
    ↓ fail
Sol High
    ↓ fail
Astra

That gives you a more useful answer than "Is Astra good?"

What kind of work in my project actually needs Astra?

That is the start of your routing rule.

The bottom line: the cheapest success path beats the best model

GPT-6.1 Sol is less a small update to GPT-6 Sol. It makes one question sharper: how cheaply can you reach near-Astra capability?

Current pricing and independent benchmarks suggest Sol can cover a wide area.

But you must confirm the final choice in your own workflow, not in a public score table.

Model price
+
Reasoning effort
+
Cache hits
+
Tool calls
+
Retries
+
Success rate
=
Real task cost

With this formula, "Which model is best?" matters less than "Which path succeeds most economically?"

Further reading

References

Go deeper with a course

If you want to build agents that stay stable inside context, rules, tools, and verification loops — not just swap models — work through a harness course.