GPT-6.1 Sol was released on September 29, 2026.
At first, the price table tells a simple story.
GPT-6.1 Sol
Input $2 / 1M
Output $10 / 1M
GPT-6 Astra
Input $10 / 1M
Output $50 / 1M
At standard prices, that is exactly a 5x gap.
So is the conclusion simple too?
"Sol is 5x cheaper than Astra, so just use Sol."
Real agent costs do not work that way.
If a model fails and you run it twice more, or you raise reasoning effort and burn far more tokens, or tool calls grow, the answer changes.
So the more interesting question for GPT-6.1 Sol is this:
Not what one token costs, but what one successful task costs.
The specs are closer than you expect
Using the OpenAI API docs, the two models look like this.
| Item | GPT-6.1 Sol | GPT-6 Astra |
|---|---|---|
| Input / 1M | $2 | $10 |
| Cached input / 1M | $0.10 | $1 |
| Output / 1M | $10 | $50 |
| Context | 1.05M | 1.05M |
| Max output | 128K | 128K |
| Knowledge cutoff | 2026-04-30 | 2026-04-30 |
Both support long context and tool use.
So Astra is not expensive just because it is "a longer model."
OpenAI's current model selection docs split the roles roughly like this.
Astra
Hardest, quality-first work
GPT-6.1 Sol
Complex work where you must also manage cost and time
Luna
Narrow, high-frequency repeat work
Model selection itself is becoming a routing problem.
One small GPT-6.1 Sol change matters more than it looks
Compared with GPT-6 Sol, base input and output prices are the same.
But cached input changed:
GPT-6 Sol $0.20 / 1M
GPT-6.1 Sol $0.10 / 1M
It dropped by half.
In agentic coding, the same repo description, system instructions, and tool schemas repeat often.
So with a high cache hit rate, a small price gap compounds.
Long shared context
+ Repeated runs
+ High cache hits
=
Bigger real cost gap
Independent tests also show "just below Astra"
In Artificial Analysis measurements from September 29, GPT-6.1 Sol Max scored 52 on the Intelligence Index, 1 point below GPT-6 Astra.
Cost per Intelligence Index task in the same analysis was:
GPT-6.1 Sol Max about $0.72
GPT-6 Astra about $3.26
So independent tests also show a zone where the task-cost gap is far bigger than the quality gap.
But do not simplify this to "Sol beat Astra."
Results change with reasoning effort.
For example, in the Artificial Analysis comparison:
Sol Medium
Intelligence Index 48
Cost per task about $0.21
Astra Xhigh
Intelligence Index 52
Cost per task about $2.31
You see a zone where buying a little more quality costs a lot more.
And on the hardest tasks, if a 4-point gap prevents real failures, Astra can still be cheaper.
Separate cost per task from cost per success
Split model cost into three stages. It becomes clearer.
Token price
↓
Cost of one run
Cost per Task
↓
Average cost when you evaluate one task
Cost per Success
↓
Cost of getting one real success
For example:
Sol
10 runs
8 successes
Total cost $2.40
Cost per Success = $0.30
Astra
10 runs
10 successes
Total cost $4.00
Cost per Success = $0.40
Here Sol is more economical.
But if Sol succeeds only 4 times:
$2.40 / 4
= $0.60 per success
Then Astra becomes cheaper.
That is why you cannot pick a model from the public price table alone.
You can calculate it with simple Python
You only need 10 to 20 real tasks.
task_id,model,cost,success,time_sec
1,gpt-6.1-sol,0.12,1,42
2,gpt-6.1-sol,0.09,1,31
3,gpt-6.1-sol,0.18,0,73
4,gpt-6-astra,0.41,1,55
5,gpt-6-astra,0.38,1,48
import pandas as pd
df = pd.read_csv("runs.csv")
summary = df.groupby("model").agg(
total_cost=("cost", "sum"),
successes=("success", "sum"),
avg_time=("time_sec", "mean"),
)
summary["cost_per_success"] = (
summary["total_cost"] / summary["successes"]
)
print(summary)
The most important column here is cost_per_success.
You get the price of a finished result, not the token price.
As a diagram, model routing feels natural
Instead of sending everything to Astra from the start, you can do this:
Task input
↓
GPT-6.1 Sol
↓
Pass completion checks?
├─ Yes → done
└─ No
↓
Escalate to Astra
↓
Re-run / review
This is not a "always use the cheap model" strategy.
It is a strategy to finish cheap-enough work in Sol and pay Astra prices only when you truly need stronger capability.
Reasoning effort matters as much as model choice
GPT-6.1 Sol supports:
low
medium
high
xhigh
max
Even with the same model name, cost, time, and quality can shift a lot when effort changes.
So do not compare like this:
Sol vs Astra
In practice, it is more accurate to see this as one set:
Sol Medium
Sol High
Sol Max
Astra Medium
Astra Xhigh
That is why effort labels matter when you read benchmark scores.
CodeBridge Mini Lab: find the tasks that need Astra
Pick about 15 tasks from your project.
Mix difficulty levels if you can.
5 — small fixes
5 — multi-file changes
5 — tasks needing analysis + implementation + tests
Then set the same completion rules.
- All tests pass
- Lint passes
- No public API changes
- No unexpected file changes
Run Sol Medium first.
Promote only failures to Sol High, and only remaining failures to Astra.
Sol Medium
↓ fail
Sol High
↓ fail
Astra
That gives you a more useful answer than "Is Astra good?"
What kind of work in my project actually needs Astra?
That is the start of your routing rule.
The bottom line: the cheapest success path beats the best model
GPT-6.1 Sol is less a small update to GPT-6 Sol. It makes one question sharper: how cheaply can you reach near-Astra capability?
Current pricing and independent benchmarks suggest Sol can cover a wide area.
But you must confirm the final choice in your own workflow, not in a public score table.
Model price
+
Reasoning effort
+
Cache hits
+
Tool calls
+
Retries
+
Success rate
=
Real task cost
With this formula, "Which model is best?" matters less than "Which path succeeds most economically?"
Further reading
- AI model cost means cost per successful task
- Reasoning high vs max: is thinking longer always better?
- Luna to Sol to Astra model routing strategy
- Why benchmark scores alone can mislead you
References
- OpenAI API: GPT-6.1 Sol
- OpenAI API: Model selection
- OpenAI API Pricing
- Artificial Analysis: GPT-6.1 Sol
Go deeper with a course
If you want to build agents that stay stable inside context, rules, tools, and verification loops — not just swap models — work through a harness course.