Model selection screens now carry one more unfamiliar option next to the model name: reasoning effort, with levels like low, medium, high, xhigh, and max.
Intuitively, max looks best. But in real work, thinking just enough can be the better choice over thinking the longest.
Reasoning effort does not change the model
OpenAI guides that lower reasoning effort saves speed and tokens, while higher effort spends budget on harder inference. GPT-6 Astra supports levels from low to max.
What matters is that the same model can change the following with effort:
- reasoning token usage
- time to first answer
- final success rate
- API cost
- answer length and review depth
So comparing only the name GPT-6 Sol is not enough:
GPT-6 Sol (high)
GPT-6 Sol (max)
From an operations view, treat these as different settings.
Where Max most often loses
Think of work with simple success conditions and a narrow answer space:
- Convert data to a JSON schema
- Fix a type error in a small function
- Summarize an email in a fixed format
- Fix a test failure with the cause already narrowed down
If high already succeeds 10 out of 10 times, the extra thinking in max cannot raise the success rate. What mostly grows is time and cost.
In contrast, higher effort can pay off for work that edits many files, interprets conflicting requirements, and verifies failure causes repeatedly.
CodeBridge Mini Lab: run High and Max on the same problem
When you compare, fix the success conditions before asking "which answer looks smarter?"
For example, run the following task 5 times each on a small codebase:
Task
1. Find the causes of 3 failing tests.
2. Edit production code only.
3. Pass all existing tests.
4. Change no unnecessary files.
A simple tracking table is enough:
run,effort,success,time_sec,cost_usd,files_changed
1,high,1,71,0.14,2
2,high,1,68,0.13,2
3,max,1,132,0.29,2
The question to ask is not did max do better? but this:
Did Max cut failures enough to justify the extra cost?
Difficulty-based escalation works well in practice
Instead of sending every request at max from the start, step it up gradually:
medium
↓ failure or verification miss
high
↓ still failing
max
This fits model routing too. Start easy work on a cheap model with low effort, and climb to pricier settings only when success conditions fail.
The key is letting test and verification results decide escalation — not the model's own "this looks hard."
Never trust a single run in quality comparisons
Agent and reasoning work can take different paths each run. Capturing one success story overstates setting differences.
Run the same problem several times at minimum, and record the following together:
| Item | Meaning |
|---|---|
| Success rate | Real completion probability |
| Median time | Felt waiting time |
| Cost per success | True cost per one success |
| Retries | Operations complexity |
| Regression | Whether existing features broke |
Conclusion: Max is closer to a last card than a default
Reasoning effort is closer to a thinking-budget slider than a quality slider.
Running even easy problems always at max grows cost fast while success rates barely move. For complex work, though, higher effort can cut retries and lower total cost instead.
So the best rule is this:
Settle on the lowest effort that meets success conditions reliably, and climb only on failure.
Further reading
- Why benchmark scores alone can mislead
- GPT-6 Sol vs Luna: when is the fast, cheap model the better pick?
- You should judge AI model cost per successful task
References
Go deeper with a course
If you want to build judgment for matching models, effort levels, and tools to each situation, a practical course on using AI tools by scenario helps.