Model selection screens now carry one more unfamiliar option next to the model name: reasoning effort, with levels like low, medium, high, xhigh, and max.

Intuitively, max looks best. But in real work, thinking just enough can be the better choice over thinking the longest.

Reasoning effort does not change the model

OpenAI guides that lower reasoning effort saves speed and tokens, while higher effort spends budget on harder inference. GPT-6 Astra supports levels from low to max.

What matters is that the same model can change the following with effort:

  • reasoning token usage
  • time to first answer
  • final success rate
  • API cost
  • answer length and review depth

So comparing only the name GPT-6 Sol is not enough:

GPT-6 Sol (high)
GPT-6 Sol (max)

From an operations view, treat these as different settings.

Where Max most often loses

Think of work with simple success conditions and a narrow answer space:

- Convert data to a JSON schema
- Fix a type error in a small function
- Summarize an email in a fixed format
- Fix a test failure with the cause already narrowed down

If high already succeeds 10 out of 10 times, the extra thinking in max cannot raise the success rate. What mostly grows is time and cost.

In contrast, higher effort can pay off for work that edits many files, interprets conflicting requirements, and verifies failure causes repeatedly.

CodeBridge Mini Lab: run High and Max on the same problem

When you compare, fix the success conditions before asking "which answer looks smarter?"

For example, run the following task 5 times each on a small codebase:

Task
1. Find the causes of 3 failing tests.
2. Edit production code only.
3. Pass all existing tests.
4. Change no unnecessary files.

A simple tracking table is enough:

run,effort,success,time_sec,cost_usd,files_changed
1,high,1,71,0.14,2
2,high,1,68,0.13,2
3,max,1,132,0.29,2

The question to ask is not did max do better? but this:

Did Max cut failures enough to justify the extra cost?

Difficulty-based escalation works well in practice

Instead of sending every request at max from the start, step it up gradually:

medium
  ↓ failure or verification miss
high
  ↓ still failing
max

This fits model routing too. Start easy work on a cheap model with low effort, and climb to pricier settings only when success conditions fail.

The key is letting test and verification results decide escalation — not the model's own "this looks hard."

Never trust a single run in quality comparisons

Agent and reasoning work can take different paths each run. Capturing one success story overstates setting differences.

Run the same problem several times at minimum, and record the following together:

Item Meaning
Success rate Real completion probability
Median time Felt waiting time
Cost per success True cost per one success
Retries Operations complexity
Regression Whether existing features broke

Conclusion: Max is closer to a last card than a default

Reasoning effort is closer to a thinking-budget slider than a quality slider.

Running even easy problems always at max grows cost fast while success rates barely move. For complex work, though, higher effort can cut retries and lower total cost instead.

So the best rule is this:

Settle on the lowest effort that meets success conditions reliably, and climb only on failure.

Further reading

References

Go deeper with a course

If you want to build judgment for matching models, effort levels, and tools to each situation, a practical course on using AI tools by scenario helps.