A sentence like hallucinations went down feels very reassuring in model evaluations.
But you must always check one thing with it:
Did the model get less wrong because it knows more, or because it answers less?
The easiest way to avoid errors is to stay silent
Imagine an extreme model.
Model A
100 questions, 100 answers
60 right / 40 wrong
Model B
100 questions, only 20 answers
18 right / 2 wrong
Model B looks low on hallucination. But it left 80 questions unhandled.
So when you look at hallucination, check these together at minimum:
Accuracy
Wrong answers
Attempt / answer rate
Abstention rate
What the GPT-6 Sol case shows
In a September 2026 Artificial Analysis report, GPT-6 Sol at max sharply cut hallucination rate on AA-Omniscience versus the prior generation.
At the same time, the model held back answers more often instead of answering everything. Wrong answers fell, but simple accuracy also dipped in places.
That case asks a good question:
"Is always answering really the better product experience?"
Think about what each metric hides. Accuracy counts how often the model is right when it answers. Answer rate counts how often it even tries. A model can lift accuracy by refusing risky questions, and it can lift answer rate by guessing more. You only see the trade when you read both.
This also explains why leaderboard comparisons can confuse buyers. Two models with the same accuracy can feel very different in production. One answers everything with shaky confidence. The other answers less but tells you clearly when it stops. For support tools, search assistants, and RAG products, the second behavior is often easier to operate.
The sweet spot depends on the job
For light brainstorming, giving an answer may matter more than avoiding mistakes.
But in these areas, a cautious answer can be more valuable:
- Company policy lookup
- Contract extraction
- Medical and finance assistance
- Production operations commands
- Research that needs sources
Here an "I don't know" is not a failure. It can be a safe state change.
CodeBridge Mini Lab: put abstention in your success criteria
Prepare 20 questions.
- 10 questions answered in the docs
- 10 questions not answered in the docs
Compare two prompts.
A:
Answer whenever you can.
B:
If you cannot verify it in the given evidence,
answer "not verified in the evidence."
Then count four outcomes:
Correct
Wrong
Correct abstention
Unnecessary abstention
This table often shows real RAG and work-assistant quality better than simple accuracy.
Run the same 20 questions through both prompts and compare the shift. Prompt A usually raises correct answers but also raises wrong ones, especially on the 10 unanswerable questions. Prompt B should convert many of those wrong answers into correct abstentions. If Prompt B only adds unnecessary abstentions on answerable questions, your instruction is too strict or your retrieval is too weak.
You can also score confidence signals. Ask the model to mark each answer as direct quote, paraphrase with source, or uncertain. Then check whether uncertain labels line up with actual errors. A model that labels well is easier to route: direct quotes can ship, paraphrases get a quick scan, and uncertain answers trigger extra search or escalation.
You must also design what happens after I don't know
Abstention alone can frustrate users.
A good system has a next step:
Low confidence
→ extra search
→ check another source
→ escalate to a stronger model
→ if still unsure, tell the user clearly
So abstention is not the end. It can be a routing signal.
Conclusion: a trustworthy AI is not one that always answers
If your goal is only "fewer wrong answers," the model can go too quiet.
If you only push answer rate, confident wrong answers can grow.
So track operations on two axes:
How well it answers + how well it stops when unsure
In your next eval, report accuracy, wrong-answer rate, answer rate, and abstention quality side by side. Reward correct abstentions on unanswerable questions and penalize confident errors. When both axes stay visible, you stop chasing silent models or chatty ones, and you start shipping assistants people can trust.
Further reading
References
Go deeper with a course
If you want to design assistants that cite sources, admit gaps, and escalate cleanly, practice role-based AI workflows.