"A judging AI is 35x faster." It makes a perfect thumbnail. But without knowing what that 35x measured and compared, adoption decisions go wrong. This post unpacks the three numbers Microsoft released on October 9: 83.5% accuracy, 35x speed, and the $0.042 price.
If you have not seen what the model is, start with the overview. This post focuses only on reading the numbers.
83.5%: an average over 36 benchmarks and about 147,000 questions
Unpacking the accuracy sentence gives this.
- An average over 36 benchmarks and about 147,137 questions
- Includes held-out questions, spanning routing, ranking, long context, multilingual, out-of-distribution, reasoning, and safety
- Compared against Quyet-1.0-Large, H2O-Lightning-4B, GPT-6 Luna Decisions, and others (Jev added later to accuracy and calibration)
- Result: Decision-1 83.5%, Quyet 81.9%, GPT-6 Luna Decisions 79.4%, H2O-Lightning-4B 77.2%
It looks clean in a table. This table copies the official figures directly.
| Model | Base | Accuracy (Microsoft-measured) | Calibration |
|---|---|---|---|
| Microsoft-Decision-1 | Qwen3.5-9B post-train | 83.5% | 92.2 |
| Quyet-1.0-Large | Gemma-4-31B-it LoRA merge | 81.9% | 93.1 |
| GPT-6 Luna Decisions | Undisclosed | 79.4% | 89.9 |
| H2O-Lightning-4B | Qwen3.5-4B | 77.2% | 91.8 |
Read calibration alongside accuracy. A 90% prediction should be right about nine times in ten on representative cases. Decision-1 scores 92.2 for second place, Quyet leads at 93.1. If you plan to gate automation on confidence, this number matters more than accuracy.
And here is the key qualifier. This is Microsoft's own evaluation on its own selection. The 36 names are not all public, and no independent replication exists yet. The Foundry catalog is also reported to lack public figures. So read this table as "Microsoft's result," not an industry ranking. It does not guarantee accuracy on your service.
35x: a median over judging latency, not overall performance
Now the 35x. Read the sentence again.
P50 latency is about 35x faster than GPT-6 Sol.
P50 is not the mean. It means half the requests finished within that time. Microsoft reports Decision-1 at 85 ms p50 and 125 ms p95. In the same announcement, Quyet measured about 380 ms and GPT-6 Sol about 3,010 ms.
85ms x 35 ≈ 2,975ms ≒ the GPT-6 Sol measurement
→ "35x" compares this specific judging slice
The wording also shifted. Early coverage named H2O-Lightning-4B v1.1 as runner-up at 2.5x, while the current text cites 4.5x over Quyet. That is why runner-up figures should not lead your summary. They can change.
The common misunderstanding needs correcting too. It does not mean long writing or coding runs 35x faster. Decision-1 never generates prose. What was compared is one structured decision's latency. Mixing different units leads to wrong conclusions.
Correct reading:
- One structured decision, p50 basis, in Microsoft's measured setup: 35x
Wrong readings:
- All coding, writing, and reasoning run 35x faster
- Mean response time drops to one thirty-fifth
- 35x guaranteed in every environment
Microsoft's own illustration lands well. Saving 100 ms on each of 20 sequential decisions saves two seconds overall. Small delays compound when decisions repeat. Put the other way, work with only one or two decisions may barely feel the 35x.
Robustness came with the same release. Perturbing the same request eight ways flipped the decision 1.3% on average, and zero times when option text was paraphrased or order was shuffled or reversed. That signals stability, but it does not prove the selected option is correct for a customer's policy.
$0.042: per million input tokens, output free
Pricing is simple. Input costs $0.042 per million tokens, output tokens are free. It matches Jev to the cent, which prompted "copied pricing" commentary.
| Item | Decision-1 | Note |
|---|---|---|
| Input | $0.042 / 1M tokens | Same unit price as Jev |
| Output | Free | Natural, since nothing is generated |
| Other | Network and operating costs separate | Actual bills vary with input length |
Microsoft's illustration compares classifying one million texts at about $11 with Decision-1 versus about $2,434 with GPT-6 Sol. That is a modeled comparison, not a customer bill. It assumes matched input volume and includes output charges for the general model.
What matters in practice is cost per completed decision, not the token unit price.
Compare this:
price per million tokens X
cost per one correct decision O
= input tokens + retries + upper-model checks + human review
The Xbox case shows that view. Sorting 10,000-plus feedback items reached competitive quality with GPT-6 Sol at over 14x the speed and one two-hundredth the cost (the Foundry post phrases it as 80 to 100x, apparently a different summary of the same internal work). The Copilot team reports competitive quality with GPT-5.6 Luna at 100x the speed, and Discovery reports 46x more consistent rubric grading at 3x the speed and nearly 4x faster replanning. All are internal Microsoft measurements, so do not generalize them. But they hint where savings came from: moving repeated picking work off the big model.
A four-box checklist for every new number
With all three numbers covered, here is the takeaway. Check these four boxes for every model launch.
| Check | Question for Decision-1 |
|---|---|
| What was measured | A 36-benchmark average, and does it resemble my work |
| Whose setup | Vendor-measured, and is there independent replication |
| What the distribution looks like | Beyond mean and median, what about p95 and the slow tail |
| What one correct answer costs | Not token price, but a correct decision including review |
This matches the five checks in the benchmark-reading post: version, effort, harness, per-task cost, and detailed makeup. For Decision AI, never skip calibration. As the per-success-cost post shows, cheap tokens turn expensive when called often.
Conclusion: take the conditions, not just the numbers
One line to close.
83.5% is Microsoft's test result, 35x compares judging-latency medians, and $0.042 is an input unit price.
None of the three numbers lies. Each has a scope. Step outside that scope and the numbers stay right while the decision goes wrong.
The next step is simple. Collect about 100 real examples from your work. Run the current method and Decision-1 on the same inputs under the same criteria, and measure accuracy, calibration, p95, and per-correct-decision cost together. That experiment design continues in the fair-comparison post and the adoption checklist.
Further reading
- An AI that only decides: Microsoft-Decision-1 overview
- Scores alone can mislead: reading 2026 model comparisons
- Judge AI cost per successful task, not per token
- Do not rank Jev, Laya, and Microsoft in one line
References
- Microsoft Command Line: Introducing Microsoft-Decision-1 (Oct 9, 2026)
- Microsoft Foundry Blog: Introducing Microsoft-Decision-1 in Microsoft Foundry (Oct 9, 2026)
- IAMag: Microsoft-Decision-1 9B Agent Decision Model (Oct 10, 2026)
- Brocker: Microsoft-Decision-1 ships with top accuracy and 35x faster latency (Oct 9, 2026)
- OrcaRouter: Microsoft-Decision-1 Is GA on Foundry, With No Benchmarks (Oct 10, 2026)
- System One Models: Microsoft-Decision-1 specs and limits
Go deeper with a course
Once you can read the numbers, only your own measurement remains. The hardest practical part is gathering 100 cases, defining correct answers, and setting the confidence level for human handoff.
The Practical Decision AI course runs Laya directly to build states and questions, then connects confidence to policy and Human Review, plus a validation frame to use before applying it to your own work. If you want to think in per-correct-decision cost instead of token price, it continues directly from this post's checklist.