Hearing 83.5% accuracy feels reassuring. But that number is Microsoft's test result, not a guarantee for your service. More important is the rule that even at 99% you must not send money or delete files automatically.
This post covers four things to measure before adoption and work that must never auto-execute. It closes the branch opened in the overview and the patterns post.
Four things to measure before adoption
Do not look at plain accuracy alone. Measure four things together per model on the same labeled data under the same request conditions.
| Measure | Question | Why it matters |
|---|---|---|
| Accuracy plus failure cost | What do we lose when wrong | The cost of being wrong beats the rate of being right |
| Calibration | Is 90% really right 90% of the time | Required before gating automation on thresholds |
| P95 latency | How bad is the slow tail | Mean and median alone break in production |
| Per-correct cost | What does one correct decision cost | Retries and review included, not just token price |
Each in turn.
Read accuracy with "what was missed." Even at 90% overall, systematic misses on payment cases block adoption. Split by task. As the fair-comparison post shows, no model wins every type.
Calibration asks whether confidence can be trusted. Predictions at 0.9 should be right about 90% of the time before thresholds can automate anything. Treat Microsoft's 92.2 and Quyet's 93.1 as references and check expected calibration error (ECE) on your data. Sysone-bench also reports score-style questions calibrating worse than choice or yes-or-no.
Read latency as p95 and distribution, not p50 alone. Numbers like Decision-1's 85 ms p50 and 125 ms p95 change under your traffic. Measure end-to-end latency with networking and orchestration, and watch how per-step delay stacks across 20 steps. As the speed post shows, experience breaks at the tail.
Read cost from the per-success-cost post. Compare input tokens plus retries, upper-model checks, and human review. $0.042 per million tokens is only the starting point.
Experiment frame (start with 100):
100 identical inputs + labels
→ current method vs Decision AI, same conditions
→ per-task accuracy + calibration + p50 and p95 + per-correct cost
→ set the threshold from the confidence distribution (raise it when low)
Work that must never auto-execute, even at 99%
Even with high accuracy, these four need separate approval.
Never auto-execute:
- sending money (payments, refunds, transfers)
- deleting (files, data removal)
- changing (permission and settings changes)
- exporting (personal data transfers)
A HIGH CONFIDENCE label is not legal approval or policy verification. A 99% refund-request probability does not mean refund eligibility. Code must check payment history, policy, and approval rules. That is why the worked gate in the patterns post sends even a 0.99 payment case to a human.
The Microsoft catalog draws the same boundary. Do not use Decision-1 as the sole automated decider for consequential decisions about credit, employment, housing, insurance, education, healthcare, or legal rights, and do not use it for surveillance, profiling, or tracking individuals. A probability score must not stand alone as a judgment about a person. Decision AI belongs inside a safe workflow.
Principle:
model = reads meaning and gives probabilities
code = enforces policy, eligibility, and risk conditions
human = final approver for risky work
The split matters because responsibility turns sharp. It is the same structure from the Jev post.
Closing the architecture: Generate, Decide, Verify, Act
Drawing the briefing's conclusion as a structure gives this.
Generate (LLM): long prose, analysis, code drafts
→ Decide (Decision AI): picks, scores, yes-or-no
→ Verify (code): policy, eligibility, thresholds, risk flags
→ Act: execute, retry, or human review
A good AI system is not one model doing everything, but work divided by design. It matches the execution-versus-verification split in the harness post.
Do not divide everything on day one. Start by finding one repeated decision in your current agent and comparing it on 100 cases. Confirm the split when match rates clear 95% with much lower latency, then set thresholds from the confidence distribution. That is the one thing to try today.
CodeBridge Mini Lab: this week's validation sheet
1. Pick 1 decision (1 of routing, retry, verification, review)
2. Collect 100 (real examples + labels + ambiguous cases included)
3. Measure 4 things:
[ ] per-task accuracy and failure cost
[ ] calibration (actual hit rate in the 90% band)
[ ] end-to-end p50 and p95 latency
[ ] per-correct cost (retries, upper checks, review included)
4. Set the gate:
[ ] threshold (for example, below 0.85 goes up)
[ ] risk conditions (payment, delete, permission, personal data always reviewed)
[ ] pinned versions and logs (model version, inputs, probabilities recorded)
With this sheet, the question changes from "should we adopt it" to "how far do we automate." That is the real conclusion of Decision AI adoption.
Conclusion: divide, measure, and let humans close
One line to close.
Complex generation goes to LLMs, repeated picks go to Decision AI, and execution decisions go to code and humans.
Do not ask one model to do everything. Move the picking work out, measure it on your data, and let humans close risky work. Those three steps are the whole point of this briefing.
Find one step in your service that only picks instead of answering. Share which decision you want to automate in the comments, and we will dig further in the next posts.
Further reading
- An AI that only decides: Microsoft-Decision-1 overview
- Three practical ways to attach Decision AI to an agent
- Do not rank Jev, Laya, and Microsoft in one line
- Judge AI cost per successful task, not per token
- From tokens to outcomes for AI agent KPIs
References
- Microsoft Command Line: Introducing Microsoft-Decision-1 (Oct 9, 2026)
- Microsoft Foundry Blog: Introducing Microsoft-Decision-1 in Microsoft Foundry (Oct 9, 2026)
- Laya Benchmarks Tracker
- sysone-bench: same-input comparison
Go deeper with a course
Once the validation sheet is clear, only attaching it to your work remains. The hardest practical part is gathering 100 cases, defining correct answers, and setting the confidence level for human handoff.
The Practical Decision AI course runs Laya directly to build states and questions, then connects confidence to policy and Human Review, plus a validation frame to use before applying it to your own work. If you want to think in per-correct-decision cost, it continues directly from this post's Mini Lab.