Three models invite a ranking. Microsoft 83.5%, Jev's figures, Laya 76.6%. Lining the numbers up feels like an answer. But those three numbers come from different tests. Put them in one line and you get an illusion, not a ranking.
This post separates two kinds of evidence: figures each side released, and experiments run on the same inputs.
Separate first: different tests versus the same test
Remember just this before comparing.
Different tests (read separately):
Microsoft's 36-benchmark average / Jev external evals and JevBench / Laya tracker
→ conditions differ, so no verdict on superiority
Same test (direct comparison allowed):
sysone-bench v2, the same-input experiment cited by the Laya tracker
→ Jev vs Laya only; Microsoft is not in that experiment
Skipping this separation causes the familiar mistake. Three numbers on screen, then a winner picked.
Figures each side released: read the conditions first
Placing each model's public figures with conditions gives this.
| Model | Public figure | Condition |
|---|---|---|
| Microsoft-Decision-1 | 83.5% mean accuracy | 36 benchmarks, about 147K questions, Microsoft's own eval with held-out questions |
| Jev | Various external evals and JevBench records | JevBench Score is not an accuracy percent; figures move with version such as jev-1.13.0 and methodology |
| Laya | 76.6% fine-tuned checkpoint, 36.1% base | 400 cases and 2,000 decisions, T4, specialized for the Typed Decisions family |
On Jev, JevBench v1.6.1 records Jev 1.13.0 Capability at 77.1 and Laya's English checkpoint at 32.5. But Capability averages Intelligence and Calibration, and Composite adds Speed and Cost as a secondary score. Do not read them as accuracy percent. Versions move the numbers and ranks.
Laya hides one more trap. The often-quoted 76.6% belongs to a checkpoint fine-tuned for Typed Decisions. In the same report, base Laya scores 36.1%. A specialized tuning figure should not describe general performance. The Laya tracker states this up front. Its 32.8 ms latency is also an in-process T4 GPU measurement, not comparable to API latency with networking.
Microsoft's side was covered in the numbers post: vendor-measured, with the 36 names not fully public and no independent replication yet.
Summary:
- Microsoft 83.5% = Microsoft's test result
- Jev figures = external and community records (check definitions and versions)
- Laya 76.6% = a specialized tuning figure (not the base model)
→ never put all three in one table
The same test: sysone-bench v2 ran Jev and Laya head to head
Jev and Laya do have a public same-input experiment. Sysone-bench v2 ran 751 byte-identical states with the same seed. It is also cited by the Laya tracker. Jev was pinned to jev-1.13.0, Laya's English checkpoint ran on a local M2 CPU, and the whole Jev run reportedly cost $0.008.
Four tasks alone tell the story.
| Task | Sample | Jev | Laya | Point |
|---|---|---|---|---|
| Support triage | n=160 | 88.8% | 80.0% | Jev ahead |
| Guardrails | n=60 | 96.7% | 88.3% | Jev ahead, small sample so read carefully |
| AG News | n=100 | 91.0% | 94.0% | Laya ahead |
| MNLI | n=60 | 86.7% | 98.3% | Laya ahead |
Across all nine suites the pattern sharpens. Jev pulls away on fine-grained multi-class work such as 6-way intent, toxicity, and multilingual intent, while Laya takes AG News and MNLI. Emotion sits around 0.54 for both, and 5-level scoring ties Jev with an open baseline.
The point to take is the pattern, not a winner. No single model wins every type. Comparison-shaped problems such as MNLI premise-hypothesis pairs suit an encoder like Laya, while many-label or non-English work favors Jev.
Caveats:
- Small samples (60 to 160). The guardrails gap may sit inside noise
- Includes hand-labeled cases by one author, English-centered
- Pinned version (jev-1.13.0). Re-run when latest moves
- Microsoft is not in this experiment. No three-way winner exists
The Laya Router note matters too. English-only checkpoint scores 36.0% on multilingual intent, while routed Laya reaches 84.0%. Jev still scores 100% under the same 25-case condition. One routing design decision more than doubles the result, which is the hint this experiment offers.
Latency makes the same mistake
Accuracy is not the only comparison people mix up.
- Microsoft p50 85 ms is a Foundry hosted measurement
- Laya 32.8 ms is an in-process T4 GPU measurement
- Sysone-bench Laya 180 to 660 ms is local M2 CPU for 5-question calls; Jev 925 to 1,068 ms is end-to-end API in the same run, region included
- Jev's official guidance says 70 to 500 ms, while a Cloudflare measurement records a 524 ms median
Different hardware, question counts, and networks cannot share one line. In particular, never mix in-process GPU time with end-to-end API time. For your service, measure end-to-end latency with networking and orchestration included.
Cost follows the same rule. Jev and Decision-1 share the $0.042 per million input tokens unit with free output. Laya costs nothing at the margin once downloaded. The more telling line is that the whole 751-state Jev run cost $0.008. Small decision models should be judged by "how cheap when bundled," not by unit price alone.
CodeBridge Mini Lab: build your own table
Do not memorize someone else's ranking. Build one small table of your own.
1. Collect 100: real inputs + labels (keep the ambiguous ones)
2. Pin everything: record model versions, byte-identical inputs, same seed and questions
3. Measure together: accuracy + calibration (ECE) + p50 and p95 + per-correct cost
4. Split by task: classification, guardrails, and inference separately
5. Check gating: at 0.85 confidence, what share auto-processes and at what accuracy
Carry over sysone-bench's lesson directly. Match question hashes first, pin versions, and write down sample sizes and limits. That is the whole of fair comparison.
Conclusion: ask for conditions, not ranks
One line to close.
Never compare top scores from different tests. Read task-by-task results from the same test.
Microsoft 83.5%, Jev's records, and Laya 76.6% each belong in their own place. On identical inputs, Jev leads some tasks and Laya leads others. Microsoft is not in that run, so no three-way winner exists.
Next comes practice. The patterns post covers where attaching pays off, and the checklist post covers what to measure before adoption.
Further reading
- An AI that only decides: Microsoft-Decision-1 overview
- Why 83.5%, 35x, and $0.042 need conditions
- What is Jev? AI that decides without writing
- Scores alone can mislead
References
- Microsoft Command Line: Introducing Microsoft-Decision-1 (Oct 9, 2026)
- Laya Benchmarks Tracker (as of Sep 23, 2026)
- Laya BENCHMARKS.md (source)
- sysone-bench: Laya vs Jev same-input comparison (GitHub)
- SDAD: byte-identical Laya vs Jev write-up (Sep 23, 2026)
- Zylver: Laya vs Jev open benchmark (Sep 22, 2026)
- JevBench: Jev vs Laya (v1.6.1, Oct 6, 2026)
- JevBench repository
Go deeper with a course
Fair-comparison instincts do not end at reading tables. They stick after running the same experiment on 100 of your own inputs and deciding where to stop when confidence is low.
The Practical Decision AI course runs Laya directly, from state design through Choice, Score, and Noul interpretation to confidence, policy, and Human Review in one flow. If you want task-by-task tables for your own work instead of someone else's ranking, it continues directly from this post's Mini Lab.