On October 9, Microsoft released a Decision AI model. It is called Microsoft-Decision-1. It does not write like GPT. It answers with probabilities over fixed options.

If you only saw the headline, "35x faster" stands out first. Before that number, this post covers what the model actually is. The number itself gets its own post here.

In one line: it picks, it does not write

Borrowing the official phrasing:

Given a fixed set of answer options, it returns a calibrated probability for each option.

Think of it as closer to a decision function. State goes in, JSON numbers come out instead of paragraphs.

General LLM:    situation → long text and reasoning (many tokens, slow)
Decision-1:     situation + options → pick + probabilities (single pass, fast)

This should feel familiar. It sits where Jev sits as a System One model, and where the open model in the Strands Decider post sits. The difference is the managed path: Microsoft Foundry.

Flow of one decision. State input, Decision-1 scoring, probability JSON output, then code and policy checks

The diagram above is the whole structure. The model computes probabilities, and application code plus policy decides what to do. That separation is the most important point in this post.

What it is built on: Qwen3.5-9B post-trained for decisions

According to Microsoft, Decision-1 starts from Alibaba's open-weight Qwen3.5-9B and is post-trained for decision scoring. Microsoft says future versions may rebase onto its own MAI or OpenAI-family models. The interface stays while the base underneath changes.

Item Detail
Base Qwen3.5-9B post-train, parameter count undisclosed
Input Text only, up to 32,768 tokens
Output Per-option probability JSON, no generated prose, no output tokens
Access Microsoft Foundry catalog public preview, OpenRouter listing
Distribution Hosted API only, no weight downloads or self-hosting

Hosted-only resembles Jev, and differs from Laya, which you can download and run locally. Azure authentication, billing, and governance come attached, with less freedom over placement.

The Foundry catalog's Benchmarks tab reportedly carries only a methodology note with no externally checkable figures yet. So every performance number in this post should be read with the qualifier "as measured by Microsoft."

Three question shapes: Choice, Yes/No, Score, plus rubrics

Calls are simpler than they look. Send one situation (state) plus the questions to ask about it. There are four shapes.

Type What it does Example
Choice Pick one of fixed options Owning team: billing, technical, account
Yes/No Probability some statement is true Is this a refund request (0 to 1)
Score Score on an ordered scale Urgency low, medium, high
Rubric Grade an AI answer or agent action against criteria Accept, revise, or reject

Sending several questions at once evaluates them in parallel, which is handy in practice. "Which team owns it, is it a refund request, is auto-routing OK" can go in one request. The classification step in the routing post fits exactly in this Decider slot.

The Foundry Playground example follows the same pattern. One customer request goes in, and urgency check, team pick, and priority score come out together. The official TechCommunity post shows the real screen, so look at that screen and demo video before implementing. This post uses the concept diagram above instead of copying the official screen for copyright reasons.

A concrete example: one support message

Take "I was charged twice. Please refund the extra charge." The system's first job is not a pretty reply. It is a decision.

state: "I was charged twice. Please refund the extra charge."
questions:
  - team: [billing, technical, account]
  - refund_requested: yes/no probability
  - auto_route_ok: yes/no probability
answers (illustrative, not real output):
  - team: billing 0.91
  - refund_requested: 0.99
  - auto_route_ok: 0.82 + confidence

The numbers above are illustrative. Real values vary with input and model version. What matters is the role split. The model reads meaning, code checks payment history, refund policy, and approval rules, and an LLM writes the final human-facing message last. Responsibility divides.

The official Python example has the same shape. It posts state and questions to the providers/microsoft/v1/systemone route and reads the choice back from answers to measure accuracy and latency. Deployment region and auth vary by account, so confirm the Foundry quickstart values before running. Keep credentials outside the source.

Where it sits inside an agent

Microsoft's listed uses make the slot clear.

Common slots:
- Agent controls: continue, stop, retry, or hand to a model, tool, or human
- Model routing: pick a cheap or strong model from the request
- Intent analysis, classification, labeling: assign messages and feedback to fixed criteria
- AI judging: grade an LLM answer against a rubric and accept, revise, or reject
- Search relevance, safety screening, computer and UI use, robot action selection

Xbox Research used it to group 10,000-plus open feedback items into researcher-defined themes, the Copilot team used it to measure reply quality, and the Discovery team used it for rubric grading in experiment replanning. All three are "pick by criteria" work, not "write long new text" work.

So the conclusion is not "it replaces LLMs" but "it divides the work." Complex generation and reasoning stay with LLMs, repeated structured decisions move to Decision AI. Code verifies behind them, and humans approve risky actions. That division continues with examples in the production-patterns post.

One thing to take away

Restating Decision-1 in one line:

Picking is becoming a layer of the stack, not a feature of a model.

Yesterday an open model called Decider arrived, and in the LocalAI post decisions became a first-class stack API. Today Microsoft adds one more managed option to that flow. Models change; the layer remains.

Today's task is one thing. Find one step in your service that picks rather than writes. Collect just 100 inputs for it, and the number-reading and comparison experiments in the next posts can start right away.

Further reading

References

Go deeper with a course

The feel for Choice and Score and reading confidence only sticks after running them yourself. The hardest practical part is deciding what stays as code rules, what goes to model judgment, and where human review enters.

The Practical Decision AI course runs Laya locally to build states and questions, then connects confidence to policy and Human Review, plus a validation frame to use before applying it to your own work. If you want to move one repeated decision off your plate, it continues directly from this post.