LocalAI 4.11.0 landed October 2. A big one: 243 merged PRs, 353 commits, 15 days of work. Audio scenes, failover, and ops pages all ship, but this post watches one thing. Decision Models and the /v1/systemone API.

If yesterday's Strands Decider post was a single model launch, today's bigger current is decision models becoming an independent primitive of a real inference stack.

What arrived: decisions is a declared use case

Trimmed to essentials from the release notes.

Item Detail
Declaration Models state known_usecases: [decisions] explicitly (never inferred, kept apart from chat and classification)
API Structured choice, score, and noul requests through POST /v1/systemone
Validation 64 KiB body cap, max 64 questions, checks on state, IDs, options, levels, and noul criteria
UI Decisions capability chip plus install guidance
Gallery models Laya, GLiNER2.5-Decide, Qwen3-VL, Tev1, kev, Nimble, CLM, and more
The flow:
a model declares decisions → install from gallery → call /v1/systemone
→ choices, scores, and extraction with no text generation

Existing /permute and /separate (NER and token classification) stay put while decisions takes its own seat. Asking a chat model to decide was a workaround. Now it is a first-class stack layer.

What came along: failover and ops pages

Decisions did not ship alone. Two ops-side arrivals pair with it.

Failover chains:
- one model name fronts an ordered local-plus-remote target list
- retries only pre-commit failures (validation errors and client cancels excluded)
- health and recovery tracking plus admin pin and unpin (MCP tools and UI)
- served model reported in response headers
→ "next target on decision failure" becomes a stack feature

Operate → This machine:
- single-node gauges for VRAM, RAM, CPU, and model disk
- log inspection and model stops
→ local ops without distributed mode

The fallback from the routing post and the execution bundling from the observability post arrive as local stack features. Ops patterns once cloud-only move down to local.

Why it matters: take choosing off generative LLMs

Many teams currently run classification, routing, moderation, and tool selection on generative LLMs. Slow and pricey. LocalAI promoting decisions to first class means the replacement sits in a standard slot.

Swap order (cheapest and easiest first):
1. moderation and spam calls (yes/no plus confidence is enough)
2. routing and classification (fixed options, scores needed)
3. tool selection (pick among candidate tools)
4. zero-shot extraction (fixed-shape pulls)
→ none of the 4 strictly needs text generation

The Mini Lab from the Decider post transfers directly: agreement rate, latency, and cost comparison. Only the model changes to a LocalAI gallery decisions model and the call site becomes /v1/systemone.

CodeBridge Mini Lab: move 1 generative choice to decisions

1. Find 1 choosing job on a generative LLM (1 of the 4 above)
2. Log 2 weeks: call counts, tokens, latency, cost
3. Run the same job on a LocalAI decisions model:
   [ ] pick agreement (target 95%+)
   [ ] latency and cost gaps
   [ ] escalation path for low confidence
4. Call it: on agreement, split it off and add a failover chain
   (local first, remote second, retrying only pre-commit failures)

It links to the local economics from the DGX Spark post. Decisions are small and run well locally. "Always-on judging" moves home.

Conclusion: judging is a stack

One line to close.

Choosing became a stack feature, not a model.

Yesterday brought Decider the model, today brings decisions the layer. The order matters. Models change, layers stay. What plugs behind /v1/systemone can wait. One task for today: find 1 choosing job among your generative calls. That one is the first candidate moving to the local stack.

Further reading

References

Go deeper with a course

To practice designing choice, delegation, and verification as a structure, this course builds harness, loop, and graph layers exactly like the decisions split here.