The open-weight field has quietly shifted.

On Artificial Analysis, the open-weight leaderboard reads MiMo-V2.6-Pro at 46, GLM-5.3 at 45, and Kimi K3 at 44. One point apart, so the three look similar on scores alone.

But when you actually deploy one, the score is the last thing to check. First comes the license you can use, whether serving fits your GPUs and budget, and how many tokens it burns. Those three come first.

The scoreboard: what a one-point gap means

Model Intelligence Index Note
MiMo-V2.6-Pro 46 Current open-weight leader
GLM-5.3 (max) 45 Strong on business workflows
Kimi K3 (max) 44 Open frontier, giant MoE

GLM-5.3 scores around 62% on AutomationBench-style tasks, and Kimi K3 made headlines as a 2.8-trillion-parameter MoE. For background, see the Kimi K3 post and the open-weight cost post.

One confusion is easy here. At 46 points, an open model sits about 12 points below the top closed model at 58, so it can feel "still far behind." Change the use case and the story changes. If you do not need peak intelligence every time, the lower serving price and data control of open models can cover the score gap.

Change the question:
"Which is smartest?" → "Does my workload run well at this score level?"

For Korea-based teams: read the Upstage Solar Mini 4 news too

On September 30, Korea's Upstage released Solar Mini 4. It is a small model at Intelligence 24, and the Artificial Analysis note is telling: token prices look similar to GPT-6 Luna, yet cost per task runs about 5x higher.

That captures the core of small and open model selection. Even with similar per-token prices, a model that burns more tokens to finish one task bills you more. With small open models, skip the price table and measure success rate and retry counts on your own tasks.

50% cheaper per token × 3x retries = actually more expensive

→ Always compute "cost per success" (method in the Mini Lab below)

Three checks before you choose

1. License: can you use it commercially?

Open weights do not share one license. Non-commercial restrictions, conditional commercial grants, and full commercial permission are all mixed together. Artificial Analysis flags commercial restrictions separately in its Openness Index.

For company projects, read the license section of the model card first. Score comparison comes after.

2. Serving: where will it run?

There are roughly three options.

A. Use a hosted API
   - Upside: no GPU worries, start immediately
   - Check: cost per task, cache discounts, speed

B. Self-host on cloud GPUs
   - Upside: data control, better unit cost at high volume
   - Check: VRAM needs, p95 latency under concurrency

C. Run on a laptop or workstation
   - Upside: fully local, free experimentation
   - Check: quality loss after quantization, speed

If option C interests you, AA-AgentPerf-Local from September 29 is worth a look. It is an open-source tool that measures agent execution speed of four open models on DGX Spark, Ryzen AI Halo, MacBook Pro (M5 Pro), and RTX 5090. Running agents on a laptop has become something you measure, not just a hobby.

Measure serving with p95 and concurrency, not averages. The method is covered in the NVIDIA AIPerf post.

3. Token efficiency: with MoE, look at active parameters

With MoE models like Kimi K3, total parameters and the active parameters used per token differ. A big total does not mean heavy compute per token. But a small active count can still mean harder serving because of routing and memory structure. See the MoE active parameters post for the concept.

Put these in your comparison table:
- Intelligence score (state whether it is max, and which effort)
- Output tokens per task
- Cost per task (counting only successes)
- Context length and quality on genuinely long inputs
- License terms

CodeBridge Mini Lab: compare two candidates on 10 tasks

Do not pick from the scoreboard. Pull 10 tasks from your own work and run them.

Materials: 10 frequent tasks (summaries, classification, code edits, table analysis)

Record per model:
- Success: Y / N (define the bar first, e.g. tests pass + cited evidence)
- Retry count
- Total tokens and cost
- Failure type: ignored instruction / hallucination / tool-call failure / format error

Decision:
Below 80% success rate, it is out (cheap does not help)
On tied success rates, pick the lower cost per success

Use the calculation from the cost-per-successful-task post directly. It is total cost until success, not per-token price.

Conclusion: picking open weights is not a scoring game

The order of work looks like this.

Check the license → measure success rate on 10 of your tasks → compare cost per success → decide how to serve it

A gap of 46 vs 45 vs 44 points feels like noise in most real work. What stops a project is one license clause, one serving failure, or one retry explosion. Scores shortlist candidates; your own task measurements decide. That one line is the whole of open-model selection.

Further reading

References

Go deeper with a course

Picking a cheap open model does not finish an agent. If you want guided practice in splitting work and verifying it with harness, loop, and graph structure, this course continues the comparison experiment directly.