Coding-agent choices used to look binary. Expensive, slow, but strong frontier models, or fast, cheap small models that choke on repo work.

JetBrains' Mellum 2.1, released October 8, shows a third path. Same size, but the training weight moved to reinforcement learning, growing the ability to explore a repo, edit it, and check the changes. Apache 2.0 licensed, so you can host it yourself.

The spec: 12B total, 2.5B awake at a time

Architecture is unchanged from Mellum 2. Everything after pre-training was reworked.

Design: Mixture-of-Experts (MoE)
Total parameters: 12B
Active per token: 2.5B (8 of 64 experts fire)
Layers: 28 / Context: 131,072 tokens
License: Apache 2.0 (weights on Hugging Face)

Think of MoE as 64 specialists on standby with 8 woken per token. Per-token compute stays near a 2.5B dense model while knowledge breadth comes from 12B. In JetBrains' words, a design for running on realistic hardware like a single H100.

What changed in this release is all post-training.

- RL promoted from a short final stage to the main part of training
- New RL tasks: math, competitive programming, science, tool use, SWE
- Open RL data filtered first (broken tests, unverifiable answers removed)
- Thousands of in-house RL environments, millions of sandbox runs

The result is new behavior. The model explores a codebase, edits files, and checks its own changes. It can take one piece of an agent's plan: find a failing test's root cause, draft a fix, verify it.

How to read the benchmarks

The announcement compares Mellum 2.1 against Mellum 2 plus similar-class open models Qwen3.5-9B and Gemma 4 E4B under one shared setup. Charts cover LiveCodeBench, AIME, GPQA, BFCL, IFEval, and SWE-verified cards.

The short version:

Biggest gain: agentic coding (largest jump vs predecessor)
Breadth: gains across coding, comp programming, math, tool calling, knowledge
Claim: holds on hard problems as well as everyday ones

Pause here. These are vendor-measured numbers. One shared setup for four models looks fair, but the tasks and settings are still the vendor's choice. Some third-party summaries cite figures like 47% on SWE-Bench, but hold that citation until you have checked the official text against the Hugging Face model card yourself.

Read the speed numbers the same way. The official wording:

Single request: ~1.6x faster with MTP (multi-token prediction) on the same H200
Heavy load: fastest in group, ~2x the token throughput of Qwen3.5-9B
Premise: post-training untouched architecture, so same speed as Mellum 2

Faster looks good. But your number must come from your repo, your hardware, your concurrency. Official charts are a starting line, not a verdict.

Where it fits: fast hands next to a big brain

Mellum 2.1 is not an "does everything" model. It is a worker that takes pieces inside agentic systems.

1. Worker inside agent systems:
   root-cause a failing test → draft a fix → self-verify
   A big model plans; Mellum clears the middle tasks

2. General assistant beyond coding:
   everyday questions, math and reasoning step by step

3. Private self-hosted deployment:
   run locally or on your own infra, keep code and data controlled

This plugs straight into the Qwen Code delegation post. When agents hand work to other agents, someone cheap and fast must clear the handed-off tasks. Same lesson as the harness-beats-model post. Wins come from verification loops, not one model shot.

Trying it locally uses familiar tools. Weights from the Hugging Face collection, served vLLM-style as an OpenAI-compatible API. GGUF builds, Ollama and LM Studio support, and the MTP head for vLLM are coming soon per the announcement.

# Without tool calling (per official instructions)
vllm serve JetBrains/Mellum2-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3
# With tool calling (per official instructions)
vllm serve JetBrains/Mellum2-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

CodeBridge Mini Lab: measure three things separately

The fastest way to turn vendor numbers into your numbers:

1. Repo exploration: 5 issues, does it find the right top-k files?
2. Bug fixing: 5 failing tests, patch success rate plus regressions?
3. Test pass rate: time to a green full suite after the fix

Record together: runtime, GPU memory, concurrency conditions
Compare against: one large model you use today, same tasks

Do not watch "fix rate" alone. Exploration accuracy, fix success, and no-regression often move separately, and binding them is the harness's job.

Conclusion: divide roles before upsizing models

One line to summarize.

Before handing everything to a giant model, design a structure that splits work to fast workers.

Mellum 2.1 is a strong candidate for that worker. No need to light up all 12B, explore-edit-verify inside repos, Apache 2.0 for your own infra. But a 1.6x speed figure is only a starting line. Measure exploration, fix, and test-pass rates on your own repo. That is the only way to make this announcement yours.

Further reading

References

Go deeper with a course

If you want to design role-splitting between big and small models and agent loops with repeated execution and verification, the step-by-step harness, loop, and graph path connects directly to this post's worker-model story.