On October 3, German Unity Day, Aleph Alpha released Kolibri. A German and English open-weight MoE under Apache 2.0 on Hugging Face. The numbers catch eyes. 78.1B total parameters with 3.46B active per token. About 4.4% awake.

Weekend evals just started, so this post covers structure and picking method over scores.

Specs: big and light

Item Detail
Shape MoE, 50 layers, 6 of 384 experts active
Parameters 78.1B total, 3.46B active per token (FP8)
Context Trained to 262,144, configured up to about 1M
Languages German and English (21.3% German in training)
Training infra Germany and Finland
Features Per-request reasoning effort, tool calling
License Apache 2.0

One fun record: on October 4 an NVIDIA forum post shared Kolibri running on DGX Spark with setup and numbers. The "100B on a desk" story from the DGX Spark post turned real on launch week.

Reading scores: workloads, not averages

The brief's viewing point is exact. Watch workload gaps over vendor scores.

- Aleph Alpha's own harness: Honeypot agentic-RAG 80.8
- Overall English benchmarks: Qwen3.8 27B scores higher
→ closer to an enterprise-workload pick than "strongest open model"

The open-weights race post covered the 46, 45, 44 one-point game. Kolibri plays another axis. Not overall rank but German, sovereign, on-prem domains. The active-parameter idea from the MoE post turns practical here. 78B talks storage and memory. 3.46B talks money and speed.

4-axis eval: your workload, VRAM, speed, boundary

The brief's action unfolds into a scorecard.

Open-model 4 axes:
[ ] your workload: 10 of your tasks, not 1 average bench (German tasks for German)
[ ] VRAM: does the full 78B load (check recommended specs around 2x A100 80GB class)
[ ] tokens/sec: measured on your hardware (time it AgentPerf style)
[ ] deployment boundary: how far data travels (sovereign and on-prem needs)

p95 and concurrency from the serving-measurement post plus the 20-case triage from the RAG failure post make the 4 axes. An 80.8 agentic-RAG figure especially means something only reproduced atop your RAG pipeline.

CodeBridge Mini Lab: 2 candidates on 4 axes

1. Pick Kolibri versus your current 1
2. Measure 4 axes (30 minutes each max):
   - 10-task success rate (domain tasks first)
   - VRAM footprint (before and after quantization)
   - tokens/sec (p50 and p95)
   - data boundary (any outside calls)
3. Call it:
   - wins domain tasks → consider adopting
   - wins averages only → hold (your jobs are not averages)

It is cost per success times the deployment boundary. Cheap and smart means nothing where data cannot travel. Winning the domain inside the boundary makes overall scores secondary.

Conclusion: open racing splits

One line to close.

Segmentation runs on active times domain times on-prem, not parameter counts.

Kolibri means 3.46B, not 78B. Store big, run light, fit one job, keep it home. Open-model picking moves from overall ranks to workload matching. One task for today: build a 10-task eval set from your work. With that set, any new model answers to the same 4 axes.

Further reading

References

Go deeper with a course

To practice workload-matched model picks and agent design, this course builds harness, loop, and graph layers exactly like the 4-axis eval here.