On October 5, Reflection AI introduced Beam, its first open-weight model:

Total:  501B parameters (sparse MoE)
Active: 23B parameters per token
Context: 1M tokens (extended in mid-training)
Modality: text-only

It pairs with Mistral Large 4 as the same trend: the question moves from "who is biggest" to "who does the same coding job with the least active compute." Beam answers that question most directly.

One accuracy note first: Beam is in final red-teaming, with weights, technical report, model card, and developer artifacts due later in October. Every score below is vendor-reported.

A model specialized from day one

Beam was designed for coding, reasoning, and agentic workloads — not as a general chatbot that happens to code, but as a workhorse for enterprise coding at low cost.

Reflection's headline results:

SWE-bench Verified      80.9
SWE-bench Pro v2-Hard   77.2
SWE-bench Pro v1        65.5
SWE-bench Multilingual  78.0
Terminal-Bench v2.1     80.1
DeepSWE v1.1            44.4
SWE Atlas Codebase QnA  34.6
MCP Atlas               78.7
GPQA Diamond            90.5
AIME 2026               97.8
HLE (no tools)          36.2

Through the lens of our coding benchmark comparison, each number asks something different: SWE-bench tests fixing issues in existing repos, Terminal-Bench tests carrying work through in a terminal, MCP Atlas tests wiring external tools. Beam targets all three axes at once.

Reflection also published its honest position: competitive with GLM-5.2, approaching Qwen 3.8-Max, while top open models like Kimi K3 lead on raw capability. Beam's bet is this sentence instead:

GLM-5.2-level reasoning at 3–4x less inference compute.

The published Figure 2 makes the claim visual: scores plotted against generation FLOPs on DeepSWE, HLE, and Terminal-Bench 2.1. Beam (dark green curve) sits up and to the left — high scores for little compute — while GLM-5.2 (orange) and the Qwen family sit to the right.

Score vs generation FLOPs. On DeepSWE, HLE, and Terminal-Bench 2.1, the Beam curve reaches high scores at low compute while GLM-5.2 and Qwen models sit further right

Source: Reflection announcement "Introducing Beam" (Oct 5, 2026), Figure 2. Generation FLOPs are estimated with the post's approximation (FLOPs ≈ 2 × active params × generated tokens), excluding prefill and serving overhead.

Why 10.5K GPUs for 4 weeks: RL as a scaling axis

The striking part of Beam's training is the RL side, more than pretraining:

Pretraining
- 23.8T tokens (web, public, licensed; code-centric curation)
- 6,144 x GB300 NVL72, under four weeks
- 52 layers, interleaved local+global attention
- fine-grained routed experts, near-uniform utilization

Mid-training
- context extended to 1M
- data pipelines built as a launchpad for RL

High-compute RL (the core)
- 10,500 x GB300 for four weeks
- 100M+ rollouts (up to 256K context)
- ~1.3B sandboxes created, ~1M environments
- 110K concurrent rollouts on average, 170K peak sandboxes
- fully asynchronous policy gradients
- stable learning even on day-old (107-version-stale) samples

Training this much RL means one thing: the model learned from massive volumes of its own attempt-fail-fix trajectories, not from memorizing answers. Reflection reports scores kept climbing with more RL compute and showed no plateau — and published the evidence as Figure 3, with DeepSWE, HLE, and Terminal-Bench scores rising against rollout count.

Scores vs cumulative RL rollouts. DeepSWE, HLE, and Terminal-Bench 2.1 scores rise as rollout count grows

Source: Reflection announcement "Introducing Beam" (Oct 5, 2026), Figure 3.

One neat mechanism: a controllable length penalty rewards successful solutions while discouraging needless tokens, teaching token efficiency first. Users then trade "short and cheap" against "long and accurate" through a reasoning-effort parameter — the same tradeoff our reasoning-effort guide covers, now shipped as a product feature.

How to verify the efficiency claim

Reflection's comparison formula is a published approximation:

FLOPs ≈ 2 x active params x mean generated tokens

MoE models plug in activated-per-token parameters rather than totals, and generated tokens include reasoning. Prompt prefill, attention, and serving overhead are excluded, so read it as a compute comparison estimate, not measured cost.

The right verification after the weights land is therefore not a FLOPs debate but measurement: run Beam, GLM, Qwen, and Kimi side by side on the same coding harness and record four things:

success  success rate (same test suite)
tokens   generated tokens per task (reasoning included)
clock    wall-clock time
memory   GPU memory (same quantization)

That experiment makes a great portfolio piece for a simple reason: many people quote benchmark tables, few own same-harness, same-condition measurement tables.

CodeBridge mini lab: prepare now, run on release day

No weights means no model testing yet. But you can prepare the harness now and run the moment they drop.

# NOTE: comparison harness sketch — run when weights release
models = ["beam-501B-A23B", "glm-5.2", "qwen-3.8", "kimi-k3"]

results = {}
for model in models:
    results[model] = run_coding_harness(
        model=model,
        tasks="swe-verified-sample-20",  # start with 20
        timeout_per_task="10min",
        record=["success", "tokens", "wall_clock", "gpu_memory"],
    )

# NOTE: rank by cost per success — not by score
rank_by_cost_per_success(results)
Do now:
1. Build a fixed 20-task set (your own repo issues recommended)
2. Freeze harness, timeouts, and recorded fields
3. Swap in models on release day
4. Put success / tokens / wall-clock / memory in one table

The Apache 2.0 license promise is another release-day checkpoint: confirm the LICENSE file says Apache 2.0 and the model card lists memory needs and quantized formats.

Conclusion: some workloads are won by light repetition, not heavy single shots

The reality Beam targets looks like this:

Enterprise coding/agent workloads
= not one hard problem,
  but thousands of medium problems run daily

In that world, cost per success decides the budget more than peak score on one problem. A model delivering GLM-5.2-class results at 23B active runs three times the attempts on the same budget — and where attempts equal coverage, efficiency is capability.

One sharp question to take away: recount your team's coding work not as "10 hard problems" but as "1,000 repeat jobs per month." If you can see the seat for an efficiency model like Beam in that table, preparing the harness today is the fastest way to be ready on release day.

Further reading

References

Go deeper with a course

To understand coding agents harness-first and run same-condition comparison experiments, practice stacking Claude Code, subagents, and MCP on real projects.