On October 5, Reflection AI introduced Beam, its first open-weight model:
Total: 501B parameters (sparse MoE)
Active: 23B parameters per token
Context: 1M tokens (extended in mid-training)
Modality: text-only
It pairs with Mistral Large 4 as the same trend: the question moves from "who is biggest" to "who does the same coding job with the least active compute." Beam answers that question most directly.
One accuracy note first: Beam is in final red-teaming, with weights, technical report, model card, and developer artifacts due later in October. Every score below is vendor-reported.
A model specialized from day one
Beam was designed for coding, reasoning, and agentic workloads — not as a general chatbot that happens to code, but as a workhorse for enterprise coding at low cost.
Reflection's headline results:
SWE-bench Verified 80.9
SWE-bench Pro v2-Hard 77.2
SWE-bench Pro v1 65.5
SWE-bench Multilingual 78.0
Terminal-Bench v2.1 80.1
DeepSWE v1.1 44.4
SWE Atlas Codebase QnA 34.6
MCP Atlas 78.7
GPQA Diamond 90.5
AIME 2026 97.8
HLE (no tools) 36.2
Through the lens of our coding benchmark comparison, each number asks something different: SWE-bench tests fixing issues in existing repos, Terminal-Bench tests carrying work through in a terminal, MCP Atlas tests wiring external tools. Beam targets all three axes at once.
Reflection also published its honest position: competitive with GLM-5.2, approaching Qwen 3.8-Max, while top open models like Kimi K3 lead on raw capability. Beam's bet is this sentence instead:
GLM-5.2-level reasoning at 3–4x less inference compute.
The published Figure 2 makes the claim visual: scores plotted against generation FLOPs on DeepSWE, HLE, and Terminal-Bench 2.1. Beam (dark green curve) sits up and to the left — high scores for little compute — while GLM-5.2 (orange) and the Qwen family sit to the right.

Source: Reflection announcement "Introducing Beam" (Oct 5, 2026), Figure 2. Generation FLOPs are estimated with the post's approximation (FLOPs ≈ 2 × active params × generated tokens), excluding prefill and serving overhead.
Why 10.5K GPUs for 4 weeks: RL as a scaling axis
The striking part of Beam's training is the RL side, more than pretraining:
Pretraining
- 23.8T tokens (web, public, licensed; code-centric curation)
- 6,144 x GB300 NVL72, under four weeks
- 52 layers, interleaved local+global attention
- fine-grained routed experts, near-uniform utilization
Mid-training
- context extended to 1M
- data pipelines built as a launchpad for RL
High-compute RL (the core)
- 10,500 x GB300 for four weeks
- 100M+ rollouts (up to 256K context)
- ~1.3B sandboxes created, ~1M environments
- 110K concurrent rollouts on average, 170K peak sandboxes
- fully asynchronous policy gradients
- stable learning even on day-old (107-version-stale) samples
Training this much RL means one thing: the model learned from massive volumes of its own attempt-fail-fix trajectories, not from memorizing answers. Reflection reports scores kept climbing with more RL compute and showed no plateau — and published the evidence as Figure 3, with DeepSWE, HLE, and Terminal-Bench scores rising against rollout count.

Source: Reflection announcement "Introducing Beam" (Oct 5, 2026), Figure 3.
One neat mechanism: a controllable length penalty rewards successful solutions while discouraging needless tokens, teaching token efficiency first. Users then trade "short and cheap" against "long and accurate" through a reasoning-effort parameter — the same tradeoff our reasoning-effort guide covers, now shipped as a product feature.
How to verify the efficiency claim
Reflection's comparison formula is a published approximation:
FLOPs ≈ 2 x active params x mean generated tokens
MoE models plug in activated-per-token parameters rather than totals, and generated tokens include reasoning. Prompt prefill, attention, and serving overhead are excluded, so read it as a compute comparison estimate, not measured cost.
The right verification after the weights land is therefore not a FLOPs debate but measurement: run Beam, GLM, Qwen, and Kimi side by side on the same coding harness and record four things:
success success rate (same test suite)
tokens generated tokens per task (reasoning included)
clock wall-clock time
memory GPU memory (same quantization)
That experiment makes a great portfolio piece for a simple reason: many people quote benchmark tables, few own same-harness, same-condition measurement tables.
CodeBridge mini lab: prepare now, run on release day
No weights means no model testing yet. But you can prepare the harness now and run the moment they drop.
# NOTE: comparison harness sketch — run when weights release
models = ["beam-501B-A23B", "glm-5.2", "qwen-3.8", "kimi-k3"]
results = {}
for model in models:
results[model] = run_coding_harness(
model=model,
tasks="swe-verified-sample-20", # start with 20
timeout_per_task="10min",
record=["success", "tokens", "wall_clock", "gpu_memory"],
)
# NOTE: rank by cost per success — not by score
rank_by_cost_per_success(results)
Do now:
1. Build a fixed 20-task set (your own repo issues recommended)
2. Freeze harness, timeouts, and recorded fields
3. Swap in models on release day
4. Put success / tokens / wall-clock / memory in one table
The Apache 2.0 license promise is another release-day checkpoint: confirm the LICENSE file says Apache 2.0 and the model card lists memory needs and quantized formats.
Conclusion: some workloads are won by light repetition, not heavy single shots
The reality Beam targets looks like this:
Enterprise coding/agent workloads
= not one hard problem,
but thousands of medium problems run daily
In that world, cost per success decides the budget more than peak score on one problem. A model delivering GLM-5.2-class results at 23B active runs three times the attempts on the same budget — and where attempts equal coverage, efficiency is capability.
One sharp question to take away: recount your team's coding work not as "10 hard problems" but as "1,000 repeat jobs per month." If you can see the seat for an efficiency model like Beam in that table, preparing the harness today is the fastest way to be ready on release day.
Further reading
- SWE-bench vs Terminal-Bench vs ProgramBench
- What do active parameters mean in MoE?
- Judge models by cost per successful task
References
Go deeper with a course
To understand coding agents harness-first and run same-condition comparison experiments, practice stacking Claude Code, subagents, and MCP on real projects.