Reading about new open-weight models, you keep meeting sentences like this.

125B total
6B active

The natural thought follows.

So is it actually as light as a 6B model?

Half right.

Active parameters describe how much of the model joins the compute path for one token. They do not mean the whole weight set is 6B.

Qwen3.8-Flash-Next is a good example.

The Qwen3.8-Flash-Next numbers first

Per official materials, the main model is about 125B parameters, plus 51B of N-gram embedding and 4B of MTP.

Per token, about 6B parameters are activated against the 125B main model.

Main model
125B total
   ↓
one token processed
   ↓
about 6B activated

How is that possible?

MoE does not use every expert at once

In a normal dense model, every token passes through most of the weights.

Token
 ↓
whole dense layer
 ↓
whole dense layer
 ↓
...

A Mixture-of-Experts (MoE) model keeps a large expert pool and lets a router pick only some experts.

                         ┌─ Expert 1
                         ├─ Expert 2  ← picked
Token → Router ──────────┼─ Expert 3
                         ├─ Expert 4  ← picked
                         └─ Expert 5

So total capacity stays large while compute per token stays capped.

That is why total parameters and active parameters split apart.

What do total parameters tell you?

Total parameters roughly describe the full weight footprint.

With more experts, a model can spread more patterns and skills across them.

But those weights must live somewhere.

So total parameters connect closely to:

checkpoint size
GPU/CPU memory
multi-GPU placement
loading time
storage

What do active parameters tell you?

Active parameters help you understand the size of the path that actually computes one token's forward pass.

So they relate more to:

FLOPs per token
inference compute
some latency behavior
training compute

But you cannot predict latency from active parameters alone.

Expert routing, memory bandwidth, communication, batch size, KV cache, and parallelism all matter too.

The most common mistake: 6B active means a 6B GPU footprint

That reading is wrong.

A model holding 125B of expert weights must keep those weights in some storage layer — a GPU cluster, CPU memory, or similar — even if one token computes through only 6B. Different tokens can pick different experts.

Token A → Expert 2 + 7
Token B → Expert 1 + 9
Token C → Expert 5 + 8

So active parameters are not the same number as model weight residency.

CodeBridge Mini Lab: compute weight-only size and active compute separately

The calculation below is not real VRAM demand.

It is a weight-only rough estimate that excludes optimizer state, activations, KV cache, and framework overhead.

def weight_gb(params_billion, bits):
    bytes_per_param = bits / 8
    return params_billion * 1e9 * bytes_per_param / 1e9

for bits in [16, 8, 4]:
    total = weight_gb(125, bits)
    active = weight_gb(6, bits)

    print(
        f"{bits:>2}-bit | "
        f"125B total weights ≈ {total:>6.1f} GB | "
        f"6B active path ≈ {active:>5.1f} GB-equivalent"
    )

On plain BF16/FP16 math:

125B weights
→ about 250GB

6B worth of parameters
→ about 12GB

But do not conclude "it runs on a 12GB GPU" from the second number.

12GB is just the plain conversion of what 6B 16-bit parameters weigh.

The real inference setup is decided by expert placement, cache, activations, and the framework.

Qwen's 51B N-gram embedding is another layer

Beyond the main MoE, Qwen3.8-Flash-Next carries 51B of N-gram embedding parameters.

The interesting part: because lookup positions can be computed ahead of time, this embedding is designed to sit in host memory with async prefetch.

So capacity grows in more than one way.

Dense matrix parameters
MoE expert parameters
Lookup embedding parameters

Each has different compute cost and memory-access behavior.

That is why comparing models by a single "how many B" number misses more and more.

The 2.4T model reads the same way

Alibaba's Qwen3.8-2.4T-A95B is, as the name says, about 2.4T total parameters with about 95B activated.

Read the numbers separately here too:

2.4T
→ total capacity and weight storage scale

95B active
→ compute path picked per token

It is neither "2.4T of compute on every token"

nor "a 95B checkpoint because 95B is active."

Why does MoE keep growing?

The appeal is fairly intuitive.

More experts
→ more capacity

Capped active experts
→ capped compute per token

Qwen3.8-Flash-Next also uses an ultra-sparse MoE style: a big expert pool with few routed experts per token.

Of course nothing is free.

  • The router can pick wrong
  • Experts need load balancing
  • Multi-GPU communication costs
  • Weight storage burden
  • Serving architecture complexity

All of that comes along.

Conclusion: read both numbers together, not one B number

When you look at an MoE model, check at least these together.

Total Parameters
Active Parameters
Number of Experts
Experts per Token
Quantization
Context Length
KV Cache structure

A figure like 6B active is a clue for understanding compute efficiency, not a replacement for total model size.

So next time you see "125B with 6B active," read it like this.

It holds 125B of capacity, but each token computes through only a selected path.

Further reading

References

Go deeper with a course

If you want to turn architecture knowledge into workload routing — which task goes to which model — a guided course on using multiple AI tools covers the decision pattern.