You see 1M context in model announcements all the time now.
Hundreds of documents, a whole codebase, or a long conversation in one pass sounds attractive.
But one key question is missing:
Is accepting 1 million tokens the same as processing 1 million tokens cheaply and quickly?
No.
That is why sparse attention is back as a keyword in the Qwen4 line.
What makes full attention expensive?
Take a very simple self-attention picture.
If every token compares itself against every other token, then with sequence length n, the attention scores grow roughly like n x n.
1K tokens
→ about 1M relations
10K tokens
→ about 100M relations
100K tokens
→ about 10B relations
Real implementations add FlashAttention, KV cache, and GQA, so this number never becomes cost directly.
But the core problem stays:
The longer the context, the heavier it gets to look at every position with equal care.
The sparse attention idea is simple
For most questions, not every past token matters equally.
Say you read 500 pages of code docs and ask:
Find why retry runs twice in payment_service.py.
Only a small slice of the context is likely to matter.
Sparse attention asks roughly this:
Do we need to inspect the whole context precisely?
↓
Find the important positions first
↓
Spend more attention on that part
Qwen Sparse Attention looks at blocks first, not tokens
Qwen Sparse Attention (QSA) in Qwen3.8-Flash-Next compresses long sequences into micro-blocks.
A lightweight indexer estimates importance at block level, then selects the most relevant regions.
Long sequence
↓
[block][block][block][block][block]...
↓
lightweight indexer
↓
select important blocks
↓
attend to selected regions
Older sparse methods could spend a lot just "finding which tokens to pick." QSA tries to shrink that indexing cost by working at block level.
Qwen pairs Gated DeltaNet with QSA.
Simplified:
GDN
→ keeps remembering the whole past as compressed state
QSA
→ searches for the exact region to re-read
What did the vendor benchmarks show?
In Qwen's experiments at 1M-token conditions, the QSA attention kernel showed up to 7.6x prefill speedup and up to 4.9x decode speedup against its baseline.
In a serving test assuming 90% prefix-cache hits, Qwen3.8-Flash-Next reported 8.6x higher 1M-context prefill throughput than Qwen3.7-Plus.
But those numbers are vendor results under specific benchmark and serving conditions.
So do not generalize:
Sparse attention is always 8.6x faster
Real speed shifts a lot with hardware, batch size, cache hits, prompt length, and framework.
CodeBridge Mini Lab: feel the scale gap with numbers
This is not a real QSA implementation.
It is a toy calculation to build intuition for sparse attention.
Full attention looks at every token pair. The sparse version looks at most at 2,048 related positions per token.
lengths = [8_192, 32_768, 131_072, 1_000_000]
budget = 2_048
for n in lengths:
full = n * n
sparse = n * min(n, budget)
ratio = full / sparse
print(
f"{n:>10,} tokens | "
f"full={full:>15,} | "
f"sparse={sparse:>15,} | "
f"ratio={ratio:>8.1f}x"
)
This code does not predict real latency.
It shows how fast the gap grows between looking at every pair and looking at a limited candidate set as length grows.
Sparse attention has a price too
Looking at only some important tokens means you can miss information when you pick wrong.
So sparse architectures face new questions:
What counts as important?
How do we find it fast?
What if we miss key positions?
Should layers share selection criteria?
That is why QSA adds a separate lightweight indexer.
The hard part is not "looking at less." It is picking correctly, then looking at less.
Why RAG survives even with 1M context
Long-context models and RAG often look like rivals.
With 1M context,
why not paste every document in?
Even with a bigger window, these problems stay:
- input token cost
- prefill latency
- retrieval accuracy for key facts
- document updates
- access permissions
- source tracking
Sparse attention is a model-side technique for lowering long-context cost.
RAG is a system-side design for choosing what to feed the model.
They solve problems on different layers.
So they will likely be used together, not as replacements.
Conclusion: long-context competition is shifting from how much you fit to how cheaply you find
Context window numbers keep growing.
But supporting a million tokens does not mean stuffing in a million tokens is good design.
The sharper question ahead is this:
Inside a long context, how accurately can you re-find what you need with how little compute?
That is exactly why QSA-style sparse attention is interesting.
Related posts
- How the Qwen4 architecture changes
- With a 1M context window, is RAG still needed?
- What is RAG? From retrieval to answer
References
Go deeper with a course
If you want to pair long-context models with retrieval, chunking, and evals in real document projects, learn by building.