A 1M context no longer surprises anyone on frontier models. GPT-6 Astra offers a context window of about 1.05M tokens.
So the natural question follows.
If everything fits, do we still need RAG?
The short answer: context size and retrieval solve different problems.
A context window is not storage
A context window is the input range a model can consult in one request.
Database / Files
↓
put needed content into the prompt
↓
Context Window
↓
Model response
1M context does not store documents inside the model permanently. If the next request needs the same material, you must supply it again or keep state in a session or harness.
Long context changes the price structure too
The GPT-6 Astra API supports 1,050,000 tokens of context, but prompts over 272K input tokens fall under a higher long-context rate in the official docs.
In other words, "it fits" and "it is economical" are different questions.
Send a 600K-token document bundle on every request:
Question 1 → 600K input
Question 2 → 600K input again
Question 3 → 600K input again
Cost and latency can both balloon.
RAG is not only for what does not fit
Narrowing the evidence is also a core value of RAG.
Full 600K tokens
↓ retrieval
relevant 8K tokens
↓
LLM
That structure helps beyond cost:
- tracing which evidence was used
- per-document access control
- updating fast-changing material
- source citations
- scaling to a large corpus
So retrieval keeps its operational advantages even as context windows grow.
When is long context the better pick?
Some work breaks when retrieval slices context apart.
Examples:
- Reviewing how clauses interact across a whole contract
- Grasping an entire repository's architecture
- Analyzing the flow of a long interview or meeting
- Comparing broadly across many documents
For these, holding a wide context can win.
CodeBridge Mini Lab: full context vs retrieval
Prepare the same 20 questions over the same document set.
Method A:
Include every document in every prompt
Method B:
retrieval top-k
→ send only relevant chunks to the model
Measure this.
answer correctness
source correctness
input tokens
latency
cost per question
Use many questions, not one. The cost problem of long context shows best across repeated questions.
Hybrid is the realistic answer in most cases
You do not have to pick one.
RAG picks candidate documents
↓
load the relevant documents fully or in wide spans
↓
a long-context model synthesizes
This balances retrieval fetching too-small chunks against stuffing the whole corpus in every time.
Conclusion: 1M context adds an option, it does not end RAG
Big context windows are powerful. But as data grows, cost, permissions, freshness, and evidence tracing stay unsolved.
So the useful question is not:
RAG or 1M context?
but:
In this task, which information should we narrow early, and which should stay wide?
Further reading
References
Go deeper with a course
If you want to practice deciding what to narrow and what to keep wide across real RAG architectures, a guided course builds the judgment step by step.