One number in Claude Opus 5.5 catches the eye first: the 1M-token context window. Per Anthropic's official docs, Opus 5.5 supports 1M tokens of context with up to 128K output tokens, and adaptive thinking stays always on.

The number alone invites a tempting thought:

"Can't I just drop the whole repo in?"

In practice, treating long context as memory capacity backfires.

A context window is not a data warehouse

Context is closer to a workspace the model consults for the current request. Fitting more information in doesn't mean every fact gets equal weight, or that conflicting requirements resolve themselves.

A large project holds these side by side:

  • Current code
  • An outdated README
  • The latest ADR (Architecture Decision Record)
  • Old issue threads
  • Test code
  • Dead implementations nobody removed

All look relevant, yet each may carry a different era's truth.

As context grows, the key skill isn't stuffing more in. It's telling the model what to trust.

CodeBridge mini experiment: plant three conflicting requirements

Try this in a sample project or your own repo.

First, prepare three files with conflicting instructions about the same feature:

README.md        : Users sign in with email.
docs/auth-v2.md  : Email + passkey are supported.
TODO-old.md      : Drop email login, keep social login only.

Then don't ask for implementation yet. Ask this instead:

I want to implement the login requirements in this repo.
Before changing code, find requirements that conflict or may be outdated.

For each claim, first summarize:
- Evidence file
- Last modified time (if known)
- Conflicting claims
- Questions to confirm with a human before implementing

This experiment doesn't grade one right answer. Watch for these:

  • Does it spot the conflict?
  • Does it avoid mashing all files together?
  • Does it refuse to crown the newest doc as truth?
  • Does it surface uncertainty before implementing?

Long context proves its value not in volume, but in managing uncertainty across volume.

Read adaptive thinking the same way

Opus 5.5 always uses adaptive thinking. The model scales its own reasoning to the task, without you ordering "think hard" every time.

But it still doesn't define done for you.

"Refactor this code" loses to verifiable conditions like these:

Goal: remove duplicated logic in PaymentService

Done means:
- Public API signatures unchanged
- All existing tests pass
- No new dependencies
- Summary of changed files with reasons
- Untested items listed separately

Even with a great model, no definition of done means you can't separate "plausible edits" from "finished work."

When does 1M tokens actually help?

Big context shines in situations like these:

  • Code migrations spanning many modules
  • Cross-checking long technical docs against code
  • Root-causing across massive logs and configs
  • Comparing requirement versions
  • Long-running agent sessions

The reverse also holds: stuffing the whole repo in to fix a bug that needs three files wastes money and focus.

Failure patterns in long context

1. Including everything "just in case"

Piling in possibly-useful material blurs the line between signal and noise.

2. Never stating priorities

When code and docs collide without guidance on which to trust, the model improvises.

3. Trying to finish in one giant request

Cramming analyze, plan, implement, and verify into one call makes a wrong early assumption expensive to fix.

This is where the loop engineering and harness engineering perspectives connect.

Related posts

References

Conclusion: bigger context demands better organizing

The 1M context in Claude Opus 5.5 is genuinely powerful. Read as "now I can paste everything at once," though, you will miss its strength.

Bigger context makes these questions matter more:

What is current? What is trustworthy? What conflicts? What must a human confirm first?

Good long-context technique is less about stuffing documents in, and more about turning big work into verifiable flows.

Go deeper with a course

If you want hands-on practice splitting big work into loops with clear completion criteria, a structured course on agent workflows helps.