Uber published its MCP Gateway. The scale differs first. Over 800 MCP servers and 5,000-plus tools in operation. Registry as control plane, Proxy Gateway as data plane, and AutoCrawler auto-discovering existing HTTP, gRPC, and TChannel APIs into the registry.

Then scale surfaced a new problem. Thousands of tool schemas cannot fit agent context. Uber's 3 answers are this post's subject.

Structure: discover automatically, register dark

AutoCrawler (Cadence workflows, IDL registry subscription)
  → scan new services, APIs, schema changes
  → create and update MCP server shapes
  → generate tool definitions and schemas (LLM assists descriptions)
  → register disabled-by-default
  → owners review, then enable

"Disabled-by-default" is the key. Discover automatically, ship dark by default. Scale without dragging service teams onto the critical path. The MCP comparison said "ship tools as a standard." Uber automated the shipping itself.

3 bottlenecks, 3 answers

Bottleneck Answer Detail
Thousands of schemas never fit context Omni MCP incremental discovery Find-it-when-needed shape
Responses run too large Response Projection Return only needed response fields
Results eat context Code Mode Agents save results to files, grep only needed parts

Code Mode delights most. Coding agents write results to files and read back only needed slices, and it is reportedly Uber coding agents' default MCP pattern. A head-on answer to the rereading problem from the 5x tokens post. Do not read it. Store it.

Numbers from the reported agentic SDLC talk add context. Model gateway at 800-plus projects and 100M-plus daily requests, Omni, CLI, and code-mode saving 40-plus percent fleetwide tokens, agent-authored PRs past 70 percent. Scale-born optimization verified at scale.

Shrink it to your size: 3 registry steps

The pattern transfers below Uber scale.

Step 1, Registry (today):
- 1 page listing MCP servers and tools in use (name, purpose, owner)
- delist zombies untouched for 2 weeks

Step 2, search (this week):
- stop full-schema injection
- switch to search, schema-on-demand, invoke
- fetch only what the moment needs

Step 3, Projection plus Code Mode (this month):
- field-pinned returns on large responses
- file-save plus grep on long results (never straight into context)

The "only what is needed" rule from the Agents API harness post and the routing post applies to discovery too.

CodeBridge Mini Lab: a 10-tool diet

1. Count injected tool schemas now (count, tokens)
2. Delist tools untouched for 2 weeks
3. Pick the 1 largest response:
   [ ] Does it return only needed fields
   [ ] Can file-save plus grep replace it
4. Compare in 1 week:
   [ ] tokens per task delta
   [ ] tool-pick accuracy delta (beef up search blurbs on drops)

Reads better beside time-to-task from the speed guide. Lighter context finishes sooner.

Conclusion: hit the next bottleneck early

One line to close.

Past connection, the MCP bottleneck is discovery times context times governance.

Set the structure before 800 tools. One registry page, search ordering, and trimming big responses work at 10 tools too. Context grows pricier as tools multiply. Build the finding structure before the price does.

Further reading

References

Go deeper with a course

To design tool discovery, invocation, and verification as a structure, this course builds harness, loop, and graph layers exactly like the Registry, search, and invoke here.