Uber published its MCP Gateway. The scale differs first. Over 800 MCP servers and 5,000-plus tools in operation. Registry as control plane, Proxy Gateway as data plane, and AutoCrawler auto-discovering existing HTTP, gRPC, and TChannel APIs into the registry.
Then scale surfaced a new problem. Thousands of tool schemas cannot fit agent context. Uber's 3 answers are this post's subject.
Structure: discover automatically, register dark
AutoCrawler (Cadence workflows, IDL registry subscription)
→ scan new services, APIs, schema changes
→ create and update MCP server shapes
→ generate tool definitions and schemas (LLM assists descriptions)
→ register disabled-by-default
→ owners review, then enable
"Disabled-by-default" is the key. Discover automatically, ship dark by default. Scale without dragging service teams onto the critical path. The MCP comparison said "ship tools as a standard." Uber automated the shipping itself.
3 bottlenecks, 3 answers
| Bottleneck | Answer | Detail |
|---|---|---|
| Thousands of schemas never fit context | Omni MCP incremental discovery | Find-it-when-needed shape |
| Responses run too large | Response Projection | Return only needed response fields |
| Results eat context | Code Mode | Agents save results to files, grep only needed parts |
Code Mode delights most. Coding agents write results to files and read back only needed slices, and it is reportedly Uber coding agents' default MCP pattern. A head-on answer to the rereading problem from the 5x tokens post. Do not read it. Store it.
Numbers from the reported agentic SDLC talk add context. Model gateway at 800-plus projects and 100M-plus daily requests, Omni, CLI, and code-mode saving 40-plus percent fleetwide tokens, agent-authored PRs past 70 percent. Scale-born optimization verified at scale.
Shrink it to your size: 3 registry steps
The pattern transfers below Uber scale.
Step 1, Registry (today):
- 1 page listing MCP servers and tools in use (name, purpose, owner)
- delist zombies untouched for 2 weeks
Step 2, search (this week):
- stop full-schema injection
- switch to search, schema-on-demand, invoke
- fetch only what the moment needs
Step 3, Projection plus Code Mode (this month):
- field-pinned returns on large responses
- file-save plus grep on long results (never straight into context)
The "only what is needed" rule from the Agents API harness post and the routing post applies to discovery too.
CodeBridge Mini Lab: a 10-tool diet
1. Count injected tool schemas now (count, tokens)
2. Delist tools untouched for 2 weeks
3. Pick the 1 largest response:
[ ] Does it return only needed fields
[ ] Can file-save plus grep replace it
4. Compare in 1 week:
[ ] tokens per task delta
[ ] tool-pick accuracy delta (beef up search blurbs on drops)
Reads better beside time-to-task from the speed guide. Lighter context finishes sooner.
Conclusion: hit the next bottleneck early
One line to close.
Past connection, the MCP bottleneck is discovery times context times governance.
Set the structure before 800 tools. One registry page, search ordering, and trimming big responses work at 10 tools too. Context grows pricier as tools multiply. Build the finding structure before the price does.
Further reading
References
- Uber: Designing MCP Gateway
- Uber: Running a Software Factory Efficiently at Uber Scale
- DEV: How Uber exposes existing APIs through its MCP Gateway (Oct 5, 2026)
Go deeper with a course
To design tool discovery, invocation, and verification as a structure, this course builds harness, loop, and graph layers exactly like the Registry, search, and invoke here.