When an agent underperforms, the model gets blamed first. Tweak the prompt, attach another tool, and if that fails, switch to a bigger model.
Postman's Agent Mode operations story, published October 9 on the AWS tech blog, flips that order. Running an agent on top of an 11-year-old product used by 40 million developers, 500,000 organizations, and 98% of the Fortune 500, the team found the sorest bottleneck was not the model but tool count and context. Past roughly 40 visible tools, tool-selection errors rose visibly, so today the system picks about 15 of 170+ tools per task.
The 40-tool cliff: what breaks as tools grow
Postman's first instinct was reasonable. Expose small atomic tools — open a request, change one field, fetch metadata — and the agent would move precisely. The result was precision that felt slow. Every step round-tripped to the model, so even fast calls felt sluggish.
Adding more tools caused a different sickness. Beyond about 40, these failures grew:
1. Calling tools that do not exist
→ inventing names outside the schema
2. Calls with wrong arguments
→ valid schema, wrong values
3. Plausible but context-wrong calls
→ a sensible-looking tool, irrelevant to this task
Bigger and newer models reduce the rate but do not remove it. The lesson connects to this post's first line: budget tools like tokens. Count itself eats performance, so tool exposure is a resource to cap, not to grow forever.
15 of 170: narrowing what the agent sees
The current design is simple. Do not show everything. Show only what the task needs.
Full catalog: 170+ tools
→ root agent queries a vector DB (tool embeddings) for candidates
→ a context-isolated sub-agent receives ~15 tools
→ the sub-agent acts within that narrow view
This is the Figure 3 flow from the AWS post. The root selects; the worker acts with a narrow field of view. It matches the harness engineering post: the control structure around the model matters more than model shopping. Narrow the room so silly mistakes cannot happen.
The pattern matters most in mature products. In Postman, tools were implicitly coupled to interface state: you needed an open tab to read, and running things opened tabs as side effects. The team decoupled that. Background request execution sends without opening a tab, while state-changing actions still need human approval. Native Git integration runs on top of this separation.
Two kinds of context: broad-shallow, narrow-deep
The second bottleneck was context. Postman splits it in two:
1. Background context (broad, shallow)
→ gathered automatically, minified, laid underneath
2. User-selected context (narrow, deep)
→ a dedicated handler per entity distills
exactly what the agent needs
One painful failure is worth quoting. Serializing rendering objects and handing them over failed — structures good for rendering were bad for reasoning. Screens and reasoning want different shapes. So each entity got its own distiller, and for open-ended fields like request descriptions, OpenAPI specs, and payloads, the team is exploring filesystem-backed handlers instead of per-handler truncation logic. Truncation is architecture, not a hack. It points the same way as the three-layer agent memory post.
One schema beats fifty read tools
The third pattern is data access. Instead of N tools for N narrow views, Postman consolidated into one catalog tool and gave the agent table schemas to query directly. The ClickHouse example in the AWS post is representative:
Table: http_events_summary_1d
Columns: service_id, total_requests(countMerge),
total_errors, error_rate_pct,
avg_latency_ms(avgMerge), p95_latency_ms(quantileMerge 0.95)
Agent writes:
WHERE bucket_1d >= today() - 7
GROUP BY service_id
HAVING p95_latency_ms < 100
ORDER BY error_rate_pct DESC
Model the data once instead of adding a tool per question. Fewer tools, more expressive power. Read it together with the Uber MCP gateway post: discovery cost is part of the context budget too.
Running on Bedrock: routing, geography, caching
That is also why this story sits on the AWS blog. Agent Mode runs on Amazon Bedrock, and the point is operations, not just "we use Claude."
Model flexibility: converse / InvokeModel APIs, no single-model lock-in
→ light models for fast work, large models for hard reasoning
→ Opus 4.6 as default; switching is a config change
Cross-Region inference:
→ modelId points at a geographic inference profile
→ Geographic: throughput + regional boundary / Global: worldwide, ~10% cheaper
→ IAM and SCPs must permit every destination
Data protection:
→ encrypted in transit and at rest, prompts never train models
→ data_retention_mode=none on supported models (check per model)
Trust controls:
→ human approval before state changes, Bedrock Guardrails PII redaction (enterprise)
→ Secret Scanner keeps keys and tokens on the machine
Prompt caching:
→ stable prefix (system, instructions, core tools, knowledge, conversation) resent each turn
→ 1-hour checkpoint (near-immutable) + 5-minute checkpoint (variable layer), 1h first
→ verify with cacheReadInputTokens / cacheWriteInputTokens and time-to-first-token
The caching note continues the prompt caching cost post. Caching is not done at install time; it is verified with read/write tokens and perceived latency.
The AWS post's seven builder takeaways are worth keeping: budget tools, prefer schema-aware reads, decouple actions from UI state, engineer context deliberately, treat the window as scarce, ship docs with features, and operate routing plus caching on Bedrock. Every one says "look at structure before swapping models."
CodeBridge Mini Lab: a 15 vs 40 vs all experiment
Turn the briefing's action item into a spec you can run today on your own agent, even without Postman:
1. Fix three conditions:
- A: expose only 15 task-relevant tools
- B: expose 40 tools
- C: expose everything
2. Run the same 5 tasks 3 times per condition:
- nonexistent-tool calls
- argument error rate
- context errors (plausible but irrelevant picks)
3. Measure together:
- selection accuracy + latency (p50/p99) + token usage
- one table: "accuracy rose, but how much did latency cost?"
Once that table exists, "let's add more tools" meetings become "which 15 do we keep" meetings. That is the real harvest of the Postman story.
Conclusion: agent improvement starts with subtraction
One line to summarize.
When an agent wanders, check exposure area before model size.
15 of 170, the 40-tool cliff, separating screen objects from reasoning context, one schema instead of fifty tools — all are techniques of showing less, not giving more. Model upgrades come later. What pays today is a catalog diet and context distillation. Count your agent's tools now. Past 40, you have walked into the minefield Postman already crossed.
Further reading
- What is harness engineering? The control structure around models
- The three-layer structure of AI agent memory
- Cutting AI API cost with prompt caching
References
- AWS ML Blog: How Postman runs Agent Mode for 40 million developers on Amazon Bedrock
- Postman: Agent Mode product page
- Postman Learning Center: Agent Mode docs
- Anthropic: Postman customer story (Claude on Bedrock)
Go deeper with a course
Tool selection, context isolation, and splitting roles through loops and graphs continue directly from this post's 170-to-15 pattern. To design agents that execute and verify — beyond prompting — start here.