On October 2, Cloudflare opened the Web Search API in beta. Ceramic.ai, Exa, and Linkup serve through one API shape, and requests live under AI Gateway logging, analytics, billing, and access control.

Short news, big meaning. Web search is becoming an infrastructure layer you observe, swap, and cost-manage like model inference, not just one more agent tool.

What shipped: 3 providers, 3 paths

Item Detail
Providers Ceramic.ai, Exa, Linkup (pick one)
Call paths AI Gateway direct, REST endpoint, Workers binding (one-line call)
Observability Shows in the same Gateway logs and analytics as inference calls
Billing Draws from existing AI Gateway credits
Data All three support Zero Data Retention via Cloudflare, BYOK available
// Workers binding shape (from official docs)
const response = await websearch({
  gatewayId: "default",
  query: "What are some fun things to do in Salt Lake City as fall approaches?",
  provider: "exa",
  limit: 5,
});
const results = await response.json();

A welcome line sits in the changelog. Ground answers in live information "instead of guessing URLs or relying on a model's training cutoff." It blocks the hallucinated-URL failure at the search layer.

Why infrastructure: RAG's search step moved outside

Recall the RAG structure. Of parse, chunk, search, rerank, and generate, search served your own docs. Now web search gets the same treatment.

Before: search = 1 function the agent calls (separate logs and bills)
Now:    search = a member inside the Gateway (unified logs, analytics, billing, access)
→ swap providers without code changes
→ compare quality, latency, and cost on the same sheet as inference

The RAG failure post said to touch embeddings only when search is the top culprit. Web search works the same. Measure provider quality, latency, and cost first, then switch.

5 picking metrics: record them together

Provider scorecard (refresh weekly):
[ ] quality: evidence hit rate on 20 of your questions
[ ] latency: p50 and p95 response times
[ ] cost: per 1k calls (in Gateway credits)
[ ] citation quality: do titles, URLs, and snippets cite straight into answers
[ ] data retention: ZDR support plus BYOK availability

The last two are very 2026. Citation quality links to "sources are part of the answer" from the finance post. ZDR and BYOK link to data control from the security post. Picking search is becoming picking security.

CodeBridge Mini Lab: 20 questions across 3 providers

1. Pick 20 of your questions (half routine, half fresh-info)
2. Run the same 20 on all 3 providers (fixed limits and options)
3. Record the 5 metrics (table above)
4. Call it:
   - tied hit rates → decide on latency and cost
   - gaps only on fresh-info questions → split by use (1 default plus 1 fresh)
   - weekly drift check in Gateway logs (does one degrade some day)

Like prompt caching, check cache hits before searching repeats. Frequent questions deserve a cache lookup first.

Conclusion: search is a slot too

The gated-models post said to slot models. Search too.

Never depend on one provider. Keep it a swappable slot behind the Gateway.

One task for today: check where your agent's search calls log and what they cost. Once that number shows, search turns from tool into infrastructure. Infrastructure can be managed.

Further reading

References

Go deeper with a course

To design production RAG through parsing, search, and citations, this course expands Classic RAG into GraphRAG and Agentic RAG, exactly like the provider comparison here.