When you pick a model, you probably compare intelligence scores and prices. I did too.

Then you wire up a voice agent or a coding agent, and a third axis pops out. Speed. A brilliant answer that arrives 10 seconds late is unusable.

This guide untangles three speed metrics: output speed, time to first token, and task completion time. Then it adds October measurements and practical ways to feel faster.

Separate the three metrics first

Request ──> [thinking + input processing] ──> first token ──> steady tokens ──> 500 tokens done
            └──── latency ────┘                                     └── output speed ──┘
            └──────────────── perceived end-to-end time ─────────────────────────────┘
Metric Meaning What it feels like
Output speed (tokens/s) Tokens per second while generating A long answer flowing smoothly
Latency (TTFT, seconds) Time to first token, including thinking for reasoning models The awkward silence before "yes?"
Task completion time (minutes) Output tokens ÷ speed Waiting for an agent to finish the job

Reasoning models think long before the first token, so latency runs high. Fast output can't save a slow-feeling start. For short answers, latency — not output speed — decides the feel.

October benchmarks: who is fast and who is nimble

Measured by Artificial Analysis.

Category Leader Figure
Output speed Celeris-1 About 1,502 tokens/sec
Output speed runner-up Mercury 2 About 811 tokens/sec
Time to first token Gemini 2.5 Flash-Lite (non-reasoning) About 0.33 sec
Latency runner-up Nemotron 3 Nano Omni 30B About 0.38 sec
Latency third Gemini 2.5 Flash (non-reasoning) About 0.45 sec

Frontier news matters too. Opus 5.5 generates output over 30% faster than its predecessor, and Fast mode for Claude Code and the Claude Platform promises up to 2.5x speed. Fast mode costs double, though: $8 input and $40 output.

Fast model ≠ all-rounder
- Celeris/Mercury class: fastest output, check intelligence separately
- Flash-Lite class: fastest first response, best for short chats and voice
- Frontier reasoning class: smartest, but thinking time slows first response
→ Different jobs have different winners

For agents, look at time-to-completion

A coding agent feels fast when the job finishes, not when tokens fly. The formula is simple:

Task completion time ≈ output tokens ÷ output speed (+ thinking time)

Examples:
- Light task using 20K tokens at 800 tok/s → about 25 sec
- Deep task using 120K tokens at 800 tok/s → about 150 sec
→ Using fewer tokens is also the shortcut to feeling faster

That's why Opus 5.5 stresses token efficiency, and why GPT-6 Astra defining the Pareto frontier at 27K tokens per task is a speed story. Models that use fewer tokens finish first at equal speed. The logic from the cost-per-successful-task post applies to time as well.

If you serve models yourself, watch p95 and concurrency — not averages. Concurrent load breaks p95 first. The AIPerf guide covers the full toolkit: TTFT, streaming gaps (ITL), throughput, and concurrency together.

Four patterns that improve perceived speed

1. Stream by default

Waiting for the full answer before showing anything halves the perceived speed. Stream from the first token. Watch for stalls (ITL) too.

2. Let the small model greet, the big model review

Fast model: first triage, drafts, progress updates (owns latency)
  ↓
Heavy model: hard analysis, final review (owns intelligence)
  ↓
Users watch progress while waiting (feels faster)

This applies the multi-AI-tools pattern and the Luna-Sol-Astra routing pattern to speed.

3. For voice, set a 1-second budget first

As the Gemini Live voice agent post shows, running tools mid-conversation is the whole game in voice. Rough human thresholds:

Under 0.5 sec: feels natural
Around 1 sec:  tolerable limit
Over 2 sec:    "did it disconnect?" checks begin
→ Piping max-effort reasoning into the first voice reply will fail
→ Open light, run deep work in the background, narrate progress aloud

4. Cache away the thinking time

Prompt caching trims input processing for repeated instructions and documents. Opus 5.5 cache reads cost $0.20 — roughly 95% off — and Gemini 4 Argon advertises 95% off caching too. You save time as well as money. See the prompt caching post for the math.

CodeBridge Mini Lab: measure a 500-token response

Like Artificial Analysis end-to-end response times, time a 500-token completion yourself:

Run the same prompt (1 code review, 1 doc summary) on 2 candidates:

- To first token: __ sec (stopwatch or logs)
- To 500 tokens done: __ sec
- Streaming smoothness: 1–5
- Stream + measure p95 over 3 runs, compare worst case, not average

Verdict: pick the better worst case, not the better average
(users remember the worst, not the average)

Related posts

References

Conclusion: summarize speed in three lines

When you pick a model, add two lines next to the intelligence score:

How many seconds to the first token, and to 500 tokens done?

And for agents, one more line:

How many minutes until the job is done?

Smart-but-slow and fast-but-shallow both earn their keep. They just belong in different seats. The moment you split first-response work from finishing work, the speed table turns into money.

Go deeper with a course

If you want guided practice splitting fast first responses from heavy analysis, a structured course on agent execution design helps.