NVIDIA released AIPerf as the successor to GenAI-Perf, and its September 18 technical blog lays out the full picture. One sentence captures it:
If the client becomes the bottleneck, the server measurement means nothing.
It is a good case of LLM benchmarks moving from model scores to real serving performance. This post covers what changed in AIPerf and what to measure if you run local models or an inference server.
Why a new tool was needed
Older benchmarkers used a single-process design, so the GIL choked them as concurrency rose. When the measuring client tires first, you measured the client limit, not server performance.
AIPerf switched to a multiprocess design. Worker processes apply load, record processors compute results, and a control plane coordinates everything over ZMQ messages. Datasets sit in memory-mapped files that workers read directly, and a timing manager issues requests with a credit system. The design goal: keep the benchmark client from becoming the bottleneck at high concurrency.
Control Plane (what to send, when, how much)
→ Timing Manager (issues credits)
→ Workers (send HTTP requests, measure response times)
→ Record Processors (compute metrics in parallel)
→ Records Manager (aggregate and export)
Coverage widened too. More than 15 endpoint types including chat, public datasets like ShareGPT, and trace-replay formats such as Mooncake, Baseten, and WEKA AgentX live in one tool. You can pick anything from a light smoke test with synthetic prompts to replaying production traffic as-is.
Being able to pick the load shape is the point
Average latency alone produces numbers far from production. Real traffic does not arrive evenly — it clumps and idles. AIPerf lets you choose the load shape.
constant — requests at fixed intervals
poisson — bursty arrivals like real queueing
gamma — bursty arrivals with adjustable clumping
ramp — slowly raise concurrency and request rate
fixed-schedule — replay a trace by its own timestamps
user-centric — think in turns per user
The usage examples work as documented.
# Measure streaming serving with ShareGPT data
aiperf profile \
--model Qwen/Qwen3-0.6B \
--endpoint-type chat \
--streaming \
--url localhost:8000 \
--public-dataset sharegpt \
--request-count 20 \
--concurrency 4
# Replay a Mooncake trace on its original timing
aiperf profile \
--model Qwen/Qwen3-0.6B \
--endpoint-type chat \
--streaming \
--url localhost:8000 \
--input-file mooncake_trace.jsonl \
--custom-dataset-type mooncake_trace \
--fixed-schedule
Here --streaming is closer to required than optional. Without streaming, the server batches the whole response at once, so there are no first-token or decode events. You cannot measure TTFT or ITL.
Read the five metrics by role
Split the recurring AIPerf metrics by role and they look like this.
TTFT (Time to First Token)
Time from request to first token.
Includes queue wait + prefill + network.
Directly tied to perceived responsiveness.
ITL (Inter-Token Latency)
Average gap between tokens. Also called TPOT.
AIPerf definition: (e2e latency - TTFT) / (output tokens - 1)
Looks at the decode span only, excluding TTFT.
Request Latency
Total time from request to last token.
e2e = TTFT + generation time.
Throughput
Tokens generated per second (whole system) and requests per second.
The baseline for capacity planning.
Percentile
avg / min / max / p50 / p90 / p95 / p99 / std.
Even with a good average, a collapsed p99 means failure in production.
The key point is not one plain tokens/sec figure. In production, how p95 latency collapses as concurrency grows is the far more important signal.
The official docs make it concrete. A 1,000-request synthetic run shows TTFT averaging 347ms with p50 289ms, p90 577ms, p99 815ms, at about 22,521 tokens/sec system throughput. A trace-based run with wider input-length spread shows a totally different picture: TTFT averaging 407ms, p99 951ms, throughput 4,675 tokens/sec. That is why a server that passes uniform synthetic tests can collapse under real traffic.
The concurrency-sweep example is textbook.
concurrency=1: TTFT p50=25ms, throughput=15 tok/s
concurrency=4: TTFT p50=35ms, throughput=55 tok/s
concurrency=8: TTFT p50=50ms, throughput=95 tok/s
concurrency=16: TTFT p50=120ms, throughput=130 tok/s ← sweet spot
concurrency=32: TTFT p50=350ms, throughput=140 tok/s ← stalling
concurrency=64: TTFT p50=900ms, throughput=135 tok/s ← saturated
Do not keep raising concurrency just because throughput still climbs. The sweet spot is where TTFT p95 starts to bend.
For reasoning models, separate TTFT from TTFO
People coming from GenAI-Perf trip on one thing: reasoning-token handling. GenAI-Perf ignored reasoning_content and marked TTFT at the first normal output token. AIPerf parses reasoning tokens too, marks TTFT at the first token of any kind, and tracks the first normal output token separately as TTFO.
GenAI-Perf TTFT = AIPerf TTFO
AIPerf TTFT ≤ AIPerf TTFO (on reasoning models)
So when you compare against old numbers, pull AIPerf's TTFO for the same definition. On reasoning models, watch both: users feel the wait until the actual answer, not the thinking tokens.
CodeBridge Mini Lab: measure all five together
If you run a local model or an inference server, start here instead of average latency.
1. TTFT — how long until the first token
2. ITL — does streaming stall
3. Request latency — how long until the end
4. Throughput — how much at once
5. Concurrency — at what concurrency p95 breaks
aiperf plot shows TTFT trends, ITL trends, throughput, and GPU usage together, so you can see which bottleneck hits first. With DCGM or pynvml, GPU telemetry joins in. This is exactly the picture behind the claim that the AI engineer role now spans evaluation plus inference plus observability plus cost engineering.
Further reading
- Why benchmark scores alone mislead you
- Judge AI model cost per successful task
- How to cut AI API cost with prompt caching
References
- NVIDIA Technical Blog: Benchmarking LLM Inference at Scale with AIPerf
- NVIDIA AIPerf Documentation
- NVIDIA AIPerf Metrics Reference
- NVIDIA: Migrating from GenAI-Perf
Go deeper with a course
If you want practice in choosing architectures by token cost and response speed — not just accuracy — the design path from classic RAG to GraphRAG and agentic RAG continues this measurement story directly.