What tokens per second and TTFT actually measure

Two numbers dominate any inference benchmark, and they measure different things. Time to first token (TTFT) is the wall-clock wait between submitting a request and receiving the first streamed token; our companion piece, Time to First Token: What Actually Determines LLM API Latency, breaks it into its four components. It bundles three costs: network round trip, queueing, and prompt prefill, where the model builds its KV cache from the input sequence before it can emit anything. Tokens per second is the opposite end of the request: how fast the model streams text after that first token arrives. NVIDIA's benchmarking documentation separates the two deliberately, defining inter-token latency (ITL) as the average time between consecutive tokens, and total tokens per second as total output tokens divided by the end-to-end latency between the first request and the last response.

The distinction matters because users do not experience them equally. A chat interface reads like a page: a human reader consumes text at roughly 238 words per minute, or about 4 to 5 tokens per second, according to a 2019 meta-analysis of 190 studies by Marc Brysbaert at Ghent University. Once streaming throughput clears that reading rate, extra tokens per second are invisible to the person waiting. What they feel is the blank screen before the first word. For an AI-native product, TTFT is the latency your customer sees; tokens per second is the cost and capacity number you carry internally.

  • TTFT: request submission to first streamed token.
  • ITL / TPOT: average gap between consecutive tokens after generation starts. NVIDIA's AIPerf excludes TTFT from this average, using (request latency minus TTFT) divided by (output tokens minus 1).
  • Tokens per second: output tokens divided by decode time. A batch-level throughput figure in most tools, not a live per-request metric.
  • What the user perceives: TTFT as application lag, tokens per second as whether the stream keeps pace with reading speed, roughly 4-5 tokens per second for a human.

How these numbers were measured

These are our own measurements, captured from our live status probe against our public API, not a compilation of figures from other providers' marketing pages. The probe calls the public API as an ordinary external client every 10 minutes, one streaming chat completion per model, with eight models running in parallel and a single stream each. This page publishes the 7-day median ending 24 Sep 2026, with the 10th to 90th percentile alongside it. The live status page at status.lyceum.technology re-measures every model on the same schedule and carries per-model latency and throughput history.

  • Prompt: a short instruction asking the model to list the integers from 1 to 200 separated by single spaces, with nothing else in the output.
  • max_tokens: 600. temperature: 0.
  • One streaming chat completion per model per probe, no concurrent load applied. Eight models are probed in parallel, but each on its own stream.
  • TTFT is wall-clock time to the first streamed token of any kind, reasoning tokens included, so it carries the full network round trip.
  • Throughput is generated tokens divided by decode time, with TTFT excluded from the denominator.
  • Window: 7 days ending 24 Sep 2026. Median with 10th to 90th percentile in brackets. Nothing rounded beyond what the probe reports.

What moves throughput and first-token latency

Prompt length is the first lever on TTFT. The attention mechanism has to build a KV cache across the full input sequence before generation can begin, so a 100-token prompt prefills in a fraction of the time a long prompt takes, before any model-specific speed enters the picture. Our probe prompt is deliberately short, so the TTFT figures below are close to the floor for each model on our infrastructure. A RAG pipeline injecting five retrieved documents into the prompt will sit well above them.

After prompt length, the serving stack's memory management becomes the dominant factor. vLLM's PagedAttention treats the KV cache the way an operating system treats virtual memory: it pages it into fixed-size blocks instead of reserving one large contiguous region per request. This near-eliminates the fragmentation and over-reservation that waste KV cache memory in naive serving, and lets many more requests share the same GPU at once. The vLLM paper reports a 2-4× throughput improvement over earlier state-of-the-art serving systems such as FasterTransformer and Orca at the same level of latency, precisely because tighter KV-cache packing supports larger batch sizes without OOM.

That capacity is the trade-off to watch. When concurrency rises and the batch grows, all the sequences in the batch contend for the same memory bandwidth on every decode step. Inter-token latency degrades first, even while TTFT stays reasonable, because each token's generation now waits behind more work on the same hardware. On a shared, multi-tenant endpoint, your traffic is not the only load on the GPU at any given moment. On an isolated, dedicated endpoint, it is.

Reading the per-model numbers for your workload

The table below is the measurement itself: 16 models, probed on identical conditions (the short instruction prompt above, max_tokens 600, temperature 0, single stream), 7-day median ending 24 Sep 2026, 10th to 90th percentile in brackets. Every model string is exactly what you would pass in the model field of an OpenAI-compatible request. Nothing is rounded beyond what the probe reports.

Model stringOutput tokens/sTTFT ms
qwen/qwen3.8-2.4t-a95b97 (87-103)906 (316-1827)
moonshotai/kimi-k2.7-code90 (59-99)714 (340-2785)
deepseek/deepseek-v4.1-flash90 (87-95)252 (219-355)
deepseek/deepseek-v4-flash-073185 (74-87)245 (208-371)
minimax/minimax-m382 (46-95)892 (454-1995)
qwen/qwen3.5-9b76 (50-97)1429 (802-3037)
z-ai/glm-5.172 (27-75)1458 (514-3449)
deepseek/deepseek-v4-pro66 (48-73)287 (237-3762)
moonshotai/kimi-k364 (57-66)315 (256-691)
qwen/qwen3.8-flash-next61 (46-108)1046 (426-3926)
z-ai/glm-5.251 (46-55)439 (352-1002)
z-ai/glm-5.2-instant48 (45-53)413 (346-1002)
z-ai/glm-5.347 (41-52)1164 (328-5128)
z-ai/glm-5.3-flash45 (29-53)292 (226-665)
qwen/qwen3.8-27b44 (40-46)968 (440-1834)
moonshotai/kimi-k2.640 (27-44)2138 (876-3742)

These are single-stream figures, one request at a time, so they are a best case, not a guarantee under load. They tell you what each model does when it has the GPU's decode bandwidth to itself against a short prompt, which is useful for comparing models against each other on identical conditions. They do not tell you what happens to your p99 TTFT or your throughput when 50 concurrent requests arrive from real product traffic, which is the next section's point.

What these measurements do not tell you

  • No concurrency curve. As noted above, one stream per model provides no data on degradation under parallel load.
  • No cold-start behaviour. These probes hit models that are warm. First-request latency after a scale-from-zero event is a different measurement, and it is not captured here.
  • Not a cross-provider benchmark. This page reports measurements from our own infrastructure only. No independent third party has benchmarked our infrastructure directly; sites like artificialanalysis.ai publish cross-provider comparisons from their own test harnesses, and their verdicts are their own. A cross-provider number is only comparable under identical prompt, max_tokens, temperature, and concurrency.
  • Not a saturated-cluster figure. This probe tracks the public API, not an isolated or dedicated endpoint under load. A saturated private cluster behaves differently, in ways this snapshot cannot show.

Reproducing the benchmark on your own traffic

The probe is not exotic. It is a streaming OpenAI-compatible chat completion, timed with a wall clock, run on a schedule. You can replicate it against your own prompt shapes and get numbers that mean something for your workload rather than for ours.

  1. Point an OpenAI-compatible client at the public API (https://api.lyceum.technology/openai/v1) with any key that has access, and set model to one of the strings from the table above.
  2. Use your own production prompt shape, not our test prompt. Prompt length moves TTFT, so a RAG-shaped prompt will sit above the numbers in the table, and that is the number you actually need.
  3. Set temperature 0 if you want runs comparable to this page, or your production temperature if you want the number your traffic produces.
  4. Capture the wall-clock time when you send the request and the wall-clock time when the first streamed chunk of any kind arrives. The difference is your TTFT.
  5. Sum the output tokens from the stream and record the wall-clock time from the first token to the last. Divide output tokens by that decode duration to get your tokens per second, excluding TTFT from the denominator, the same way ours is.
  6. Run this on a schedule, not once. Our status page re-probes every model every 10 minutes, so a single snapshot can catch a transient. A rolling median over hours gives you the same percentile band ours does, and the same early warning when a figure drifts.
  • Size your own workload against the model you plan to run, then test it on your traffic. The table above is the baseline, your prompt shape is the correction, and an isolated endpoint takes other tenants' load out of the picture. For the economics of that step, see Pay Per Token vs Dedicated GPU Inference.
  • A throughput or TTFT figure without its prompt, output length, temperature, and concurrency beside it is not a measurement.