AI This article was created with the help of AI.

What time to first token actually measures

When building interactive AI applications or low-latency agentic systems, the perceived responsiveness of your product depends entirely on how quickly the model starts generating output. Time to First Token (TTFT) is the definitive metric for this initial interaction. It measures the total wall-clock time elapsed from the millisecond your application dispatches an HTTP request to the moment it receives the very first output token from the server. If your time to first token is elevated, users experience a jarring freeze, and downstream systems like text-to-speech pipelines or tool calling loops remain blocked.

The four stages of request delivery

Many developers treat TTFT as a monolithic hardware execution time, but it represents the cumulative sum of multiple distinct systems operations. In modern AI infrastructure, a request traverses four specific phases before yielding its initial token:

  • Network and API overhead: The duration required for your client payload to serialize, traverse public routing or private interconnects, pass TLS termination, and be parsed by the API gateway.
  • Queueing delay: The idle waiting time a request spends inside the inference engine's scheduling backlog while waiting for available GPU memory or execution slots.
  • Prompt prefilling: The compute-intensive phase where the transformer model computes attention across all input tokens in parallel and populates the Key-Value (KV) cache.
  • Initial token decoding: The first autoregressive decode step where the model computes logits over the vocabulary, samples the first token ID, and de-tokenizes the result.

Understanding this separation is essential for isolating performance bottlenecks. While end-to-end latency measures the total time required to generate an entire completion, and inter-token latency (ITL, also called time per output token) measures the average gap between consecutive tokens during generation, TTFT defines the initial barrier to entry. Even if an inference engine achieves high throughput across thousands of tokens, a sluggish prefill or extended queue time will make the API feel unresponsive.

The hidden cost of queueing delay

In production environments handling bursty traffic, queueing delay is frequently the largest single contributor to tail TTFT spikes. When multiple clients dispatch requests simultaneously, the inference engine cannot allocate execution resources instantly without exceeding hardware limits. If the engine lacks immediate capacity, new requests sit in a scheduler backlog, accumulating latency before any matrix multiplication begins.

Continuous batching vs decode starvation

Modern inference engines use continuous batching (iteration-level scheduling) to maximize GPU compute utilization. Rather than waiting for an entire batch to finish generating, new requests are injected dynamically into active execution loops. However, combining compute-heavy prefill operations with memory-bandwidth-bound decode steps introduces fundamental scheduling tensions.

To prevent decoding operations from stalling, traditional schedulers often prioritize active decode tasks. Under bursty workloads, this decode-prioritizing policy can cause significant prefill delays, inflating TTFT while maintaining stable token generation rates. Research into inference scheduling, such as the FairBatching framework, demonstrates that rigidly prioritizing decoding tasks to protect Time-Per-Output-Token (TPOT) targets leaves decode slack unused while prefill requests queue, and that reclaiming that slack for prefill surges reduces TTFT tail latency by up to 2.29x while still holding TPOT service-level objectives. Managing queue depth and using balanced chunked prefill scheduling are critical for keeping queue times minimal.

The three common scheduling policies trade TTFT against decode smoothness in different ways. Prefill-prioritizing schedulers admit new requests as soon as they arrive, which keeps queue time short but pauses in-flight decode steps whenever a large prompt is being processed, so token generation becomes jittery. Stall-free decode-prioritizing schedulers do the opposite: decode steps are guaranteed compute slots in every batch and the per-token rate stays steady, but prefill work is deferred and tail TTFT climbs during concurrency spikes. Fair dynamic batching breaks with that decode-first paradigm by adapting batch capacity to the available slack, reclaiming compute from bursting decode tasks to serve prefill surges so that both TTFT and per-token latency stay predictable.

Cold starts versus warm inference

In serverless or dynamically scaled GPU environments, the physical state of the underlying instance determines whether a request executes immediately or incurs an initialization penalty. A warm inference request hits an engine that already holds the model weights in GPU High Bandwidth Memory (HBM) and maintains an active CUDA context. In contrast, a cold start occurs when an engine must spin up a container, allocate VRAM, and stream tens or hundreds of gigabytes of weights from storage.

Storage bottlenecks and weights-to-VRAM streaming

Model weight loading is fundamentally an I/O bottleneck governed by the interconnect between storage tiers, host system RAM, and GPU memory over PCIe buses. In standard architectures, serialized weights are read from network-attached storage or NVMe drives into CPU memory before being copied across PCIe lanes to GPU VRAM.

Optimizing storage throughput and overlapping tensor deserialization drastically compresses cold start duration. For example, NVIDIA benchmarks evaluate that using concurrent weight loaders on high-throughput IO2 SSD storage allows a 15 GB model to load into GPU memory in 7.53 seconds at concurrency 8, compared to 47 seconds using standard single-threaded loaders. Understanding your provider's weight-caching and container lifecycle policies is vital to preventing cold start latency from degrading baseline TTFT.

Prompt prefilling and the KV cache

Once a request clears the queue and lands on active GPU compute, TTFT enters the prompt prefill phase. Unlike autoregressive decoding, which generates one token at a time in memory-bandwidth-bound operations, prefilling is a compute-bound operation that processes the entire prompt in parallel. The self-attention mechanism computes pairwise token interactions, scaling quadratically with input sequence length.

Prefix caching and prompt structure

During prefilling, the engine calculates key and value vector representations for every prompt token and writes them into the KV cache. If your application sends large system prompts, extensive few-shot examples, or long retrieval documents, computing this attention from scratch on every call creates substantial inference latency.

Prefix caching mitigates this overhead by retaining immutable KV blocks in GPU memory across requests that share identical prompt prefixes. When a request matches an existing prefix, the engine skips attention recalculation for those tokens and executes prefill only for the delta. In production setups, monitoring systems track this using the engine's own counters: vLLM's production metrics endpoint exposes prefix cache query and prefix cache hit counters, both measured in tokens, so teams can verify that common prompt templates achieve high cache reuse.

  • Place static content first: Keep system prompts, schema definitions, and constant tool descriptions at the beginning of the prompt to maximize cache block matching.
  • Isolate dynamic variables: Append variable data, such as user inputs and dynamic timestamps, strictly at the end of the prompt sequence.
  • Align prompt boundaries: Standardize whitespace and token boundaries so the tokenizer produces consistent sub-word IDs across queries.

Token streaming over Server-Sent Events

For user-facing products, waiting for an entire multi-paragraph completion before returning a response destroys the interactive experience. Token streaming over Server-Sent Events (SSE) allows the server to push each generated token to the client as an independent HTTP chunk the instant it is sampled, effectively masking end-to-end generation latency behind a fast TTFT.

Chunked transfer and de-tokenization

In streaming mode, the inference engine performs an iterative pipeline: computing logits for the current token, sampling a token ID, executing a de-tokenization step to convert the ID into text, and framing the payload as an SSE data event. As NVIDIA's LLM benchmarking guidance notes, in streaming mode the de-tokenization step can be performed multiple times as partial results are returned to the user, so text arrives incrementally rather than waiting for complete sequence termination.

When client applications consume streamed chunks, UI layers can render text token-by-token or buffer small windows for smoother rendering. For downstream agent systems, streaming enables early speculative processing, allowing validation routines or tool routing logic to begin while the model finishes generating the remaining payload.

Model architecture and smart routing

Beyond infrastructure scheduling and storage bottlenecks, model architecture dictates the baseline compute cost of the prefill stage. A model's total parameter count, attention architecture, and activation structure set a physical floor on how fast it can process input sequences.

Isolating routing overhead from compute latency

Dense architectures activate every parameter in the network for each token, demanding significant floating-point operations during prefilling. In contrast, sparse Mixture-of-Experts (MoE) architectures route tokens to a subset of specialized experts, reducing the active parameter count per token and yielding faster prefill execution despite large overall model weights on disk.

  • Dense models: Require uniform computation across all parameters, leading to higher prefill times as parameter counts scale.
  • Sparse MoE models: Activate only a fraction of total parameters per token, enabling lower prefill latency and reduced TTFT for equivalent model capability.
  • Smart routing layers: Lightweight classification routers analyze prompt intent and route traffic dynamically, adding minimal routing latency while directing queries to the most cost-effective model.

When evaluating API responsiveness, distinguish between routing network hops and engine-level compute bottlenecks. A well-designed routing proxy adds negligible overhead while ensuring that simple prompts route to compact, fast-prefill models and complex reasoning tasks land on high-capacity architectures.

Managing latency with Serverless Inference

For AI-native product companies, managing the infrastructure stack required for low TTFT (optimizing continuous batching schedulers, maintaining warm GPU pools, configuring prefix caches, and tuning network gateways) introduces significant engineering overhead. Lyceum provides Serverless Inference to deliver access to pre-hosted open-source models without the burden of provisioning or managing physical GPU clusters.

Built on high-performance open stacks like vLLM and TensorRT-LLM with hosting in the European Union (eu-north1), Serverless Inference eliminates infrastructure provisioning delays and operates on per-token pricing with zero base fees. Developers can monitor live per-model latency and throughput metrics directly on the public status page at https://status.lyceum.technology, noting that self-serve serverless endpoints operate without formal availability tiers or service-level agreements.

  • Measure TTFT precisely: Isolate network latency, queueing delay, and prefill compute in your telemetry to identify true bottlenecks.
  • Structure prompts for caching: Keep static system instructions identical across calls to leverage prefix caching.
  • Stream over SSE: Always enable token streaming in user-facing endpoints to deliver sub-second responsiveness.

By understanding the mechanics governing queueing, prefill computation, and streaming delivery, engineering teams can optimize their prompt architectures and select the right serving configurations to keep TTFT consistently low.