AI This article was created with the help of AI.

The Leaderboard Illusion: Why Engine Benchmarks Disagree

Engineering teams evaluating LLM serving infrastructure routinely encounter contradictory throughput rankings across tech blogs and documentation. One benchmark shows SGLang dominating throughput, another crowns vLLM as the overall leader, and vendor releases claim TensorRT-LLM achieves untouchable token generation speeds. These numbers are rarely fabricated, but they are almost always measured under hyper-specific runtime conditions that do not reflect production traffic.

The root cause of these discrepancies lies in benchmark configurations that isolate specific execution phases. For instance, the original SGLang paper, submitted in December 2023, reports up to 6.4x higher throughput and up to 3.7x lower latency than the state-of-the-art inference systems of that period across tasks including agent control, few-shot learning, JSON decoding, retrieval-augmented generation and multi-turn chat, measured against vLLM v0.2.5 on 24 GB NVIDIA A10G GPUs at float16. That headline multiple comes from the workloads where prefix reuse was maximal, not from generic single-turn serving. Conversely, vLLM's v0.6.0 performance update recorded a 2.7x throughput improvement and a 5x reduction in time per output token (TPOT) on Llama 3 8B compared to its own v0.5.3 release, measured on NVIDIA H100 GPUs using the ShareGPT dataset. Meanwhile, NVIDIA publishes no direct head-to-head framework comparison against vLLM or SGLang on its primary documentation channels.

Every performance metric shifts radically depending on the chosen model architecture, GPU generation, quantization format, concurrency level, and prompt-to-generation ratio. We explored the two-way comparison in our vLLM vs TensorRT-LLM production benchmark, but adding a third engine makes it clear that chasing peak tokens per second on a static leaderboard is an architectural dead end.

  • Prompt-to-output ratios: Prefix caching engines dominate prefill-heavy workloads but offer no advantage during long autoregressive decoding loops.
  • Batching and scheduling: Multi-step scheduling and continuous batching yield different latency curves depending on request concurrency.
  • Kernel maturity: Point releases and compiler updates routinely invert benchmark rankings across hardware architectures.

vLLM: The Safe Default for High Model Churn

vLLM has established itself as the standard execution engine for teams that prioritize agility and broad architectural support. When an open-weight model drops on Hugging Face, vLLM typically provides functional serving support within hours. For engineering teams managing a dynamic catalogue of checkpoints, this velocity outweighs marginal latency differences.

The engine relies on PagedAttention to eliminate memory fragmentation in the Key-Value (KV) cache, enabling continuous batching that dynamically injects new requests into running iterations without waiting for prior sequences to complete. In modern deployments following our vLLM production deployment patterns, automatic prefix caching (APC) and chunked prefill operate out of the box to stabilize memory allocation and prevent compute spikes.

vLLM is the default choice when model churn is high or when infrastructure must support diverse open-source architectures without custom compilation pipelines. Its developer experience minimizes engineering overhead, allowing teams to deploy reliable endpoints with minimal operational friction.

  • Rapid model support: Immediate compatibility with hundreds of open-source architectures and custom fine-tunes.
  • PagedAttention architecture: Dynamic memory allocation that prevents out-of-memory errors caused by KV cache fragmentation.
  • Operational simplicity: Native Python integration and standardized CLI flags that simplify cluster rollout.

SGLang: Optimising for Repeated Prefixes and Prefill

SGLang takes a specialized approach to memory management by treating the KV cache as an execution tree rather than an ephemeral buffer. Built around RadixAttention, the engine maintains an LRU cache of intermediate KV states organized in a radix tree. This design enables automatic prefix caching across independent incoming requests without requiring manual session management.

This architecture delivers massive throughput gains for workloads that exhibit structural prefix sharing: multi-turn agent loops, few-shot prompting pipelines, structured JSON decoding, and complex Retrieval-Augmented Generation (RAG) systems. By matching the longest common prefix in the radix tree, SGLang bypasses the prefill phase for cached prompt segments, significantly reducing time to first token (TTFT).

However, prefix caching is not a universal performance lever. As vLLM's official documentation states, automatic prefix caching only reduces the time of processing the queries (the prefill phase) and does not reduce the time of generating new tokens during decode. Workloads dominated by short system prompts and long generation lengths see minimal benefit from RadixAttention.

The workload shape decides how much of this matters. Multi-turn agent loops re-send the whole conversation plus the latest tool output on every call, so almost the entire prompt is a cached prefix and the remaining work is the new suffix. Structured JSON decoding shares a common grammar and schema prefix across requests, with constrained token masking rather than prefill as the limiting factor. Long-form document Q&A caches a static context but pays for a long decode on detailed answers. And a short prompt with a long generation has almost no prefill to cache at all, leaving autoregressive memory bandwidth as the bottleneck, which is exactly the case where RadixAttention earns nothing.

TensorRT-LLM: Maximum Tuning for Frozen Configurations

TensorRT-LLM is NVIDIA's comprehensive open-source library for accelerating and optimising the inference performance of the latest large language models on NVIDIA GPUs. It incorporates advanced runtime primitives, including in-flight batching, paged KV caching with block reuse, chunked prefill, and native support for speculative decoding algorithms such as EAGLE and Medusa.

Historically, TensorRT-LLM required complex ahead-of-time (AOT) engine builds that produced rigid binary artifacts tied to specific tensor parallelism and batch dimensions. As of Release 1.2, NVIDIA removed the legacy TensorRT engine-compilation backend, establishing PyTorch as the sole execution backend. While this improves developer ergonomics, deploying TensorRT-LLM still demands extensive per-model, per-GPU, and per-precision kernel tuning to achieve optimal throughput.

The operational trade-off is clear: TensorRT-LLM excels when serving a stable, unchanging model on dedicated NVIDIA silicon where engineers can invest time in custom FP8 or FP4 quantization and kernel profiling. If your architecture changes frequently, the operational cost of managing custom build pipelines quickly overtakes the runtime gains.

  • Hardware-level optimization: Deep integration with Hopper and Blackwell architectures, including FP8 and FP4 execution paths.
  • In-flight batching: Dynamic request scheduling that maximizes SM occupancy across mixed prefill and decode phases.
  • Configuration rigidity: Tuning effort is tightly coupled to specific GPU counts, precision formats, and model weights.

The Engine Selection Rule: A Reproducible Decision Procedure

Choosing the right serving engine should not be based on vendor leaderboards. Instead, platform teams must evaluate two structural variables: model churn frequency and prefix sharing density. By mapping your workload across these dimensions, the optimal engine choice becomes clear.

Workload CharacteristicRecommended EngineKey Operational Trade-Off
High model churn, broad open-source cataloguevLLMSlightly lower theoretical peak throughput on frozen dense models
High prefix reuse, agent loops, multi-turn chatSGLangRequires prefill-heavy traffic to realize radix caching benefits
Frozen production model, dedicated NVIDIA siliconTensorRT-LLMRequires specialized configuration and deep hardware profiling

Before committing to an engine, run an internal benchmark protocol using your production request distribution rather than synthetic uniform sequences. Measure time-to-first-token (TTFT) and inter-token latency (ITL) under realistic concurrency.

  1. Record a representative sample of real production prompts to capture authentic prefix distributions and prompt-to-generation ratios.
  2. Replay traffic against candidate engines deployed on identical GPU hardware with matched parallelism settings.
  3. Sweep concurrency from a single stream up to the level your service actually runs at, logging p95 TTFT, p99 ITL, and total tokens per second.
  4. Evaluate operational overhead, including cold-start duration, driver compatibility, and failure recovery mechanics.

NVIDIA Dynamo: The Orchestrator, Not a Fourth Engine

A common category error in AI infrastructure discussions is evaluating NVIDIA Dynamo as a direct competitor to vLLM, SGLang, or TensorRT-LLM. Dynamo is not an inference execution engine; it is a distributed orchestration layer that coordinates multi-node inference across heterogeneous clusters.

As NVIDIA documentation specifies, Dynamo does not replace underlying execution engines like TensorRT-LLM or vLLM; it transforms them into a coordinated, multi-node inference fabric. Dynamo manages global request scheduling, disaggregates prefill and decode stages across separate worker pools, and handles cross-node tensor routing.

  • Cluster-wide scheduling: Balances incoming request streams across multiple GPU worker nodes dynamically.
  • Prefill and decode disaggregation: Separates compute-heavy prompt evaluation from memory-bound autoregressive decoding.
  • Engine integration: Operates above the execution backend, allowing platform teams to leverage optimized engines underneath.

Declining the Engine Choice: Serverless and Dedicated Options

For most platform and application teams, managing bare-metal GPU runtimes, CUDA kernel updates, and continuous batching schedulers introduces unnecessary operational friction. Maintaining custom serving infrastructure diverts senior engineering time away from building differentiated product features. The most pragmatic architecture is often to decline the engine choice entirely and rely on managed inference infrastructure.

At Lyceum, we build our Serverless Inference on an open inference stack combining vLLM, NVIDIA Dynamo, and TensorRT-LLM. This open-stack foundation contrasts with the proprietary engines used by providers like Together and Fireworks. Our serverless endpoints are OpenAI SDK compatible, with per-token pricing and no data retention, so developers switch models by changing the base URL, the API key and the model string and keep the rest of their code unchanged.

When workloads demand isolated compute or custom container environments, our Dedicated Inference provides private endpoints with automated scaling and scale-to-zero capabilities across European data centres in Paris and Finland. For engineering teams that genuinely require full runtime control, our On-demand GPU VM instances deliver raw SSH access to NVIDIA L40S (48 GB), A100 (80 GB), and H100 (80 GB) instances with per-second billing and 18-second provisioning.

  • Serverless Inference: Pay-per-token access to open-weight models via an OpenAI-compatible API without managing infrastructure.
  • Dedicated Inference: Isolated endpoints running custom Hugging Face weights or Docker containers on dedicated GPUs.
  • On-demand GPU VM: Raw GPU instances with NVLink interconnects for teams executing custom kernel profiling.