Engineering
How KV cache-aware routing cut our GLM-5.3 prices
What rebuilding our serverless engine on NVIDIA Dynamo taught us about routing, and how the saved GPU time reached our prices.
TL;DR On a shared serverless API, an agent's next turn can land on a copy of the model that must read its whole history again. Our router now sends it, where it can, to a copy that already holds that history. Since 1 October 2026, GLM-5.3 costs 41% less for cached input, 20% less for input and 2% less for output.
GLM-5.3 at a glance
- Developer
- Z.ai
- Model ID
z-ai/glm-5.3- Context window
- 1M tokens
- Max output
- 65,536 tokens
Price per million tokens
- Input
- $1.40
- Cached input
- $0.26
- Output
- $4.40
- Billing
- Billed per token. No base fee.
- API
- OpenAI-compatible
- Runs in
- European data centres in Paris and Finland
- Serving engine
- NVIDIA Dynamo with KV cache-aware routing
A coding agent repeats itself. Every turn, it sends us the system prompt, the tool definitions and the conversation so far, with a small new step at the end. The model has read almost all of it before. Whether we read it again or reuse the work from the last turn comes down mostly to one decision: where the request lands.
Why a shared API reads the same prompt twice
When a model reads a prompt, it stores a compact record of every token in every attention layer: the KV cache (key-value cache). Building it is called prefill. If the next request starts with exactly the same text, a copy of the model that still holds that cache only processes what is new. A copy of GLM-5.3 runs on one server with several GPUs, so we call each copy a replica; the figures simplify and say GPU.
On a dedicated deployment, one customer's requests go to the same few replicas, so reuse comes almost for free. A shared serverless API spreads requests across several replicas. Conventional balancers choose without knowing what each replica holds: round robin takes the next one in line, least-busy the shortest queue.
If turn 12 of a session lands on a replica that never saw turn 11, that replica prefills the history again. In the example agent turn priced below, that means up to 42,000 input tokens instead of 2,000, to write 1,000 tokens of output. The waste is also hard to see: a replica recomputing a prefix just looks busy, so a least-busy balancer reads a cache miss as load.
Routing by prefix, weighed against load
When we added replicas for GLM-5.3, we put NVIDIA Dynamo, an open-source framework for serving models across many GPUs, behind our serverless API and turned on its KV cache-aware router. The router tracks which prefixes each replica holds. Our own changes sit around it: for example, our API recognises a conversation and keeps it on the same deployment.
Affinity alone fails in a different way. Sending every request to its best match would pile traffic onto a few replicas while others sat idle. So the router weighs how much prefill each replica would save against how busy it already is, and gives up some reuse when the best match is overloaded.
Time sets the other limit. A prefix stays in GPU memory only until newer work pushes it out, so how long it lasts depends on load. In our own test calls, quick follow-ups almost always found the cache, and after a pause of one to three minutes it was hit or miss. Agents that call a tool and come straight back gain the most.
Moving the first models
The router takes a paragraph to describe. Moving onto it took our inference team several weeks. Our serverless models have different architectures that need different engine versions, so the team moves them one model at a time, and GLM-5.3 was among the first.
For GLM-5.3, the model ID and the endpoint stayed the same.
What a simulation shows
To isolate the mechanism, we ran both routers on the same made-up traffic.
The same 600 requests, routed two ways
Six coding agents on four GPUs, each GPU holding two conversations.
Round robin
32% of requests hit the cache
1.00× GPU time per request
Cache-aware
99% of requests hit the cache
0.34× GPU time per request
Source: Lyceum simulation, October 2026. The grids show 16 of the 600 steps; the percentages and GPU time cover all 600, with GPU time relative to round robin.
Under cache-aware routing, each conversation stays on one GPU, and the only misses are each agent's first request. Read this as the direction of the effect, not its size in production.
The set-up is generous: there is cache space for every conversation at once, so nothing is evicted. The simulated router always picks a GPU that holds the conversation, however busy, so it never makes the trade-off above. That leaves 2 GPUs serving about twice as many requests as the other 2. And a real miss costs more the longer the prefix.
From GPU time to price
Less repeated prefill means less GPU time per request. Running GLM-5.3 this way costs us less, and the new prices pass that on. Cached input fell the most.
What GLM-5.3 costs
US dollars per million tokens, excluding VAT. Before and after 1 October 2026.
| Token type | Before | Now | Change |
|---|---|---|---|
| Input | $1.75 | $1.40 | −20% |
| Cached input | $0.44 | $0.26 | −41% |
| Output | $4.50 | $4.40 | −2% |
Source: Lyceum price list, 1 October 2026.
Each workload's saving is a blend of the three cuts, weighted by what it spends on each line, not by how much of its prompt repeats. In the example coding agent turn, resent context is the largest line, bigger than the answer, and it carries $0.0072 of the $0.0080 saving. A chat turn also reuses most of its input but spends most on output, which barely moved, so it saves the least. A retrieval-augmented generation (RAG) request is mostly new context, so its saving sits close to the input cut.
What one request costs
Three example workloads, before and now. Lower is better.
| Workload | Cached input tokens | New input tokens | Output tokens | Before | Now | Saving |
|---|---|---|---|---|---|---|
| Chat | 2,000 | 400 | 500 | $0.0038 | $0.0033 | 14% less |
| Coding agent | 40,000 | 2,000 | 1,000 | $0.0256 | $0.0176 | 31% less |
| RAG | 1,000 | 8,000 | 600 | $0.0171 | $0.0141 | 18% less |
Source: Lyceum price list, 1 October 2026. Token counts are examples, not measurements.
What we learned
In LLM serving, load balancing is a cache decision, and clients can help make it. If you build agents, keep the start of each prompt stable. Put the system prompt and tool definitions first, append new turns at the end, and keep changing values such as timestamps out of the early text.
The open problem is the case our simulation leaves out: more live conversations than cache space. That is where routing gets harder.
Thanks to the inference team for the move, and to the maintainers of NVIDIA Dynamo, whose open-source work we build on.
Run GLM-5.3 on Lyceum
Model ID z-ai/glm-5.3. $1.40 per million input tokens and $0.26 for cached input. OpenAI-compatible, billed per token with no base fee.
app.py · Install openai and set LYCEUM_API_KEY
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LYCEUM_API_KEY"],
base_url="https://api.lyceum.technology/openai/v1",
)
response = client.chat.completions.create(
model="z-ai/glm-5.3",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)app.ts · Install openai and set LYCEUM_API_KEY in Node.js
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LYCEUM_API_KEY,
baseURL: "https://api.lyceum.technology/openai/v1",
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.3",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0]?.message.content);Terminal · Set LYCEUM_API_KEY in your terminal
curl https://api.lyceum.technology/openai/v1/chat/completions \
-H "Authorization: Bearer $LYCEUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-5.3",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'