Where each model runs, and why it is per model
Picking an inference endpoint comes down to three questions: which open model strings can I call, where does each one run, and what does it cost per token. Most providers answer the first and leave you to infer the other two. This page answers all three for one endpoint, as the roster stood on 10 September 2026, and includes the strings that stopped answering this week and what to call instead.
Residency on this endpoint is a per-model fact. It comes from each model's own record, never from a platform-wide statement. The roster itself is machine-readable: the documentation lists GET /models as the call that returns the models currently available on the OpenAI-compatible base URL, so nothing on this page has to be taken on trust.
One hosting statement sits outside the table below: the launch newsletter, issue 008 of 3 Sep 2026, says z-ai/glm-5.3 and z-ai/glm-5.3-flash run on "our own B200s in Paris". That sentence is the published claim, attributed to the newsletter, and this page adds nothing to it.
- The model strings whose own record publishes a hosting location, and what each record states
- The published prices and context windows, per string
- The strings that stopped answering between 3 and 9 September 2026, with a replacement for each
- A rerunnable roster check, so you can verify every claim here against the live API
The model strings with a published hosting location
A row appears here only because that model's own record publishes a hosting location. These are the strings whose record states where they run; no other string on this page carries one. For a data-residency question, the answer has to come from the model's record, and this table is that answer, verbatim.
| model string | hosting as the model's own record states it |
|---|---|
| deepseek/deepseek-v4-flash-0731 | eu-north1, European Union |
| z-ai/glm-5.1 | eu-north1, European Union |
| z-ai/glm-5.2 | eu-north1, European Union |
| z-ai/glm-5.2-instant | eu-north1, European Union |
| moonshotai/kimi-k2.7-code | eu-north1, European Union |
| moonshotai/kimi-k3 | eu-north1, European Union |
| minimax/minimax-m3 | eu-north1, European Union |
| qwen/qwen3-235b-a22b-instruct-2507 | eu-north1, European Union |
| qwen/qwen3-embedding-8b | eu-north1, European Union |
| qwen/qwen3.5-9b | eu-north1, European Union |
The whole table sits behind one OpenAI-compatible endpoint and one key, available through early access. The documentation puts it plainly: point any OpenAI client at the base URL, swap in an lk_ API key and a model name, and everything else, streaming, tool calling and usage accounting, works the same.
Published prices and context windows
Prices are per 1M tokens and publish per string, alongside the context window. The table includes strings for which the approved offering records used for this article supply a context window and both ends of the price. Examples include qwen/qwen3.5-9b at $0.15 input and $0.20 output, moonshotai/kimi-k2.7-code at $1.25 input and $4.50 output, and z-ai/glm-5.2-instant at $1.50 input and $4.50 output:
| model string | context | input /1M | output /1M |
|---|---|---|---|
| qwen/qwen3.5-9b | 256K | $0.15 | $0.20 |
| z-ai/glm-5.3-flash | 1M | $0.20 | $0.50 |
| qwen/qwen3.8-flash-next | 256K | $0.20 | $0.50 |
| minimax/minimax-m3 | 1M | $0.40 | $2.00 |
| qwen/qwen3.8-27b | 256K | $0.40 | $2.40 |
| z-ai/glm-5.2-instant | 1M | $1.50 | $4.50 |
| z-ai/glm-5.3 | 1M | $1.40 | $4.40 |
| moonshotai/kimi-k2.7-code | 256K | $1.25 | $4.50 |
| qwen/qwen3.8-2.4t-a95b | 256K | $2.50 | $6.00 |
| moonshotai/kimi-k3 | 1M | $3.00 | $15.00 |
Prompt caching is automatic: when consecutive requests share an identical leading prefix, the cached portion can be billed at the reduced input rate published for that string. The selected strings below have complete cached-input figures in the source data:
| model string | cached input /1M |
|---|---|
| z-ai/glm-5.3 | $0.26 |
| z-ai/glm-5.3-flash | $0.05 |
| qwen/qwen3.8-2.4t-a95b | $0.63 |
| qwen/qwen3.8-27b | $0.10 |
| qwen/qwen3.8-flash-next | $0.05 |
The live roster also includes the following strings. They are outside the complete price-and-context table above because the approved offering records used for this article omit at least one of those fields:
- deepseek/deepseek-v4-flash-0731
- deepseek/deepseek-v4-pro
- z-ai/glm-5.1
- z-ai/glm-5.2
- z-ai/glm-5.3-flash-instant
- z-ai/glm-5.3-instant
- moonshotai/kimi-k2.6
- minimax/minimax-m2.5
- qwen/qwen3-235b-a22b-instruct-2507
- qwen/qwen3-embedding-8b
- qwen/qwen3.8-27b-instant
- qwen/qwen3.8-flash-next-instant
What changed on the roster, and what to call instead
The following strings were absent from GET /models on 10 September 2026. The date records the first absence in the source snapshots. Each replacement is a configuration suggestion with its matching basis shown, not a capability recommendation.
| model string | no longer available since | switch to | basis |
|---|---|---|---|
| nousresearch/hermes-4-70b | 3 September 2026 | z-ai/glm-5.3-flash | nearest published output price: $0.40 against $0.50 per 1M |
| qwen/qwen2.5-vl-72b-instruct | 3 September 2026 | deepseek/deepseek-v4-flash-0731 | the only live model whose record documents image input |
| qwen/qwen3-32b | 3 September 2026 | qwen/qwen3.5-9b | nearest published output price: $0.30 against $0.20 per 1M |
| nvidia/cosmos3-super-reasoner | 4 September 2026 | qwen/qwen3.5-9b | nearest published output price: $0.30 against $0.20 per 1M |
| nvidia/llama-3_1-nemotron-ultra-253b-v1 | 4 September 2026 | minimax/minimax-m3 | nearest published output price: $1.80 against $2.00 per 1M |
| nvidia/nemotron-3-nano-omni | 4 September 2026 | qwen/qwen3.5-9b | nearest published output price: $0.24 against $0.20 per 1M |
| qwen/qwen3-next-80b-a3b-thinking | 4 September 2026 | z-ai/glm-5.2-instant | no published price for the removed model; matched on the mid parameter band read off its id, not on a benchmark |
| google/gemma-3-27b-it | 9 September 2026 | qwen/qwen3.5-9b | nearest published output price: $0.30 against $0.20 per 1M |
| openai/gpt-oss-120b | 9 September 2026 | z-ai/glm-5.3-flash | nearest published output price: $0.60 against $0.50 per 1M |
| nousresearch/hermes-4-405b | 9 September 2026 | qwen/qwen3.8-27b | nearest published output price: $3.00 against $2.40 per 1M |
| meta-llama/llama-3.3-70b-instruct | 9 September 2026 | z-ai/glm-5.3-flash | nearest published output price: $0.40 against $0.50 per 1M |
| openbmb/minicpm-v-4_5 | 9 September 2026 | deepseek/deepseek-v4-flash-0731 | the only live model whose record documents image input |
| nvidia/nemotron-3-super-120b-a12b | 9 September 2026 | z-ai/glm-5.3-flash | nearest published output price: $0.90 against $0.50 per 1M |
| nvidia/nemotron-3-ultra-550b-a55b | 9 September 2026 | qwen/qwen3.8-27b | nearest published output price: $3.00 against $2.40 per 1M |
| nvidia/nvidia-nemotron-3-nano-30b-a3b | 9 September 2026 | qwen/qwen3.5-9b | nearest published output price: $0.24 against $0.20 per 1M |
| qwen/qwen3-30b-a3b-instruct-2507 | 9 September 2026 | qwen/qwen3.5-9b | nearest published output price: $0.30 against $0.20 per 1M |
| qwen/qwen3.5-397b-a17b | 9 September 2026 | z-ai/glm-5.2-instant | nearest published output price: $3.60 against $4.50 per 1M |
Most destinations are matched on the nearest published output price. Two vision entries use documented image input, and one entry uses the parameter band because the source dataset supplied no price. None is a benchmark or capability match. Test the candidate against your workload before changing production traffic.
Changing the model field in your client
The fix is a one-field change. The base URL and the API key stay as they are; only the string in model changes. First, check what the roster answers right now, with BASE_URL set to the serverless base URL shown in the client sample below:
curl https://api.lyceum.technology/openai/v1/models \
-H "Authorization: Bearer lk_..."Then change the one field in an OpenAI-SDK client:
from openai import OpenAI
client = OpenAI(
api_key="lk_...",
base_url="https://api.lyceum.technology/openai/v1", # unchanged
)
resp = client.chat.completions.create(
model="z-ai/glm-5.3-flash", # the only field that changes
messages=[{"role": "user", "content": "Explain prompt caching in one sentence."}],
)
print(resp.choices[0].message.content)If you carry several removed strings across services, apply the full mapping in one pass:
REPLACEMENTS = {
"nousresearch/hermes-4-70b": "z-ai/glm-5.3-flash",
"qwen/qwen2.5-vl-72b-instruct": "deepseek/deepseek-v4-flash-0731",
"qwen/qwen3-32b": "qwen/qwen3.5-9b",
"nvidia/cosmos3-super-reasoner": "qwen/qwen3.5-9b",
"nvidia/llama-3_1-nemotron-ultra-253b-v1": "minimax/minimax-m3",
"nvidia/nemotron-3-nano-omni": "qwen/qwen3.5-9b",
"qwen/qwen3-next-80b-a3b-thinking": "z-ai/glm-5.2-instant",
"google/gemma-3-27b-it": "qwen/qwen3.5-9b",
"openai/gpt-oss-120b": "z-ai/glm-5.3-flash",
"nousresearch/hermes-4-405b": "qwen/qwen3.8-27b",
"meta-llama/llama-3.3-70b-instruct": "z-ai/glm-5.3-flash",
"openbmb/minicpm-v-4_5": "deepseek/deepseek-v4-flash-0731",
"nvidia/nemotron-3-super-120b-a12b": "z-ai/glm-5.3-flash",
"nvidia/nemotron-3-ultra-550b-a55b": "qwen/qwen3.8-27b",
"nvidia/nvidia-nemotron-3-nano-30b-a3b": "qwen/qwen3.5-9b",
"qwen/qwen3-30b-a3b-instruct-2507": "qwen/qwen3.5-9b",
"qwen/qwen3.5-397b-a17b": "z-ai/glm-5.2-instant",
}Model IDs are case-insensitive, and where a client rejects a model name containing a slash, the slash-free variant such as z-ai-glm-5.2 is accepted instead.
Keeping your config current
Before trusting any page, including this one, verify against the live roster. The check takes one request with your lk_ key as a Bearer token:
- Send GET /models on the base URL shown above, with your lk_ key as a Bearer token.
- Read the returned model IDs and confirm the string you call is present.
- If a string is missing, treat the replacement table as a shortlist, not a capability match, and test the candidate against your workload.
- Rerun the check after any deploy that touches model configuration.
Read errors defensively. Gateway errors, a bad key, an unknown model ID or insufficient credit, come back as a detail string; upstream errors use the OpenAI-compatible error.message shape. The documentation's own test client reads detail first and falls back to error.message, which is the order to copy. An unknown model ID returns a detail message that points at GET /models, the fastest route back to the live list.
For follow-up reading, the documentation index sits at docs.lyceum.technology/llms.txt, which lists every documentation page including Serverless Inference and the coding-agent guides.
Serverless Inference: one endpoint, one key
Serverless Inference is the product behind every string on this page: pre-hosted open models behind one OpenAI-compatible endpoint, metered per token, with no GPU to provision. For a team whose product runs on inference, the operational questions are the ones this page answered: which strings answer, where each runs, what each costs, and what to change when a string goes away.
- Residency is a per-model fact, stated in each model's own record.
- The price tables include strings for which the source data supplies the fields shown.
- Replacement rows use explicit heuristics and require workload testing before a production change.
If you are evaluating endpoints for a European product, or replacing strings that stopped answering this week, start at Serverless Inference and run the roster check against your own key before you commit a single config line. A Lyceum video comparison puts these prices side by side with five other European providers and reports 120 measured API requests across Kimi K3, GLM-5.2 and Qwen3.5-9B.