Where each model runs, and why it is per model

Picking an inference endpoint comes down to three questions: which open model strings can I call, where does each one run, and what does it cost per token. Most providers answer the first and leave you to infer the other two. This page answers all three for one endpoint, as the roster stood on 10 September 2026, and includes the strings that stopped answering this week and what to call instead.

Residency on this endpoint is a per-model fact. It comes from each model's own record, never from a platform-wide statement. The roster itself is machine-readable: the documentation lists GET /models as the call that returns the models currently available on the OpenAI-compatible base URL, so nothing on this page has to be taken on trust.

One hosting statement sits outside the table below: the launch newsletter, issue 008 of 3 Sep 2026, says z-ai/glm-5.3 and z-ai/glm-5.3-flash run on "our own B200s in Paris". That sentence is the published claim, attributed to the newsletter, and this page adds nothing to it.

  • The model strings whose own record publishes a hosting location, and what each record states
  • The published prices and context windows, per string
  • The strings that stopped answering between 3 and 9 September 2026, with a replacement for each
  • A rerunnable roster check, so you can verify every claim here against the live API

The model strings with a published hosting location

A row appears here only because that model's own record publishes a hosting location. These are the strings whose record states where they run; no other string on this page carries one. For a data-residency question, the answer has to come from the model's record, and this table is that answer, verbatim.

model stringhosting as the model's own record states it
deepseek/deepseek-v4-flash-0731eu-north1, European Union
z-ai/glm-5.1eu-north1, European Union
z-ai/glm-5.2eu-north1, European Union
z-ai/glm-5.2-instanteu-north1, European Union
moonshotai/kimi-k2.7-codeeu-north1, European Union
moonshotai/kimi-k3eu-north1, European Union
minimax/minimax-m3eu-north1, European Union
qwen/qwen3-235b-a22b-instruct-2507eu-north1, European Union
qwen/qwen3-embedding-8beu-north1, European Union
qwen/qwen3.5-9beu-north1, European Union

The whole table sits behind one OpenAI-compatible endpoint and one key, available through early access. The documentation puts it plainly: point any OpenAI client at the base URL, swap in an lk_ API key and a model name, and everything else, streaming, tool calling and usage accounting, works the same.

Published prices and context windows

Prices are per 1M tokens and publish per string, alongside the context window. The table includes strings for which the approved offering records used for this article supply a context window and both ends of the price. Examples include qwen/qwen3.5-9b at $0.15 input and $0.20 output, moonshotai/kimi-k2.7-code at $1.25 input and $4.50 output, and z-ai/glm-5.2-instant at $1.50 input and $4.50 output:

model stringcontextinput /1Moutput /1M
qwen/qwen3.5-9b256K$0.15$0.20
z-ai/glm-5.3-flash1M$0.20$0.50
qwen/qwen3.8-flash-next256K$0.20$0.50
minimax/minimax-m31M$0.40$2.00
qwen/qwen3.8-27b256K$0.40$2.40
z-ai/glm-5.2-instant1M$1.50$4.50
z-ai/glm-5.31M$1.40$4.40
moonshotai/kimi-k2.7-code256K$1.25$4.50
qwen/qwen3.8-2.4t-a95b256K$2.50$6.00
moonshotai/kimi-k31M$3.00$15.00

Prompt caching is automatic: when consecutive requests share an identical leading prefix, the cached portion can be billed at the reduced input rate published for that string. The selected strings below have complete cached-input figures in the source data:

model stringcached input /1M
z-ai/glm-5.3$0.26
z-ai/glm-5.3-flash$0.05
qwen/qwen3.8-2.4t-a95b$0.63
qwen/qwen3.8-27b$0.10
qwen/qwen3.8-flash-next$0.05

The live roster also includes the following strings. They are outside the complete price-and-context table above because the approved offering records used for this article omit at least one of those fields:

  • deepseek/deepseek-v4-flash-0731
  • deepseek/deepseek-v4-pro
  • z-ai/glm-5.1
  • z-ai/glm-5.2
  • z-ai/glm-5.3-flash-instant
  • z-ai/glm-5.3-instant
  • moonshotai/kimi-k2.6
  • minimax/minimax-m2.5
  • qwen/qwen3-235b-a22b-instruct-2507
  • qwen/qwen3-embedding-8b
  • qwen/qwen3.8-27b-instant
  • qwen/qwen3.8-flash-next-instant

What changed on the roster, and what to call instead

The following strings were absent from GET /models on 10 September 2026. The date records the first absence in the source snapshots. Each replacement is a configuration suggestion with its matching basis shown, not a capability recommendation.

model stringno longer available sinceswitch tobasis
nousresearch/hermes-4-70b3 September 2026z-ai/glm-5.3-flashnearest published output price: $0.40 against $0.50 per 1M
qwen/qwen2.5-vl-72b-instruct3 September 2026deepseek/deepseek-v4-flash-0731the only live model whose record documents image input
qwen/qwen3-32b3 September 2026qwen/qwen3.5-9bnearest published output price: $0.30 against $0.20 per 1M
nvidia/cosmos3-super-reasoner4 September 2026qwen/qwen3.5-9bnearest published output price: $0.30 against $0.20 per 1M
nvidia/llama-3_1-nemotron-ultra-253b-v14 September 2026minimax/minimax-m3nearest published output price: $1.80 against $2.00 per 1M
nvidia/nemotron-3-nano-omni4 September 2026qwen/qwen3.5-9bnearest published output price: $0.24 against $0.20 per 1M
qwen/qwen3-next-80b-a3b-thinking4 September 2026z-ai/glm-5.2-instantno published price for the removed model; matched on the mid parameter band read off its id, not on a benchmark
google/gemma-3-27b-it9 September 2026qwen/qwen3.5-9bnearest published output price: $0.30 against $0.20 per 1M
openai/gpt-oss-120b9 September 2026z-ai/glm-5.3-flashnearest published output price: $0.60 against $0.50 per 1M
nousresearch/hermes-4-405b9 September 2026qwen/qwen3.8-27bnearest published output price: $3.00 against $2.40 per 1M
meta-llama/llama-3.3-70b-instruct9 September 2026z-ai/glm-5.3-flashnearest published output price: $0.40 against $0.50 per 1M
openbmb/minicpm-v-4_59 September 2026deepseek/deepseek-v4-flash-0731the only live model whose record documents image input
nvidia/nemotron-3-super-120b-a12b9 September 2026z-ai/glm-5.3-flashnearest published output price: $0.90 against $0.50 per 1M
nvidia/nemotron-3-ultra-550b-a55b9 September 2026qwen/qwen3.8-27bnearest published output price: $3.00 against $2.40 per 1M
nvidia/nvidia-nemotron-3-nano-30b-a3b9 September 2026qwen/qwen3.5-9bnearest published output price: $0.24 against $0.20 per 1M
qwen/qwen3-30b-a3b-instruct-25079 September 2026qwen/qwen3.5-9bnearest published output price: $0.30 against $0.20 per 1M
qwen/qwen3.5-397b-a17b9 September 2026z-ai/glm-5.2-instantnearest published output price: $3.60 against $4.50 per 1M

Most destinations are matched on the nearest published output price. Two vision entries use documented image input, and one entry uses the parameter band because the source dataset supplied no price. None is a benchmark or capability match. Test the candidate against your workload before changing production traffic.

Changing the model field in your client

The fix is a one-field change. The base URL and the API key stay as they are; only the string in model changes. First, check what the roster answers right now, with BASE_URL set to the serverless base URL shown in the client sample below:

curl https://api.lyceum.technology/openai/v1/models \
  -H "Authorization: Bearer lk_..."

Then change the one field in an OpenAI-SDK client:

from openai import OpenAI

client = OpenAI(
 api_key="lk_...",
 base_url="https://api.lyceum.technology/openai/v1", # unchanged
)

resp = client.chat.completions.create(
 model="z-ai/glm-5.3-flash", # the only field that changes
 messages=[{"role": "user", "content": "Explain prompt caching in one sentence."}],
)
print(resp.choices[0].message.content)

If you carry several removed strings across services, apply the full mapping in one pass:

REPLACEMENTS = {
 "nousresearch/hermes-4-70b": "z-ai/glm-5.3-flash",
 "qwen/qwen2.5-vl-72b-instruct": "deepseek/deepseek-v4-flash-0731",
 "qwen/qwen3-32b": "qwen/qwen3.5-9b",
 "nvidia/cosmos3-super-reasoner": "qwen/qwen3.5-9b",
 "nvidia/llama-3_1-nemotron-ultra-253b-v1": "minimax/minimax-m3",
 "nvidia/nemotron-3-nano-omni": "qwen/qwen3.5-9b",
 "qwen/qwen3-next-80b-a3b-thinking": "z-ai/glm-5.2-instant",
 "google/gemma-3-27b-it": "qwen/qwen3.5-9b",
 "openai/gpt-oss-120b": "z-ai/glm-5.3-flash",
 "nousresearch/hermes-4-405b": "qwen/qwen3.8-27b",
 "meta-llama/llama-3.3-70b-instruct": "z-ai/glm-5.3-flash",
 "openbmb/minicpm-v-4_5": "deepseek/deepseek-v4-flash-0731",
 "nvidia/nemotron-3-super-120b-a12b": "z-ai/glm-5.3-flash",
 "nvidia/nemotron-3-ultra-550b-a55b": "qwen/qwen3.8-27b",
 "nvidia/nvidia-nemotron-3-nano-30b-a3b": "qwen/qwen3.5-9b",
 "qwen/qwen3-30b-a3b-instruct-2507": "qwen/qwen3.5-9b",
 "qwen/qwen3.5-397b-a17b": "z-ai/glm-5.2-instant",
}

Model IDs are case-insensitive, and where a client rejects a model name containing a slash, the slash-free variant such as z-ai-glm-5.2 is accepted instead.

Keeping your config current

Before trusting any page, including this one, verify against the live roster. The check takes one request with your lk_ key as a Bearer token:

  1. Send GET /models on the base URL shown above, with your lk_ key as a Bearer token.
  2. Read the returned model IDs and confirm the string you call is present.
  3. If a string is missing, treat the replacement table as a shortlist, not a capability match, and test the candidate against your workload.
  4. Rerun the check after any deploy that touches model configuration.

Read errors defensively. Gateway errors, a bad key, an unknown model ID or insufficient credit, come back as a detail string; upstream errors use the OpenAI-compatible error.message shape. The documentation's own test client reads detail first and falls back to error.message, which is the order to copy. An unknown model ID returns a detail message that points at GET /models, the fastest route back to the live list.

For follow-up reading, the documentation index sits at docs.lyceum.technology/llms.txt, which lists every documentation page including Serverless Inference and the coding-agent guides.

Serverless Inference: one endpoint, one key

Serverless Inference is the product behind every string on this page: pre-hosted open models behind one OpenAI-compatible endpoint, metered per token, with no GPU to provision. For a team whose product runs on inference, the operational questions are the ones this page answered: which strings answer, where each runs, what each costs, and what to change when a string goes away.

  • Residency is a per-model fact, stated in each model's own record.
  • The price tables include strings for which the source data supplies the fields shown.
  • Replacement rows use explicit heuristics and require workload testing before a production change.

If you are evaluating endpoints for a European product, or replacing strings that stopped answering this week, start at Serverless Inference and run the roster check against your own key before you commit a single config line. A Lyceum video comparison puts these prices side by side with five other European providers and reports 120 measured API requests across Kimi K3, GLM-5.2 and Qwen3.5-9B.