# Lyceum: open models

> 23 open models behind one OpenAI-compatible API. Billed per token. No base fee.

Source: https://lyceum.technology/models/. Prices in US dollars per million tokens, from the platform on 5 October 2026.

## How to call them

- OpenAI-compatible base URL: https://api.lyceum.technology/openai/v1
- Anthropic-compatible base URL, for Claude Code: https://api.lyceum.technology/anthropic
- Authentication: an API key from https://dashboard.lyceum.technology, sent as a Bearer token.
- The `model` field takes the ID from the table below.
- Docs: https://docs.lyceum.technology
- Data: Prompts and outputs are processed, not stored, and never used for training. GPUs run in European data centres.

## Models

| Model | ID | Provider | Best for | Accepts | Context | Input | Cached input | Output |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| DeepSeek V4.1 Flash | `deepseek/deepseek-v4.1-flash` | DeepSeek | General | Text, Images | 1M | $0.50 | $0.13 | $1.50 |
| GLM-5.3 | `z-ai/glm-5.3` | Z.ai | Coding, Agents | Text | 1M | $1.40 | $0.26 | $4.40 |
| Kimi K3 | `moonshotai/kimi-k3` | Moonshot | Coding, Agents | Text, Images | 1M | $3.00 | $0.75 | $15.00 |
| GLM-5.3 Flash | `z-ai/glm-5.3-flash` | Z.ai | Fast | Text, Images | 1M | $0.20 | $0.05 | $0.50 |
| DeepSeek V4 Pro 0813 | `deepseek/deepseek-v4-pro-0813` | DeepSeek | Reasoning, Coding | Text | 1M | $2.00 | $0.50 | $4.00 |
| Kimi K2.7 Code | `moonshotai/kimi-k2.7-code` | Moonshot | Coding, Agents | Text, Images | 256K | $1.25 | $0.31 | $4.50 |
| Qwen3.8 2.4T A95B | `qwen/qwen3.8-2.4t-a95b` | Qwen | Agents | Text | 256K | $2.50 | $0.63 | $6.00 |
| Qwen3.8 Flash Next | `qwen/qwen3.8-flash-next` | Qwen | Fast | Text, Images | 256K | $0.20 | $0.05 | $0.50 |
| MiniMax M3 | `minimax/minimax-m3` | MiniMax | Long documents | Text, Images | 1M | $0.40 | $0.10 | $2.00 |
| GLM-5.2 | `z-ai/glm-5.2` | Z.ai | Coding, Reasoning | Text | 1M | $1.50 | $0.38 | $4.50 |
| Qwen3 Embedding 8B | `qwen/qwen3-embedding-8b` | Qwen | Embeddings | Text | 32K | $0.02 | $0.00 | Not listed |
| Qwen3.8 27B | `qwen/qwen3.8-27b` | Qwen | General | Text, Images | 256K | $0.40 | $0.10 | $2.40 |
| Lyceum Router | `lyceum/router` | Lyceum | Routing | Text, Images | 1M | $0.40 | $0.10 | $2.00 |
| Lyceum Simple | `lyceum/simple` | Lyceum | Routing | Text, Images | 256K | $0.20 | $0.05 | $0.50 |
| Lyceum Complex | `lyceum/complex` | Lyceum | Routing | Text, Images | 1M | $0.40 | $0.10 | $2.00 |
| Lyceum Reasoning | `lyceum/reasoning` | Lyceum | Routing | Text | 1M | $1.50 | $0.38 | $4.50 |
| DeepSeek V4 Flash | `deepseek/deepseek-v4-flash-0731` | DeepSeek | Fast, Agents | Text | 1M | $0.25 | $0.06 | $0.30 |
| deepseek/deepseek-v4-pro | `deepseek/deepseek-v4-pro` | DeepSeek | General |  | Not confirmed | $1.75 | $0.44 | $3.50 |
| Qwen3.8 27B Instant | `qwen/qwen3.8-27b-instant` | Qwen | Fast, General | Text, Images | 256K | $0.40 | $0.10 | $2.40 |
| Qwen3.8 Flash Next Instant | `qwen/qwen3.8-flash-next-instant` | Qwen | Fast | Text, Images | 256K | $0.20 | $0.05 | $0.50 |
| GLM-5.2 Instant | `z-ai/glm-5.2-instant` | Z.ai | Fast, Coding | Text | 1M | $1.50 | $0.38 | $4.50 |
| GLM-5.3 Flash Instant | `z-ai/glm-5.3-flash-instant` | Z.ai | Fast | Text, Images | 1M | $0.20 | $0.05 | $0.50 |
| GLM-5.3 Instant | `z-ai/glm-5.3-instant` | Z.ai | Fast, Coding | Text | 1M | $1.40 | $0.26 | $4.40 |

## Model details

### DeepSeek V4.1 Flash

A Flash model that reads text and images, built to make long, input-heavy agent workloads cheaper to run.

- Parameters: 552B backbone, plus 196B of Engram memory, 8B per token in prefill, 16B in decode active
- Licence: MIT (https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/LICENSE)
- Released: 10 September 2026
- Model card: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- Page: https://lyceum.technology/models/deepseek/deepseek-v4.1-flash/

### GLM-5.3

Z.ai’s large model for complex coding and long agent tasks, built on the same base model as GLM-5.2.

- Parameters: 744B, 40B active
- Licence: GLM-5.3 License (https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE)
- Released: 14 August 2026
- Model card: https://huggingface.co/zai-org/GLM-5.3
- Page: https://lyceum.technology/models/z-ai/glm-5.3/

### Kimi K3

Moonshot’s largest open model, for long coding sessions, agentic knowledge work and reasoning, with image input and a 1M-token context.

- Parameters: 2.8T, 104B active
- Licence: Kimi K3 License (https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE)
- Released: 16 July 2026
- Model card: https://huggingface.co/moonshotai/Kimi-K3
- Page: https://lyceum.technology/models/moonshotai/kimi-k3/

### GLM-5.3 Flash

A smaller, lower-cost GLM model that reads text and images, for coding and agent work such as browser and computer use.

- Parameters: 320B, 18B active
- Licence: MIT (https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE)
- Released: 26 August 2026
- Model card: https://huggingface.co/zai-org/GLM-5.3-Flash
- Page: https://lyceum.technology/models/z-ai/glm-5.3-flash/

### DeepSeek V4 Pro 0813

DeepSeek’s largest V4 model for reasoning, coding and agent work, in its general availability release of August 2026.

- Parameters: 1.6T, 49B active
- Licence: MIT (https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813/blob/main/LICENSE)
- Released: 13 August 2026
- Model card: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
- Page: https://lyceum.technology/models/deepseek/deepseek-v4-pro-0813/

### Kimi K2.7 Code

A coding model built on Kimi K2.6 for long software engineering tasks, which Moonshot says uses about 30% fewer thinking tokens than K2.6.

- Parameters: 1T, 32B active
- Licence: Modified MIT License (https://huggingface.co/moonshotai/Kimi-K2.7-Code/blob/main/LICENSE)
- Released: 12 June 2026
- Model card: https://huggingface.co/moonshotai/Kimi-K2.7-Code
- Page: https://lyceum.technology/models/moonshotai/kimi-k2.7-code/

### Qwen3.8 2.4T A95B

The open-weight Qwen3.8 flagship: text only, always in thinking mode, built for coding, research and long agentic tasks.

- Parameters: 2.4T, 95B active
- Licence: Qwen3.8-Max License (https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE)
- Released: 12 August 2026
- Model card: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
- Page: https://lyceum.technology/models/qwen/qwen3.8-2.4t-a95b/

### Qwen3.8 Flash Next

A low-cost preview of the Qwen4 architecture, for high-volume apps, tool-driven workflows and coding assistants.

- Parameters: 125B, plus 51B of n-gram embeddings and 4B for multi-token prediction, 6B active
- Licence: Qwen Community License 1.0 (https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE)
- Released: 26 August 2026
- Model card: https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- Page: https://lyceum.technology/models/qwen/qwen3.8-flash-next/

### MiniMax M3

MiniMax’s model for long coding and office tasks, with a 1M-token context and image input.

- Parameters: About 428B, About 23B active
- Licence: MiniMax Community License (https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE)
- Released: 1 June 2026
- Model card: https://huggingface.co/MiniMaxAI/MiniMax-M3
- Page: https://lyceum.technology/models/minimax/minimax-m3/

### GLM-5.2

Z.ai’s model for long agentic coding tasks, with a 1M-token context and a choice of how long it thinks.

- Parameters: 744B, 40B active
- Licence: MIT (https://huggingface.co/zai-org/GLM-5.2/blob/main/LICENSE)
- Released: 16 June 2026
- Model card: https://huggingface.co/zai-org/GLM-5.2
- Page: https://lyceum.technology/models/z-ai/glm-5.2/

### Qwen3 Embedding 8B

A multilingual text embedding model for search and retrieval, including code, and for classification and clustering.

- Parameters: 8B
- Licence: Apache 2.0 (https://huggingface.co/Qwen/Qwen3-Embedding-8B/blob/main/LICENSE)
- Released: 5 June 2025
- Model card: https://huggingface.co/Qwen/Qwen3-Embedding-8B
- Page: https://lyceum.technology/models/qwen/qwen3-embedding-8b/

### Qwen3.8 27B

A compact, dense Qwen3.8 model that reads text and images, bringing the family’s coding and agent gains to a smaller size.

- Parameters: 27B
- Licence: Apache 2.0 (https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE)
- Released: 14 August 2026
- Model card: https://huggingface.co/Qwen/Qwen3.8-27B
- Page: https://lyceum.technology/models/qwen/qwen3.8-27b/

### Lyceum Router

A routing alias: we choose the model for each request, and bill it as that model.

- Page: https://lyceum.technology/models/lyceum/router/

### Lyceum Simple

A routing alias for light requests, billed as the model it routes to.

- Page: https://lyceum.technology/models/lyceum/simple/

### Lyceum Complex

A routing alias for demanding requests, billed as the model it routes to.

- Page: https://lyceum.technology/models/lyceum/complex/

### Lyceum Reasoning

A routing alias for requests that need reasoning, billed as the model it routes to.

- Page: https://lyceum.technology/models/lyceum/reasoning/

### DeepSeek V4 Flash

The July 2026 update of DeepSeek’s smaller V4 model, tuned for stronger agentic coding and tool use.

- Parameters: 284B, 13B active
- Licence: MIT (https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/LICENSE)
- Released: 31 July 2026
- Model card: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
- Page: https://lyceum.technology/models/deepseek/deepseek-v4-flash-0731/

### deepseek/deepseek-v4-pro

See the API documentation for supported features

- Page: https://lyceum.technology/models/deepseek/deepseek-v4-pro/

### Qwen3.8 27B Instant

A compact, dense Qwen3.8 model that reads text and images, bringing the family’s coding and agent gains to a smaller size. Answers without a reasoning phase.

- Parameters: 27B
- Licence: Apache 2.0 (https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE)
- Released: 14 August 2026
- Model card: https://huggingface.co/Qwen/Qwen3.8-27B
- Page: https://lyceum.technology/models/qwen/qwen3.8-27b-instant/

### Qwen3.8 Flash Next Instant

A low-cost preview of the Qwen4 architecture, for high-volume apps, tool-driven workflows and coding assistants. Answers without a reasoning phase.

- Parameters: 125B, plus 51B of n-gram embeddings and 4B for multi-token prediction, 6B active
- Licence: Qwen Community License 1.0 (https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE)
- Released: 26 August 2026
- Model card: https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- Page: https://lyceum.technology/models/qwen/qwen3.8-flash-next-instant/

### GLM-5.2 Instant

Z.ai’s model for long agentic coding tasks, with a 1M-token context and a choice of how long it thinks. Answers without a reasoning phase.

- Parameters: 744B, 40B active
- Licence: MIT (https://huggingface.co/zai-org/GLM-5.2/blob/main/LICENSE)
- Released: 16 June 2026
- Model card: https://huggingface.co/zai-org/GLM-5.2
- Page: https://lyceum.technology/models/z-ai/glm-5.2-instant/

### GLM-5.3 Flash Instant

A smaller, lower-cost GLM model that reads text and images, for coding and agent work such as browser and computer use. Answers without a reasoning phase.

- Parameters: 320B, 18B active
- Licence: MIT (https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE)
- Released: 26 August 2026
- Model card: https://huggingface.co/zai-org/GLM-5.3-Flash
- Page: https://lyceum.technology/models/z-ai/glm-5.3-flash-instant/

### GLM-5.3 Instant

Z.ai’s large model for complex coding and long agent tasks, built on the same base model as GLM-5.2. Answers without a reasoning phase.

- Parameters: 744B, 40B active
- Licence: GLM-5.3 License (https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE)
- Released: 14 August 2026
- Model card: https://huggingface.co/zai-org/GLM-5.3
- Page: https://lyceum.technology/models/z-ai/glm-5.3-instant/

## Retired models

Lyceum no longer serves these models. Requests to their model IDs return `404 model not found`. Switch to the model listed next to each one: it keeps what the old model could do, such as image input or reasoning. Source: https://lyceum.technology/models/#retired-models

| Retired model | Provider | Former model ID | Use instead |
| --- | --- | --- | --- |
| DeepSeek V3.2 | DeepSeek | Not published | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) |
| GLM-5 | Z.ai | Not published | GLM-5.3 (`z-ai/glm-5.3`) |
| Kimi K2.5 | Moonshot | Not published | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Qwen3.5 397B A17B | Qwen | `qwen/qwen3.5-397b-a17b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Qwen3 235B Thinking | Qwen | Not published | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) |
| Qwen3 Next 80B A3B Thinking | Qwen | `qwen/qwen3-next-80b-a3b-thinking` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Qwen3 32B | Qwen | `qwen/qwen3-32b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Qwen3 30B A3B | Qwen | `qwen/qwen3-30b-a3b-instruct-2507` | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) |
| Qwen3 Coder 30B A3B | Qwen | Not published | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) |
| Qwen2.5 VL 72B | Qwen | `qwen/qwen2.5-vl-72b-instruct` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| INTELLECT-3 | PrimeIntellect | Not published | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Llama 3.3 70B | Meta | `meta-llama/llama-3.3-70b-instruct` | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) |
| Gemma 3 27B | Google | `google/gemma-3-27b-it` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| gpt-oss 120B | OpenAI | `openai/gpt-oss-120b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Hermes 4 405B | NousResearch | `nousresearch/hermes-4-405b` | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) |
| Hermes 4 70B | NousResearch | `nousresearch/hermes-4-70b` | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) |
| Nemotron 3 Ultra 550B | NVIDIA | `nvidia/nemotron-3-ultra-550b-a55b` | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) |
| Nemotron Ultra 253B | NVIDIA | `nvidia/llama-3_1-nemotron-ultra-253b-v1` | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) |
| Nemotron 3 Super 120B | NVIDIA | `nvidia/nemotron-3-super-120b-a12b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Nemotron 3 Nano 30B | NVIDIA | `nvidia/nvidia-nemotron-3-nano-30b-a3b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Nemotron 3 Nano Omni | NVIDIA | `nvidia/nemotron-3-nano-omni` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| Cosmos 3 Super Reasoner | NVIDIA | `nvidia/cosmos3-super-reasoner` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| MiniCPM-V 4.5 | OpenBMB | `openbmb/minicpm-v-4_5` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) |
| MiniMax M2.5 | MiniMax | `minimax/minimax-m2.5` | MiniMax M3 (`minimax/minimax-m3`) |
| Kimi K2.6 | Moonshot | `moonshotai/kimi-k2.6` | Kimi K3 (`moonshotai/kimi-k3`) |
| GLM-5.1 | Z.ai | `z-ai/glm-5.1` | GLM-5.3 (`z-ai/glm-5.3`) |
| Qwen3.5 9B | Qwen | `qwen/qwen3.5-9b` | Qwen3.8 Flash Next (`qwen/qwen3.8-flash-next`) |
| Qwen3 235B A22B | Qwen | `qwen/qwen3-235b-a22b-instruct-2507` | Qwen3.8 2.4T A95B (`qwen/qwen3.8-2.4t-a95b`) |
