Find your model
23 open models behind one OpenAI-compatible API. Search for what you need, and compare prices per million tokens.
Search by name, provider, task, context window or price. Typos are fine.
23 models
Copy for agents# Lyceum: open models > 23 open models behind one OpenAI-compatible API. Billed per token. No base fee. Source: https://lyceum.technology/models/. Prices in US dollars per million tokens, from the platform on 5 October 2026. ## How to call them - OpenAI-compatible base URL: https://api.lyceum.technology/openai/v1 - Anthropic-compatible base URL, for Claude Code: https://api.lyceum.technology/anthropic - Authentication: an API key from https://dashboard.lyceum.technology, sent as a Bearer token. - The `model` field takes the ID from the table below. - Docs: https://docs.lyceum.technology - Data: Prompts and outputs are processed, not stored, and never used for training. GPUs run in European data centres. ## Models | Model | ID | Provider | Best for | Accepts | Context | Input | Cached input | Output | | --- | --- | --- | --- | --- | --- | --- | --- | --- | | DeepSeek V4.1 Flash | `deepseek/deepseek-v4.1-flash` | DeepSeek | General | Text, Images | 1M | $0.50 | $0.13 | $1.50 | | GLM-5.3 | `z-ai/glm-5.3` | Z.ai | Coding, Agents | Text | 1M | $1.40 | $0.26 | $4.40 | | Kimi K3 | `moonshotai/kimi-k3` | Moonshot | Coding, Agents | Text, Images | 1M | $3.00 | $0.75 | $15.00 | | GLM-5.3 Flash | `z-ai/glm-5.3-flash` | Z.ai | Fast | Text, Images | 1M | $0.20 | $0.05 | $0.50 | | DeepSeek V4 Pro 0813 | `deepseek/deepseek-v4-pro-0813` | DeepSeek | Reasoning, Coding | Text | 1M | $2.00 | $0.50 | $4.00 | | Kimi K2.7 Code | `moonshotai/kimi-k2.7-code` | Moonshot | Coding, Agents | Text, Images | 256K | $1.25 | $0.31 | $4.50 | | Qwen3.8 2.4T A95B | `qwen/qwen3.8-2.4t-a95b` | Qwen | Agents | Text | 256K | $2.50 | $0.63 | $6.00 | | Qwen3.8 Flash Next | `qwen/qwen3.8-flash-next` | Qwen | Fast | Text, Images | 256K | $0.20 | $0.05 | $0.50 | | MiniMax M3 | `minimax/minimax-m3` | MiniMax | Long documents | Text, Images | 1M | $0.40 | $0.10 | $2.00 | | GLM-5.2 | `z-ai/glm-5.2` | Z.ai | Coding, Reasoning | Text | 1M | $1.50 | $0.38 | $4.50 | | Qwen3 Embedding 8B | `qwen/qwen3-embedding-8b` | Qwen | Embeddings | Text | 32K | $0.02 | $0.00 | Not listed | | Qwen3.8 27B | `qwen/qwen3.8-27b` | Qwen | General | Text, Images | 256K | $0.40 | $0.10 | $2.40 | | Lyceum Router | `lyceum/router` | Lyceum | Routing | Text, Images | 1M | $0.40 | $0.10 | $2.00 | | Lyceum Simple | `lyceum/simple` | Lyceum | Routing | Text, Images | 256K | $0.20 | $0.05 | $0.50 | | Lyceum Complex | `lyceum/complex` | Lyceum | Routing | Text, Images | 1M | $0.40 | $0.10 | $2.00 | | Lyceum Reasoning | `lyceum/reasoning` | Lyceum | Routing | Text | 1M | $1.50 | $0.38 | $4.50 | | DeepSeek V4 Flash | `deepseek/deepseek-v4-flash-0731` | DeepSeek | Fast, Agents | Text | 1M | $0.25 | $0.06 | $0.30 | | deepseek/deepseek-v4-pro | `deepseek/deepseek-v4-pro` | DeepSeek | General | | Not confirmed | $1.75 | $0.44 | $3.50 | | Qwen3.8 27B Instant | `qwen/qwen3.8-27b-instant` | Qwen | Fast, General | Text, Images | 256K | $0.40 | $0.10 | $2.40 | | Qwen3.8 Flash Next Instant | `qwen/qwen3.8-flash-next-instant` | Qwen | Fast | Text, Images | 256K | $0.20 | $0.05 | $0.50 | | GLM-5.2 Instant | `z-ai/glm-5.2-instant` | Z.ai | Fast, Coding | Text | 1M | $1.50 | $0.38 | $4.50 | | GLM-5.3 Flash Instant | `z-ai/glm-5.3-flash-instant` | Z.ai | Fast | Text, Images | 1M | $0.20 | $0.05 | $0.50 | | GLM-5.3 Instant | `z-ai/glm-5.3-instant` | Z.ai | Fast, Coding | Text | 1M | $1.40 | $0.26 | $4.40 | ## Model details ### DeepSeek V4.1 Flash A Flash model that reads text and images, built to make long, input-heavy agent workloads cheaper to run. - Parameters: 552B backbone, plus 196B of Engram memory, 8B per token in prefill, 16B in decode active - Licence: MIT (https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/LICENSE) - Released: 10 September 2026 - Model card: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash - Page: https://lyceum.technology/models/deepseek/deepseek-v4.1-flash/ ### GLM-5.3 Z.ai’s large model for complex coding and long agent tasks, built on the same base model as GLM-5.2. - Parameters: 744B, 40B active - Licence: GLM-5.3 License (https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE) - Released: 14 August 2026 - Model card: https://huggingface.co/zai-org/GLM-5.3 - Page: https://lyceum.technology/models/z-ai/glm-5.3/ ### Kimi K3 Moonshot’s largest open model, for long coding sessions, agentic knowledge work and reasoning, with image input and a 1M-token context. - Parameters: 2.8T, 104B active - Licence: Kimi K3 License (https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE) - Released: 16 July 2026 - Model card: https://huggingface.co/moonshotai/Kimi-K3 - Page: https://lyceum.technology/models/moonshotai/kimi-k3/ ### GLM-5.3 Flash A smaller, lower-cost GLM model that reads text and images, for coding and agent work such as browser and computer use. - Parameters: 320B, 18B active - Licence: MIT (https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE) - Released: 26 August 2026 - Model card: https://huggingface.co/zai-org/GLM-5.3-Flash - Page: https://lyceum.technology/models/z-ai/glm-5.3-flash/ ### DeepSeek V4 Pro 0813 DeepSeek’s largest V4 model for reasoning, coding and agent work, in its general availability release of August 2026. - Parameters: 1.6T, 49B active - Licence: MIT (https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813/blob/main/LICENSE) - Released: 13 August 2026 - Model card: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813 - Page: https://lyceum.technology/models/deepseek/deepseek-v4-pro-0813/ ### Kimi K2.7 Code A coding model built on Kimi K2.6 for long software engineering tasks, which Moonshot says uses about 30% fewer thinking tokens than K2.6. - Parameters: 1T, 32B active - Licence: Modified MIT License (https://huggingface.co/moonshotai/Kimi-K2.7-Code/blob/main/LICENSE) - Released: 12 June 2026 - Model card: https://huggingface.co/moonshotai/Kimi-K2.7-Code - Page: https://lyceum.technology/models/moonshotai/kimi-k2.7-code/ ### Qwen3.8 2.4T A95B The open-weight Qwen3.8 flagship: text only, always in thinking mode, built for coding, research and long agentic tasks. - Parameters: 2.4T, 95B active - Licence: Qwen3.8-Max License (https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE) - Released: 12 August 2026 - Model card: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B - Page: https://lyceum.technology/models/qwen/qwen3.8-2.4t-a95b/ ### Qwen3.8 Flash Next A low-cost preview of the Qwen4 architecture, for high-volume apps, tool-driven workflows and coding assistants. - Parameters: 125B, plus 51B of n-gram embeddings and 4B for multi-token prediction, 6B active - Licence: Qwen Community License 1.0 (https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE) - Released: 26 August 2026 - Model card: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - Page: https://lyceum.technology/models/qwen/qwen3.8-flash-next/ ### MiniMax M3 MiniMax’s model for long coding and office tasks, with a 1M-token context and image input. - Parameters: About 428B, About 23B active - Licence: MiniMax Community License (https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE) - Released: 1 June 2026 - Model card: https://huggingface.co/MiniMaxAI/MiniMax-M3 - Page: https://lyceum.technology/models/minimax/minimax-m3/ ### GLM-5.2 Z.ai’s model for long agentic coding tasks, with a 1M-token context and a choice of how long it thinks. - Parameters: 744B, 40B active - Licence: MIT (https://huggingface.co/zai-org/GLM-5.2/blob/main/LICENSE) - Released: 16 June 2026 - Model card: https://huggingface.co/zai-org/GLM-5.2 - Page: https://lyceum.technology/models/z-ai/glm-5.2/ ### Qwen3 Embedding 8B A multilingual text embedding model for search and retrieval, including code, and for classification and clustering. - Parameters: 8B - Licence: Apache 2.0 (https://huggingface.co/Qwen/Qwen3-Embedding-8B/blob/main/LICENSE) - Released: 5 June 2025 - Model card: https://huggingface.co/Qwen/Qwen3-Embedding-8B - Page: https://lyceum.technology/models/qwen/qwen3-embedding-8b/ ### Qwen3.8 27B A compact, dense Qwen3.8 model that reads text and images, bringing the family’s coding and agent gains to a smaller size. - Parameters: 27B - Licence: Apache 2.0 (https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE) - Released: 14 August 2026 - Model card: https://huggingface.co/Qwen/Qwen3.8-27B - Page: https://lyceum.technology/models/qwen/qwen3.8-27b/ ### Lyceum Router A routing alias: we choose the model for each request, and bill it as that model. - Page: https://lyceum.technology/models/lyceum/router/ ### Lyceum Simple A routing alias for light requests, billed as the model it routes to. - Page: https://lyceum.technology/models/lyceum/simple/ ### Lyceum Complex A routing alias for demanding requests, billed as the model it routes to. - Page: https://lyceum.technology/models/lyceum/complex/ ### Lyceum Reasoning A routing alias for requests that need reasoning, billed as the model it routes to. - Page: https://lyceum.technology/models/lyceum/reasoning/ ### DeepSeek V4 Flash The July 2026 update of DeepSeek’s smaller V4 model, tuned for stronger agentic coding and tool use. - Parameters: 284B, 13B active - Licence: MIT (https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/LICENSE) - Released: 31 July 2026 - Model card: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 - Page: https://lyceum.technology/models/deepseek/deepseek-v4-flash-0731/ ### deepseek/deepseek-v4-pro See the API documentation for supported features - Page: https://lyceum.technology/models/deepseek/deepseek-v4-pro/ ### Qwen3.8 27B Instant A compact, dense Qwen3.8 model that reads text and images, bringing the family’s coding and agent gains to a smaller size. Answers without a reasoning phase. - Parameters: 27B - Licence: Apache 2.0 (https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE) - Released: 14 August 2026 - Model card: https://huggingface.co/Qwen/Qwen3.8-27B - Page: https://lyceum.technology/models/qwen/qwen3.8-27b-instant/ ### Qwen3.8 Flash Next Instant A low-cost preview of the Qwen4 architecture, for high-volume apps, tool-driven workflows and coding assistants. Answers without a reasoning phase. - Parameters: 125B, plus 51B of n-gram embeddings and 4B for multi-token prediction, 6B active - Licence: Qwen Community License 1.0 (https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE) - Released: 26 August 2026 - Model card: https://huggingface.co/Qwen/Qwen3.8-Flash-Next - Page: https://lyceum.technology/models/qwen/qwen3.8-flash-next-instant/ ### GLM-5.2 Instant Z.ai’s model for long agentic coding tasks, with a 1M-token context and a choice of how long it thinks. Answers without a reasoning phase. - Parameters: 744B, 40B active - Licence: MIT (https://huggingface.co/zai-org/GLM-5.2/blob/main/LICENSE) - Released: 16 June 2026 - Model card: https://huggingface.co/zai-org/GLM-5.2 - Page: https://lyceum.technology/models/z-ai/glm-5.2-instant/ ### GLM-5.3 Flash Instant A smaller, lower-cost GLM model that reads text and images, for coding and agent work such as browser and computer use. Answers without a reasoning phase. - Parameters: 320B, 18B active - Licence: MIT (https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE) - Released: 26 August 2026 - Model card: https://huggingface.co/zai-org/GLM-5.3-Flash - Page: https://lyceum.technology/models/z-ai/glm-5.3-flash-instant/ ### GLM-5.3 Instant Z.ai’s large model for complex coding and long agent tasks, built on the same base model as GLM-5.2. Answers without a reasoning phase. - Parameters: 744B, 40B active - Licence: GLM-5.3 License (https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE) - Released: 14 August 2026 - Model card: https://huggingface.co/zai-org/GLM-5.3 - Page: https://lyceum.technology/models/z-ai/glm-5.3-instant/ ## Retired models Lyceum no longer serves these models. Requests to their model IDs return `404 model not found`. Switch to the model listed next to each one: it keeps what the old model could do, such as image input or reasoning. Source: https://lyceum.technology/models/#retired-models | Retired model | Provider | Former model ID | Use instead | | --- | --- | --- | --- | | DeepSeek V3.2 | DeepSeek | Not published | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) | | GLM-5 | Z.ai | Not published | GLM-5.3 (`z-ai/glm-5.3`) | | Kimi K2.5 | Moonshot | Not published | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Qwen3.5 397B A17B | Qwen | `qwen/qwen3.5-397b-a17b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Qwen3 235B Thinking | Qwen | Not published | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) | | Qwen3 Next 80B A3B Thinking | Qwen | `qwen/qwen3-next-80b-a3b-thinking` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Qwen3 32B | Qwen | `qwen/qwen3-32b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Qwen3 30B A3B | Qwen | `qwen/qwen3-30b-a3b-instruct-2507` | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) | | Qwen3 Coder 30B A3B | Qwen | Not published | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) | | Qwen2.5 VL 72B | Qwen | `qwen/qwen2.5-vl-72b-instruct` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | INTELLECT-3 | PrimeIntellect | Not published | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Llama 3.3 70B | Meta | `meta-llama/llama-3.3-70b-instruct` | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) | | Gemma 3 27B | Google | `google/gemma-3-27b-it` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | gpt-oss 120B | OpenAI | `openai/gpt-oss-120b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Hermes 4 405B | NousResearch | `nousresearch/hermes-4-405b` | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) | | Hermes 4 70B | NousResearch | `nousresearch/hermes-4-70b` | DeepSeek V4.1 Flash (`deepseek/deepseek-v4.1-flash`) | | Nemotron 3 Ultra 550B | NVIDIA | `nvidia/nemotron-3-ultra-550b-a55b` | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) | | Nemotron Ultra 253B | NVIDIA | `nvidia/llama-3_1-nemotron-ultra-253b-v1` | DeepSeek V4 Pro 0813 (`deepseek/deepseek-v4-pro-0813`) | | Nemotron 3 Super 120B | NVIDIA | `nvidia/nemotron-3-super-120b-a12b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Nemotron 3 Nano 30B | NVIDIA | `nvidia/nvidia-nemotron-3-nano-30b-a3b` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Nemotron 3 Nano Omni | NVIDIA | `nvidia/nemotron-3-nano-omni` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | Cosmos 3 Super Reasoner | NVIDIA | `nvidia/cosmos3-super-reasoner` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | MiniCPM-V 4.5 | OpenBMB | `openbmb/minicpm-v-4_5` | GLM-5.3 Flash (`z-ai/glm-5.3-flash`) | | MiniMax M2.5 | MiniMax | `minimax/minimax-m2.5` | MiniMax M3 (`minimax/minimax-m3`) | | Kimi K2.6 | Moonshot | `moonshotai/kimi-k2.6` | Kimi K3 (`moonshotai/kimi-k3`) | | GLM-5.1 | Z.ai | `z-ai/glm-5.1` | GLM-5.3 (`z-ai/glm-5.3`) | | Qwen3.5 9B | Qwen | `qwen/qwen3.5-9b` | Qwen3.8 Flash Next (`qwen/qwen3.8-flash-next`) | | Qwen3 235B A22B | Qwen | `qwen/qwen3-235b-a22b-instruct-2507` | Qwen3.8 2.4T A95B (`qwen/qwen3.8-2.4t-a95b`) |Prices in USD per 1M tokens. Click a price to copy that model’s prices.
| Model | Best for | Accepts | Context | Input | Cached input | Output |
|---|---|---|---|---|---|---|
| 1M | $0.50 | $0.13 | $1.50 | |||
About DeepSeek V4.1 Flash A Flash model that reads text and images, built to make long, input-heavy agent workloads cheaper to run.
API details and quickstartModel card on Hugging FaceDeepSeek’s announcement | ||||||
| 1M | $1.40 | $0.26 | $4.40 | |||
About GLM-5.3 Z.ai’s large model for complex coding and long agent tasks, built on the same base model as GLM-5.2.
API details and quickstartModel card on Hugging FaceZ.ai’s announcement | ||||||
| 1M | $3.00 | $0.75 | $15.00 | |||
About Kimi K3 Moonshot’s largest open model, for long coding sessions, agentic knowledge work and reasoning, with image input and a 1M-token context.
API details and quickstartModel card on Hugging FaceMoonshot’s announcement | ||||||
| 1M | $0.20 | $0.05 | $0.50 | |||
About GLM-5.3 Flash A smaller, lower-cost GLM model that reads text and images, for coding and agent work such as browser and computer use.
API details and quickstartModel card on Hugging FaceZ.ai’s announcement | ||||||
| 1M | $2.00 | $0.50 | $4.00 | |||
About DeepSeek V4 Pro 0813 DeepSeek’s largest V4 model for reasoning, coding and agent work, in its general availability release of August 2026.
API details and quickstartModel card on Hugging FaceDeepSeek’s announcement | ||||||
| 256K | $1.25 | $0.31 | $4.50 | |||
About Kimi K2.7 Code A coding model built on Kimi K2.6 for long software engineering tasks, which Moonshot says uses about 30% fewer thinking tokens than K2.6.
API details and quickstartModel card on Hugging FaceMoonshot’s announcement | ||||||
| 256K | $2.50 | $0.63 | $6.00 | |||
About Qwen3.8 2.4T A95B The open-weight Qwen3.8 flagship: text only, always in thinking mode, built for coding, research and long agentic tasks.
API details and quickstartModel card on Hugging FaceQwen’s announcement | ||||||
| 256K | $0.20 | $0.05 | $0.50 | |||
About Qwen3.8 Flash Next A low-cost preview of the Qwen4 architecture, for high-volume apps, tool-driven workflows and coding assistants.
API details and quickstartModel card on Hugging FaceQwen’s announcement | ||||||
| 1M | $0.40 | $0.10 | $2.00 | |||
About MiniMax M3 MiniMax’s model for long coding and office tasks, with a 1M-token context and image input.
API details and quickstartModel card on Hugging FaceMiniMax’s announcement | ||||||
| 1M | $1.50 | $0.38 | $4.50 | |||
About GLM-5.2 Z.ai’s model for long agentic coding tasks, with a 1M-token context and a choice of how long it thinks.
API details and quickstartModel card on Hugging FaceZ.ai’s announcement | ||||||
| 32K | $0.02 | $0.00 | Input only | |||
About Qwen3 Embedding 8B A multilingual text embedding model for search and retrieval, including code, and for classification and clustering.
API details and quickstartModel card on Hugging FaceQwen’s announcement | ||||||
| 256K | $0.40 | $0.10 | $2.40 | |||
About Qwen3.8 27B A compact, dense Qwen3.8 model that reads text and images, bringing the family’s coding and agent gains to a smaller size.
| ||||||
| 1M | $0.40 | $0.10 | $2.00 | |||
About Lyceum Router A routing alias: we choose the model for each request, and bill it as that model.
| ||||||
| 256K | $0.20 | $0.05 | $0.50 | |||
About Lyceum Simple A routing alias for light requests, billed as the model it routes to.
| ||||||
| 1M | $0.40 | $0.10 | $2.00 | |||
About Lyceum Complex A routing alias for demanding requests, billed as the model it routes to.
| ||||||
| 1M | $1.50 | $0.38 | $4.50 | |||
About Lyceum Reasoning A routing alias for requests that need reasoning, billed as the model it routes to.
| ||||||
| 1M | $0.25 | $0.06 | $0.30 | |||
About DeepSeek V4 Flash The July 2026 update of DeepSeek’s smaller V4 model, tuned for stronger agentic coding and tool use.
API details and quickstartModel card on Hugging FaceDeepSeek’s announcement | ||||||
| Not confirmed | Not confirmed | Not confirmed | $1.75 | $0.44 | $3.50 | |
About deepseek/deepseek-v4-pro See the API documentation for supported features
| ||||||
| 256K | $0.40 | $0.10 | $2.40 | |||
About Qwen3.8 27B Instant A compact, dense Qwen3.8 model that reads text and images, bringing the family’s coding and agent gains to a smaller size. Answers without a reasoning phase.
| ||||||
| 256K | $0.20 | $0.05 | $0.50 | |||
About Qwen3.8 Flash Next Instant A low-cost preview of the Qwen4 architecture, for high-volume apps, tool-driven workflows and coding assistants. Answers without a reasoning phase.
API details and quickstartModel card on Hugging FaceQwen’s announcement | ||||||
| 1M | $1.50 | $0.38 | $4.50 | |||
About GLM-5.2 Instant Z.ai’s model for long agentic coding tasks, with a 1M-token context and a choice of how long it thinks. Answers without a reasoning phase.
API details and quickstartModel card on Hugging FaceZ.ai’s announcement | ||||||
| 1M | $0.20 | $0.05 | $0.50 | |||
About GLM-5.3 Flash Instant A smaller, lower-cost GLM model that reads text and images, for coding and agent work such as browser and computer use. Answers without a reasoning phase.
API details and quickstartModel card on Hugging FaceZ.ai’s announcement | ||||||
| 1M | $1.40 | $0.26 | $4.40 | |||
About GLM-5.3 Instant Z.ai’s large model for complex coding and long agent tasks, built on the same base model as GLM-5.2. Answers without a reasoning phase.
API details and quickstartModel card on Hugging FaceZ.ai’s announcement | ||||||
No model matches all of that.
Remove one of the parts above, or .
Retired models (28)
Lyceum no longer serves these models. Requests to their model IDs return 404 model not found. Switch to the model listed next to each one: it keeps what the old model could do, such as image input or reasoning.
| Retired model | Former model ID | Use instead |
|---|---|---|
| DeepSeek V3.2DeepSeek | Not published | DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash |
| GLM-5Z.ai | Not published | GLM-5.3z-ai/glm-5.3 |
| Kimi K2.5Moonshot | Not published | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Qwen3.5 397B A17BQwen | qwen/qwen3.5-397b-a17b | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Qwen3 235B ThinkingQwen | Not published | DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 |
| Qwen3 Next 80B A3B ThinkingQwen | qwen/qwen3-next-80b-a3b-thinking | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Qwen3 32BQwen | qwen/qwen3-32b | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Qwen3 30B A3BQwen | qwen/qwen3-30b-a3b-instruct-2507 | DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash |
| Qwen3 Coder 30B A3BQwen | Not published | DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash |
| Qwen2.5 VL 72BQwen | qwen/qwen2.5-vl-72b-instruct | GLM-5.3 Flashz-ai/glm-5.3-flash |
| INTELLECT-3PrimeIntellect | Not published | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Llama 3.3 70BMeta | meta-llama/llama-3.3-70b-instruct | DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash |
| Gemma 3 27BGoogle | google/gemma-3-27b-it | GLM-5.3 Flashz-ai/glm-5.3-flash |
| gpt-oss 120BOpenAI | openai/gpt-oss-120b | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Hermes 4 405BNousResearch | nousresearch/hermes-4-405b | DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 |
| Hermes 4 70BNousResearch | nousresearch/hermes-4-70b | DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash |
| Nemotron 3 Ultra 550BNVIDIA | nvidia/nemotron-3-ultra-550b-a55b | DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 |
| Nemotron Ultra 253BNVIDIA | nvidia/llama-3_1-nemotron-ultra-253b-v1 | DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 |
| Nemotron 3 Super 120BNVIDIA | nvidia/nemotron-3-super-120b-a12b | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Nemotron 3 Nano 30BNVIDIA | nvidia/nvidia-nemotron-3-nano-30b-a3b | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Nemotron 3 Nano OmniNVIDIA | nvidia/nemotron-3-nano-omni | GLM-5.3 Flashz-ai/glm-5.3-flash |
| Cosmos 3 Super ReasonerNVIDIA | nvidia/cosmos3-super-reasoner | GLM-5.3 Flashz-ai/glm-5.3-flash |
| MiniCPM-V 4.5OpenBMB | openbmb/minicpm-v-4_5 | GLM-5.3 Flashz-ai/glm-5.3-flash |
| MiniMax M2.5MiniMax | minimax/minimax-m2.5 | MiniMax M3minimax/minimax-m3 |
| Kimi K2.6Moonshot | moonshotai/kimi-k2.6 | Kimi K3moonshotai/kimi-k3 |
| GLM-5.1Z.ai | z-ai/glm-5.1 | GLM-5.3z-ai/glm-5.3 |
| Qwen3.5 9BQwen | qwen/qwen3.5-9b | Qwen3.8 Flash Nextqwen/qwen3.8-flash-next |
| Qwen3 235B A22BQwen | qwen/qwen3-235b-a22b-instruct-2507 | Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b |