What's new
ZAI has officially released GLM-5.3, a frontier Mixture-of-Experts (MoE) model engineered specifically for complex software engineering, multi-turn agentic workflows, and vulnerability analysis. Rather than altering the underlying parameter count or pre-training dataset from scratch, GLM-5.3 retains the exact same base model architecture as GLM-5.2. Every measured performance gain across coding benchmarks and multi-step reasoning tasks stems exclusively from scaled post-training pipelines.
The core engineering focus behind GLM-5.3 centers on expanding reinforcement learning (RL) environments so they resemble real units of expert work rather than isolated coding exercises. ZAI says a useful task environment has to be executable, verifiable and close to real professional work, and that it built pipelines to synthesize such environments end to end, plus the RL reward signal for a subset of tasks. That matters for product teams building autonomous coding tools, refactoring engines or security auditing agents.
- Architecture parity with GLM-5.2: the same base model is used, so every reported gain comes from post-training.
- Scaled RL infrastructure: built on the GLM-5.2 stack, including slime for large-scale asynchronous training and SAO for RL on long-horizon tasks.
- Executable developer environments: moving beyond static unit tests to dynamic environments requiring compiler diagnostics, dependency resolution, and multi-turn shell interactions.
- Mandatory thinking mode: reasoning tokens are now active by default, replacing binary toggles with configurable effort levels to match target compute budgets.
Specs and context window
Under the hood, GLM-5.3 is a sparse Mixture-of-Experts (MoE) model, so only a fraction of its parameters are active per token. On Lyceum it is served on vLLM, the same serving stack used across the supported-models catalogue.
A primary strength of GLM-5.3 is its native 1M-token context window. That much context lets developers ingest entire enterprise repositories, complex system architecture documents, or exhaustive technical specifications in a single prompt. It eliminates brittle retrieval fragmentation when building code-review agents or complex technical summarization tools.
| Architectural Parameter | Specification | Engineering Impact |
|---|---|---|
| Architecture | Mixture-of-Experts (MoE) | Deep parametric capacity with sparse per-token compute |
| Base model | Same base model as GLM-5.2 | All reported gains come from post-training, not new pre-training |
| Context Length | 1M tokens | Ingests massive codebases, logs, and technical documentation in one pass |
| Serving Engine | vLLM | Standardized tensor parallelism, continuous batching, and PagedAttention |
| Reasoning Control | Configurable reasoning effort | Reasoning token budget can be matched to task complexity |
Long-context serving builds on IndexShare, the efficient long-context processing component ZAI introduced with GLM-5.2 and carried into GLM-5.3, together with SAO for RL on long-horizon tasks.
Coding and agentic benchmarks
In vendor-reported benchmarks published by ZAI, GLM-5.3 demonstrates substantial generational gains over GLM-5.2, particularly when tasked with multi-step agentic coding workflows. On ZAI's internal Z.ai Code Bench, GLM-5.3 recorded a 50% relative improvement over GLM-5.2. This internal benchmark measures end-to-end task completion and fine-grained checklist accuracy in isolated local developer setups rather than grading snippet-level syntax.
Public evaluations reflect a similar leap in capabilities. On Terminal Bench 3.0, which evaluates interactive shell commands, environment navigation, and multi-step terminal workflows, GLM-5.3 achieved a score of 28.3, jumping from the 4.6 recorded by GLM-5.2. This makes GLM-5.3 the leading open-weights model on this evaluation, outperforming Kimi K3 (17.4) and approaching closed frontier models.
| Benchmark | GLM-5.3 (Vendor-Reported) | GLM-5.2 (Vendor-Reported) | Kimi K3 (Vendor-Reported) | Metric Definition |
|---|---|---|---|---|
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | Interactive terminal operations and tool execution |
| DeepSWE v1.1 | 66.9 | 46.2 | 67.5 | Multi-file software engineering bug resolution |
| Z.ai Code Bench (Max Effort) | 34.5% | 23.4% | Not Reported | End-to-end task completion in real developer environments |
| PostTrainBench | 39.8 | 31.7 | 32.0 | Long-horizon RL and optimization reasoning |
| Agents' Last Exam | 28.5 | 23.8 | Not Reported | Complex multi-disciplinary agentic reasoning |
Token efficiency improved alongside accuracy: ZAI reports markedly stronger agentic coding results than GLM-5.2 at every effort level while consuming fewer output tokens on Z.ai Code Bench. For production systems, shorter outputs directly lower per-task latency and execution cost.
Emergent cyber capability
A notable development resulting from ZAI's post-training expansion is the emergence of advanced cybersecurity reasoning. By incorporating vulnerability analysis environments and exploit verification data into the reinforcement learning loop, the model began autonomously synthesizing complete, multi-stage exploitation sequences across complex architectures.
On CyberGym, a white-box source code evaluation that tests whether an AI can identify and trigger software faults, GLM-5.3 achieved a vendor-reported score of 84.5 (source: zai). ZAI's own comparison table puts that above GLM-5.2 (77.2), Kimi K3 (80.0) and GPT-5.6 Sol (83.6) on the same benchmark. On ExploitBench, which demands deeper reasoning over real-world Common Vulnerabilities and Exposures (CVEs), GLM-5.3 scored 54.4, more than doubling GLM-5.2's 24.4.
- CyberGym fault triggering: 84.5 vendor-reported score on white-box source code vulnerability analysis, state of the art on that benchmark.
- ExploitBench reasoning: 54.4 on multi-stage exploit planning, more than doubling GLM-5.2's 24.4.
- Gains up the exploitation chain: ZAI reports its largest improvements further along the exploitation chain, including on ExploitGym.
- Staged weight release: ZAI said it would publish the open weights two weeks after launch, once safety evaluation and hardening were complete.
These capabilities make GLM-5.3 a potent foundation model for automated penetration testing tools, static code analysis platforms, and continuous DevSecOps scanning pipelines.
Pricing
Understanding the cost per million tokens is critical when deploying frontier MoE models in production. GLM-5.3 is available through Serverless Inference with transparent, metered per-token billing and zero idle compute charges. Teams pay strictly for the input, cached input, and output tokens consumed by their applications.
| Token Type | Price per 1 Million Tokens | Billing Mechanism |
|---|---|---|
| Input Tokens | $1.40 | Standard un-cached prompt tokens processed |
| Cached Input Tokens | $0.26 | Pre-cached prompt prefix tokens, billed well below the standard input rate |
| Output Tokens | $4.40 | Tokens generated during reasoning and response completion |
Prompt caching provides significant cost savings for agentic architectures. When executing multi-turn agent loops where repository context or system instructions persist across calls, repeated prompt prefixes are billed at $0.26 per million tokens instead of the $1.40 standard input rate.
Regarding hosting location and data residency: residency is a per-model fact, and GLM-5.3 currently carries no specific EU-hosted region claim or sovereign hosting designation. For compliance tracking, engineering teams should evaluate each model's verified specification record individually.
How to call it
Deploying GLM-5.3 requires no custom SDKs or specialized orchestration wrappers. Serverless Inference exposes an OpenAI-compatible API endpoint, so teams integrate the model by pointing standard client libraries at the serverless base URL and passing the official API model string: z-ai/glm-5.3.
Because GLM-5.3 enforces active thinking mode by default, requests must configure the reasoning_effort parameter. The API accepts three effort levels: low, high, and max. For complex coding and vulnerability detection tasks, setting reasoning_effort to max is strongly recommended to allocate sufficient reasoning depth before output generation.
- Authentication: Obtain your API key (prefixed with lk_) from the console.
- Base URL configuration: Point your OpenAI client or HTTP client at the serverless base URL published in the docs.
- Model selection: Set the model parameter to z-ai/glm-5.3.
- Reasoning configuration: Include the reasoning_effort parameter (low, high, or max) within your payload based on task requirements.
Below is an integration example in Python demonstrating how to initialize the client and invoke GLM-5.3 with max reasoning effort:
- Base URL: https://api.lyceum.technology/openai/v1
- Model String: z-ai/glm-5.3
- Header: Authorization: Bearer $LYCEUM_API_KEY
- Payload Parameters: model, messages, and reasoning_effort
Where it fits in the Lyceum catalogue
Within the broader model catalogue, GLM-5.3 serves as a dedicated flagship option for complex reasoning, autonomous engineering agents, and deep code analysis. While models such as GLM-5.3 Flash ($0.20 input / $0.50 output) or DeepSeek V4 Flash ($0.15 input / $0.30 output) excel at high-volume, low-latency triage tasks, GLM-5.3 is purpose-built for heavy problem-solving where multi-turn accuracy and environment execution matter most.
| Model Offering | Target Workload | Input Price / 1M | Output Price / 1M | Context Window |
|---|---|---|---|---|
| GLM-5.3 | Complex agentic coding, vulnerability discovery, deep reasoning | $1.40 | $4.40 | 1M tokens |
| GLM-5.2 | General bilingual reasoning, long-context document analysis | $1.50 | $4.50 | 1M tokens |
| GLM-5.3 Flash | High-throughput text generation, lightweight coding assistance | $0.20 | $0.50 | 1M tokens |
| DeepSeek V4 Flash | Cost-sensitive tool calling, classification, and summarization | $0.15 | $0.30 | 1M tokens |
For AI-native product companies designing production-grade coding assistants or automated security platforms, running GLM-5.3 on Serverless Inference eliminates the operational overhead of provisioning dedicated GPU clusters. You get instantaneous access to pre-hosted vLLM infrastructure, sub-second time-to-first-token, and pay-per-token pricing with no minimum commitments or base fees.
- Zero idle compute costs: Metered per-token pricing ensures you never pay for unutilized GPU memory between agent execution runs.
- OpenAI SDK compatibility: Swap your existing base URL and model parameter to z-ai/glm-5.3 without refactoring production application logic.
- Immediate scalability: Send requests directly to the Serverless Inference endpoint to benchmark GLM-5.3 against your existing evaluation datasets today.
Running GLM-5.3 in Claude Code
Lyceum exposes an Anthropic-compatible endpoint alongside the OpenAI-compatible one, so Claude Code can be pointed at GLM-5.3 without a separate adapter. The Lyceum CLI does the setup for you: lyceum code launches Claude Code preconfigured in its own profile, which leaves an existing Anthropic login untouched. To wire it up by hand, set three environment variables and start Claude Code:
export ANTHROPIC_BASE_URL=https://api.lyceum.technology/anthropic
export ANTHROPIC_AUTH_TOKEN=lk_your_api_key_here
export ANTHROPIC_MODEL=z-ai/glm-5.3
claudeUse ANTHROPIC_AUTH_TOKEN rather than ANTHROPIC_API_KEY. Both authenticate, but ANTHROPIC_API_KEY makes Claude Code ask for approval of a custom key on first run, which fails outright in non-interactive use such as claude -p or CI with "Not logged in - Please run /login". The tier variables ANTHROPIC_DEFAULT_OPUS_MODEL, ANTHROPIC_DEFAULT_SONNET_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL map the Claude tier names in the /model menu onto Lyceum models, so picking a tier selects the model you mapped to it.