What's new

ZAI has officially released GLM-5.3, a frontier Mixture-of-Experts (MoE) model engineered specifically for complex software engineering, multi-turn agentic workflows, and vulnerability analysis. Rather than altering the underlying parameter count or pre-training dataset from scratch, GLM-5.3 retains the exact same base model architecture as GLM-5.2. Every measured performance gain across coding benchmarks and multi-step reasoning tasks stems exclusively from scaled post-training pipelines.

The core engineering focus behind GLM-5.3 centers on expanding reinforcement learning (RL) environments so they resemble real units of expert work rather than isolated coding exercises. ZAI says a useful task environment has to be executable, verifiable and close to real professional work, and that it built pipelines to synthesize such environments end to end, plus the RL reward signal for a subset of tasks. That matters for product teams building autonomous coding tools, refactoring engines or security auditing agents.

  • Architecture parity with GLM-5.2: the same base model is used, so every reported gain comes from post-training.
  • Scaled RL infrastructure: built on the GLM-5.2 stack, including slime for large-scale asynchronous training and SAO for RL on long-horizon tasks.
  • Executable developer environments: moving beyond static unit tests to dynamic environments requiring compiler diagnostics, dependency resolution, and multi-turn shell interactions.
  • Mandatory thinking mode: reasoning tokens are now active by default, replacing binary toggles with configurable effort levels to match target compute budgets.

Specs and context window

Under the hood, GLM-5.3 is a sparse Mixture-of-Experts (MoE) model, so only a fraction of its parameters are active per token. On Lyceum it is served on vLLM, the same serving stack used across the supported-models catalogue.

A primary strength of GLM-5.3 is its native 1M-token context window. That much context lets developers ingest entire enterprise repositories, complex system architecture documents, or exhaustive technical specifications in a single prompt. It eliminates brittle retrieval fragmentation when building code-review agents or complex technical summarization tools.

Architectural ParameterSpecificationEngineering Impact
ArchitectureMixture-of-Experts (MoE)Deep parametric capacity with sparse per-token compute
Base modelSame base model as GLM-5.2All reported gains come from post-training, not new pre-training
Context Length1M tokensIngests massive codebases, logs, and technical documentation in one pass
Serving EnginevLLMStandardized tensor parallelism, continuous batching, and PagedAttention
Reasoning ControlConfigurable reasoning effortReasoning token budget can be matched to task complexity

Long-context serving builds on IndexShare, the efficient long-context processing component ZAI introduced with GLM-5.2 and carried into GLM-5.3, together with SAO for RL on long-horizon tasks.

Coding and agentic benchmarks

In vendor-reported benchmarks published by ZAI, GLM-5.3 demonstrates substantial generational gains over GLM-5.2, particularly when tasked with multi-step agentic coding workflows. On ZAI's internal Z.ai Code Bench, GLM-5.3 recorded a 50% relative improvement over GLM-5.2. This internal benchmark measures end-to-end task completion and fine-grained checklist accuracy in isolated local developer setups rather than grading snippet-level syntax.

Public evaluations reflect a similar leap in capabilities. On Terminal Bench 3.0, which evaluates interactive shell commands, environment navigation, and multi-step terminal workflows, GLM-5.3 achieved a score of 28.3, jumping from the 4.6 recorded by GLM-5.2. This makes GLM-5.3 the leading open-weights model on this evaluation, outperforming Kimi K3 (17.4) and approaching closed frontier models.

BenchmarkGLM-5.3 (Vendor-Reported)GLM-5.2 (Vendor-Reported)Kimi K3 (Vendor-Reported)Metric Definition
Terminal Bench 3.028.34.617.4Interactive terminal operations and tool execution
DeepSWE v1.166.946.267.5Multi-file software engineering bug resolution
Z.ai Code Bench (Max Effort)34.5%23.4%Not ReportedEnd-to-end task completion in real developer environments
PostTrainBench39.831.732.0Long-horizon RL and optimization reasoning
Agents' Last Exam28.523.8Not ReportedComplex multi-disciplinary agentic reasoning

Token efficiency improved alongside accuracy: ZAI reports markedly stronger agentic coding results than GLM-5.2 at every effort level while consuming fewer output tokens on Z.ai Code Bench. For production systems, shorter outputs directly lower per-task latency and execution cost.

Emergent cyber capability

A notable development resulting from ZAI's post-training expansion is the emergence of advanced cybersecurity reasoning. By incorporating vulnerability analysis environments and exploit verification data into the reinforcement learning loop, the model began autonomously synthesizing complete, multi-stage exploitation sequences across complex architectures.

On CyberGym, a white-box source code evaluation that tests whether an AI can identify and trigger software faults, GLM-5.3 achieved a vendor-reported score of 84.5 (source: zai). ZAI's own comparison table puts that above GLM-5.2 (77.2), Kimi K3 (80.0) and GPT-5.6 Sol (83.6) on the same benchmark. On ExploitBench, which demands deeper reasoning over real-world Common Vulnerabilities and Exposures (CVEs), GLM-5.3 scored 54.4, more than doubling GLM-5.2's 24.4.

  • CyberGym fault triggering: 84.5 vendor-reported score on white-box source code vulnerability analysis, state of the art on that benchmark.
  • ExploitBench reasoning: 54.4 on multi-stage exploit planning, more than doubling GLM-5.2's 24.4.
  • Gains up the exploitation chain: ZAI reports its largest improvements further along the exploitation chain, including on ExploitGym.
  • Staged weight release: ZAI said it would publish the open weights two weeks after launch, once safety evaluation and hardening were complete.

These capabilities make GLM-5.3 a potent foundation model for automated penetration testing tools, static code analysis platforms, and continuous DevSecOps scanning pipelines.

Pricing

Understanding the cost per million tokens is critical when deploying frontier MoE models in production. GLM-5.3 is available through Serverless Inference with transparent, metered per-token billing and zero idle compute charges. Teams pay strictly for the input, cached input, and output tokens consumed by their applications.

Token TypePrice per 1 Million TokensBilling Mechanism
Input Tokens$1.40Standard un-cached prompt tokens processed
Cached Input Tokens$0.26Pre-cached prompt prefix tokens, billed well below the standard input rate
Output Tokens$4.40Tokens generated during reasoning and response completion

Prompt caching provides significant cost savings for agentic architectures. When executing multi-turn agent loops where repository context or system instructions persist across calls, repeated prompt prefixes are billed at $0.26 per million tokens instead of the $1.40 standard input rate.

Regarding hosting location and data residency: residency is a per-model fact, and GLM-5.3 currently carries no specific EU-hosted region claim or sovereign hosting designation. For compliance tracking, engineering teams should evaluate each model's verified specification record individually.

How to call it

Deploying GLM-5.3 requires no custom SDKs or specialized orchestration wrappers. Serverless Inference exposes an OpenAI-compatible API endpoint, so teams integrate the model by pointing standard client libraries at the serverless base URL and passing the official API model string: z-ai/glm-5.3.

Because GLM-5.3 enforces active thinking mode by default, requests must configure the reasoning_effort parameter. The API accepts three effort levels: low, high, and max. For complex coding and vulnerability detection tasks, setting reasoning_effort to max is strongly recommended to allocate sufficient reasoning depth before output generation.

  1. Authentication: Obtain your API key (prefixed with lk_) from the console.
  2. Base URL configuration: Point your OpenAI client or HTTP client at the serverless base URL published in the docs.
  3. Model selection: Set the model parameter to z-ai/glm-5.3.
  4. Reasoning configuration: Include the reasoning_effort parameter (low, high, or max) within your payload based on task requirements.

Below is an integration example in Python demonstrating how to initialize the client and invoke GLM-5.3 with max reasoning effort:

  • Base URL: https://api.lyceum.technology/openai/v1
  • Model String: z-ai/glm-5.3
  • Header: Authorization: Bearer $LYCEUM_API_KEY
  • Payload Parameters: model, messages, and reasoning_effort

Where it fits in the Lyceum catalogue

Within the broader model catalogue, GLM-5.3 serves as a dedicated flagship option for complex reasoning, autonomous engineering agents, and deep code analysis. While models such as GLM-5.3 Flash ($0.20 input / $0.50 output) or DeepSeek V4 Flash ($0.15 input / $0.30 output) excel at high-volume, low-latency triage tasks, GLM-5.3 is purpose-built for heavy problem-solving where multi-turn accuracy and environment execution matter most.

Model OfferingTarget WorkloadInput Price / 1MOutput Price / 1MContext Window
GLM-5.3Complex agentic coding, vulnerability discovery, deep reasoning$1.40$4.401M tokens
GLM-5.2General bilingual reasoning, long-context document analysis$1.50$4.501M tokens
GLM-5.3 FlashHigh-throughput text generation, lightweight coding assistance$0.20$0.501M tokens
DeepSeek V4 FlashCost-sensitive tool calling, classification, and summarization$0.15$0.301M tokens

For AI-native product companies designing production-grade coding assistants or automated security platforms, running GLM-5.3 on Serverless Inference eliminates the operational overhead of provisioning dedicated GPU clusters. You get instantaneous access to pre-hosted vLLM infrastructure, sub-second time-to-first-token, and pay-per-token pricing with no minimum commitments or base fees.

  • Zero idle compute costs: Metered per-token pricing ensures you never pay for unutilized GPU memory between agent execution runs.
  • OpenAI SDK compatibility: Swap your existing base URL and model parameter to z-ai/glm-5.3 without refactoring production application logic.
  • Immediate scalability: Send requests directly to the Serverless Inference endpoint to benchmark GLM-5.3 against your existing evaluation datasets today.

Running GLM-5.3 in Claude Code

Lyceum exposes an Anthropic-compatible endpoint alongside the OpenAI-compatible one, so Claude Code can be pointed at GLM-5.3 without a separate adapter. The Lyceum CLI does the setup for you: lyceum code launches Claude Code preconfigured in its own profile, which leaves an existing Anthropic login untouched. To wire it up by hand, set three environment variables and start Claude Code:

export ANTHROPIC_BASE_URL=https://api.lyceum.technology/anthropic
export ANTHROPIC_AUTH_TOKEN=lk_your_api_key_here
export ANTHROPIC_MODEL=z-ai/glm-5.3
claude

Use ANTHROPIC_AUTH_TOKEN rather than ANTHROPIC_API_KEY. Both authenticate, but ANTHROPIC_API_KEY makes Claude Code ask for approval of a custom key on first run, which fails outright in non-interactive use such as claude -p or CI with "Not logged in - Please run /login". The tier variables ANTHROPIC_DEFAULT_OPUS_MODEL, ANTHROPIC_DEFAULT_SONNET_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL map the Claude tier names in the /model menu onto Lyceum models, so picking a tier selects the model you mapped to it.