Pay for what you run

Inference is billed per token with no base fee. GPUs are billed per second with no subscription.

Inference

USD per million tokens: input, cached input and output, billed per token with no base fee

How serverless inference works

Prompt caching

Repeated prompt starts still in GPU memory are billed at the cached input price, where listed.

14 models

ModelInputCached inputOutput
QwenQwen3 Embedding 8B$0.02Not offeredInput only
Z.aiGLM-5.3 FlashGLM-5.3 Flash Instant is billed as GLM-5.3 Flash.$0.20$0.05$0.50
QwenQwen3.8 Flash NextLyceum Simple and Qwen3.8 Flash Next Instant are billed as Qwen3.8 Flash Next.$0.20$0.05$0.50
DeepSeekDeepSeek V4 Flash$0.25$0.06$0.30
MiniMaxMiniMax M3Lyceum Complex and Lyceum Router are billed as MiniMax M3.$0.40$0.10$2.00
QwenQwen3.8 27BQwen3.8 27B Instant is billed as Qwen3.8 27B.$0.40$0.10$2.40
DeepSeekDeepSeek V4.1 Flash$0.50$0.13$1.50
MoonshotKimi K2.7 Code$1.25$0.31$4.50
Z.aiGLM-5.3GLM-5.3 Instant is billed as GLM-5.3.$1.40$0.26$4.40
Z.aiGLM-5.2GLM-5.2 Instant and Lyceum Reasoning are billed as GLM-5.2.$1.50$0.38$4.50
DeepSeekdeepseek/deepseek-v4-pro$1.75$0.44$3.50
DeepSeekDeepSeek V4 Pro 0813$2.00$0.50$4.00
QwenQwen3.8 2.4T A95B$2.50$0.63$6.00
MoonshotKimi K3$3.00$0.75$15.00

Sorted by input price, cheapest first, on a logarithmic scale.

GPU compute

USD per GPU hour, billed per second with no subscription

Discuss your workload
GPUMemoryOn demandSpotAvailable nowReserved blocks
NVIDIA B300288 GB$7.99$2.99Check dashboard ↗On request
NVIDIA B200192 GB$6.49$2.40Check dashboard ↗On request
NVIDIA H200141 GB$4.29$1.60Check dashboard ↗On request
NVIDIA H10080 GB$2.79$1.10Check dashboard ↗On request
NVIDIA A10080 GB$1.59$0.80Check dashboard ↗On request
NVIDIA L40S48 GB$1.19Not offeredCheck dashboard ↗On request

More ways to run

Capacity agreed with you, or training without a machine to manage

Dedicated endpoints

GPUs reserved behind your own endpoint. Commitment pricing, agreed per contract.

See dedicated endpoints

Serverless training

Your training code in a container, on the GPU you choose. Check its price before you launch.

See serverless training

Reserved GPUs and clusters

GPUs or a whole cluster held for your workload. Commitment pricing, agreed per contract.

See reserved GPUs and clusters

Your next workload starts here