Pay for what you run
Inference is billed per token with no base fee. GPUs are billed per second with no subscription.
Inference
USD per million tokens: input, cached input and output, billed per token with no base fee
Prompt caching
Repeated prompt starts still in GPU memory are billed at the cached input price, where listed.
14 models
ModelInputCached inputOutput
Qwen3.8 Flash NextLyceum Simple and Qwen3.8 Flash Next Instant are billed as Qwen3.8 Flash Next.$0.20$0.05$0.50
Sorted by input price, cheapest first, on a logarithmic scale.
GPU compute
USD per GPU hour, billed per second with no subscription
GPUMemoryOn demandSpotAvailable nowReserved blocks
More ways to run
Capacity agreed with you, or training without a machine to manage