Coming from Kimi K2.6
Lyceum retired Kimi K2.6 on 5 October 2026. Requests that still send the model id moonshotai/kimi-k2.6 are now routed to Kimi K3 and billed at Kimi K3 rates: $3.00 per million input tokens, $0.75 cached and $15.00 output, against the $1.00, $0.25 and $4.00 that Kimi K2.6 cost. Change the model id in your code to moonshotai/kimi-k3, so your logs and cost reports name the model you actually run, and recheck your budget for output-heavy jobs.
Two things work differently. Kimi K3 accepts text and images but not PDFs, so a PDF sent on the old id now returns an error. And if your Kimi K2.6 traffic was coding-only, the cheaper Moonshot option on Lyceum is Kimi K2.7 Code (moonshotai/kimi-k2.7-code) at $1.25 input, $0.31 cached and $4.50 output per million tokens, with a 256K-token context; it also still reads PDFs. Keep Kimi K3 for work that needs the 1M-token window, or for agentic and knowledge work beyond code. Switching reasoning off with reasoning_effort set to "none" works on Kimi K3 as it did on Kimi K2.6.
The 1M-Token Context Shift in Production Inference
Engineering teams building production LLM applications spend substantial engineering cycles maintaining Retrieval-Augmented Generation (RAG) pipelines. Chunking source code, synchronizing vector databases, and tuning embedding retrieval thresholds can fail when queries require full repository context or multi-document cross-referencing. When a model context window expands to 1M tokens, developers can send a whole codebase or documentation set that fits in the window in a single inference pass instead of splitting them across a vector index.
Deploying 1M-token context models introduces specific operational trade-offs in inference execution. Moonshot's Kimi K3 provides a 1M-token context window for text and image input, and Lyceum lists it as EU-hosted (eu-north1). Processing full-context prompts avoids missing context boundaries, but raw input processing costs scale linearly with prompt length. At $3.00 per million input tokens and $15.00 per million output tokens, unoptimized long-context requests quickly dominate compute expenditure for AI-native platforms.
- Pipeline complexity: Can remove vector database maintenance, embedding re-indexing, and chunk overlap tuning for corpora that fit in the window.
- Time-to-first-token latency: Prefill phase execution time increases with sequence length, requiring high-throughput attention kernels.
- Cost per call: Ingesting a large codebase scales directly on base input rates, significantly exceeding traditional vector retrieval costs.
To make 1M-token context economical for agentic workflows and multi-turn developer tools, Lyceum Serverless Inference caches repeated prompt prefixes, so production systems can reuse static codebase context without re-paying the full prefill cost on every turn. Cached prompts are never written to a database, so zero data retention still holds on Lyceum's serverless capacity in Paris and Finland.
Kimi K3 Technical Profile: 2.8T MoE and 1M Context Window
Moonshot AI's Kimi K3 architecture uses a sparse Mixture-of-Experts (MoE) design containing 2.8 trillion total parameters. The published Kimi K3 model card, checked on 5 October 2026, lists 2.8T total parameters, 104B activated per token, 16 of 896 experts routed per token, and a 1,048,576-token context window. During inference, the router sends each token to 16 of the 896 experts, plus 2 shared experts, so about 104 billion parameters are active per token, a figure that also covers attention and the model's one dense layer. For AI-native product teams, this means far less compute per token than a dense model of the same total size.
Architectural Specifications and Context Limits
- Total parameter count: 2.8 trillion, in a sparse Mixture-of-Experts design with 896 experts.
- Active parameters: 104 billion parameters activated per token, with 16 of 896 experts routed per token.
- Context window limit: 1,048,576 tokens (1M context) for long-sequence tasks.
- Regional hosting: Listed by Lyceum as EU-hosted (eu-north1).
- Token pricing model: $3.00 per 1M input tokens, $15.00 per 1M output tokens, and $0.75 per 1M cached input tokens.
- Input types on Lyceum: Text and images. PDF input is rejected on Kimi K3.
- License: Open weights under the Kimi K3 License, which adds conditions for large commercial deployments.
Processing full 1M-token context windows presents substantial memory management challenges, particularly around Key-Value (KV) cache allocation across long conversation threads or codebase analysis. On Lyceum Serverless Inference, Kimi K3 runs on Lyceum's serverless capacity in Paris and Finland, so you do not provision that GPU memory yourself.
Running a 1M-token MoE model yourself means sizing GPU memory for the weights plus the KV cache of every concurrent request. On serverless inference that sizing is Lyceum's job: you are billed per token, and per-model latency and throughput are visible on Lyceum's public status page.
Kimi K3 Rate Card: $3.00 Input and $15.00 Output Breakdown
At $3.00 per million input tokens and $15.00 per million output tokens, Kimi K3 represents the top tier of open-weights pricing on Lyceum Serverless Inference. For AI-native product companies running long-horizon autonomous agents, multi-document analysis, or dense codebase processing across a 1M-token context window, calculating unit economics requires evaluating both baseline token rates and dynamic caching performance.
| Model / Provider Tier | Context Window | Input (per 1M) | Cached Input (per 1M) | Output (per 1M) |
|---|---|---|---|---|
| Kimi K3 (Lyceum, EU-hosted) | 1M tokens | $3.00 | $0.75 | $15.00 |
| Kimi K2.7 Code (Lyceum, EU-hosted) | 256K tokens | $1.25 | $0.31 | $4.50 |
| Claude Fable 5.1 (Anthropic API) | 1M tokens | $10.00 | $0.25 (cache hit) | $50.00 |
Although Kimi K3 is the highest per-token entry in Lyceum's open model catalogue, it is not the most expensive option on the market. Anthropic's models overview lists Claude Fable 5.1 at $10.00 input and $50.00 output per million tokens and Claude Opus 5.5 at $4.00 and $20.00, and OpenAI's API pricing page lists gpt-6-astra at $10.00 and $50.00 (USD, standard tier, short context; both read 5 October 2026). Not every closed model costs more: the same OpenAI page lists gpt-6.1-sol at $2.00 input and $10.00 output. Cached input is where the comparison turns the other way. Kimi K3 bills cache hits at $0.75 per million tokens, while Anthropic lists $0.25 for Claude Fable 5.1 and $0.20 for Claude Opus 5.5, plus a separate charge for cache writes. Compare the blended cost on your own traffic mix; our Kimi K3 vs Claude Fable 5 and GLM-5.2 vs Claude Opus 5 comparisons, written for the previous Claude generation, go deeper.
On Lyceum, Kimi K3 runs on Serverless Inference as moonshotai/kimi-k3, billed per token with no base fee, no dedicated GPU node allocation and no long-term contract. Lyceum lists the model as EU-hosted (eu-north1).
Prompt Caching Economics: Lowering Effective Input Costs to $0.75
At $3.00 per million input tokens, filling Kimi K3's full 1M-token context window for every request scales API spend rapidly. AI-native product teams building agentic loops or processing static document corpora rarely send entirely fresh context on every turn.
KV Cache Reuse in Multi-Turn Workloads
Prompt caching on Lyceum is automatic, with nothing to enable. When consecutive requests share an identical leading prefix, such as a fixed system prompt, agent tool definitions or earlier conversation turns, the cached portion is reused and billed at the cached input rate. Keep the stable part of the prompt at the front and vary only the tail. Caching is best-effort: a request may or may not hit a warm cache depending on recent traffic, and the usage object reports how many prompt tokens came from the cache. Cached prompts sit in GPU memory only, per session, for a few minutes at most, and are never written to a database, which is what keeps the zero-retention policy intact.
- 75% lower rate on cache hits: Prompt tokens served from the cache bill at $0.75 instead of $3.00 per million tokens; uncached tokens and cache misses still pay the full rate.
- Latency and throughput benefits: Reusing cached prefill state cuts redundant prefix processing across multi-turn agent execution and long-document queries.
In multi-turn agent workflows, the full $3.00 rate applies to the initial prefill; later turns that hit the cache bill the shared prefix at $0.75 per million, and only the new tail of each turn pays the full rate. Illustrative arithmetic, not a measurement: a turn with 900,000 input and 10,000 output tokens costs $2.70 + $0.15 = $2.85 uncached. If 880,000 of those input tokens hit the cache, it costs $0.66 + $0.06 + $0.15 = $0.87. A cache hit is not guaranteed.
Comparing 1M-Context Models in the EU: Kimi K3 vs GLM-5.2 and MiniMax-M3
When building production agents or repository-scale analysis workflows, selecting a million-token model requires balancing raw capability against token expenditure. In the Lyceum model catalogue, several 1M-context models run in the EU (eu-north1), among them Kimi K3, GLM-5.2, and MiniMax-M3. Kimi K3 has the highest output rate in the catalogue at $15.00 per million tokens. For engineering teams evaluating high-frequency inference, understanding when to deploy Kimi K3 versus lighter 1M-context alternatives directly dictates unit economics. For coding-only work, Kimi K2.7 Code is the cheaper Moonshot model on Lyceum.
| Model | Model ID | Provider | Input Rate (Uncached) | Input Rate (Cached) | Output Rate |
|---|---|---|---|---|---|
| Kimi K3 | moonshotai/kimi-k3 | Moonshot | $3.00 / 1M | $0.75 / 1M | $15.00 / 1M |
| GLM-5.2 | z-ai/glm-5.2 | ZAI | $1.50 / 1M | $0.38 / 1M | $4.50 / 1M |
| MiniMax-M3 | minimax/minimax-m3 | MiniMax | $0.40 / 1M | $0.10 / 1M | $2.00 / 1M |
Workload Profiles and Cost Optimization
The substantial output price variance between Kimi K3 ($15.00/1M) and GLM-5.2 ($4.50/1M) shapes workload placement across your stack. Kimi K3's higher rate pays off where the job needs its 1M-token window or long-horizon agentic and knowledge work; for coding-only traffic within 256K tokens, Kimi K2.7 Code ($4.50/1M output) is the cheaper Moonshot option. For document extraction or long-form output pipelines, MiniMax-M3 ($2.00/1M output) and GLM-5.2 deliver substantially lower operational cost.
On Lyceum Serverless Inference, all three are listed as EU-hosted (eu-north1) and run with zero data retention: prompts and outputs are not stored or used for training.
Data Handling and Hosting Region for Long-Context Workloads
Feeding a 1M-token context window into an LLM can mean sending several megabytes of text per call, including full source code repositories and sensitive financial records. Before payloads like that go to any endpoint, European product teams usually need two answers: where the model processes them, and whether anything is kept afterwards.
Lyceum lists Kimi K3 as EU-hosted (eu-north1), and Lyceum's serverless models run in Paris and Finland. Prompts and outputs are processed but not stored, and never used for training. The DPA is part of Lyceum's terms. Data centre operators hold ISO certifications at facility level; those certifications belong to the operators, not to Lyceum.
- Hosting stated per model: Lyceum lists Kimi K3 as EU-hosted (eu-north1). Check the processing region of every other model you route to.
- No retention: Prompts and outputs are processed but not stored, and never used for training.
- Transient in-memory caching: Cached prompts stay in GPU memory only and are never written to a database.
An EU-hosted model with no prompt retention lets technical leads answer both questions before agentic workflows and document analytics reach production. For the GPU memory side of long contexts, see our guide to long-context inference GPU requirements.
Integrating Kimi K3 via OpenAI-Compatible Serverless Inference
Lyceum serves Kimi K3 through an OpenAI-compatible API, so an existing OpenAI SDK client needs three settings: the base URL shown in your Lyceum dashboard, a Lyceum API key and the model id moonshotai/kimi-k3. Test tools, output formats and limits on your own workload before you move production traffic. The example below uses the placeholder key lk_your_api_key_here.
from openai import OpenAI
client = OpenAI(
base_url="https://api.lyceum.technology/openai/v1",
api_key="lk_your_api_key_here",
)
response = client.chat.completions.create(
model="moonshotai/kimi-k3",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)Use Kimi K3 with Claude Code
Lyceum also exposes an Anthropic-compatible endpoint for Claude Code. Lyceum's recommended route is the Lyceum CLI, where lyceum code launches Claude Code preconfigured in its own profile. To try Kimi K3 in a single terminal without installing anything, set the variables below, replace the placeholder with your Lyceum key and start Claude Code. Use ANTHROPIC_AUTH_TOKEN rather than ANTHROPIC_API_KEY, which makes Claude Code ask for approval of a custom key on first run. Claude Code's tier names map to other Lyceum models by default (Fable and Opus to GLM-5.2, Sonnet to Kimi K2.7 Code, Haiku to MiniMax-M3), so select Kimi K3 explicitly: with ANTHROPIC_MODEL, or with /model moonshotai/kimi-k3 in a session that already runs on Lyceum, which switches from the next message. Do not add the Anthropic-only [1m] suffix; Claude Code will not connect through Lyceum with it.
export ANTHROPIC_BASE_URL="https://api.lyceum.technology/anthropic"
export ANTHROPIC_AUTH_TOKEN="lk_your_api_key_here"
export ANTHROPIC_MODEL="moonshotai/kimi-k3"
claudeStreaming long Kimi K3 responses
For real-time applications such as interactive coding agents or document analysis workflows, streaming output over Server-Sent Events (SSE) delivers each token to the client as it is produced, so the application can render output while generation is still running rather than waiting for the whole completion. Streaming also protects long requests: Lyceum allows a single request up to 300 seconds of generation time, and a long non-streaming call can hit a client-side timeout before the response finishes.
- Base URL endpoint: Point your SDK client at the base URL shown in your Lyceum dashboard.
- Authentication: Pass your Lyceum API key as a Bearer token in the authorization header.
- Model identifier: moonshotai/kimi-k3.
- Response streaming: Set stream=True to receive continuous SSE delta chunks as tokens are generated.
Because the API follows the OpenAI format, the same client code can later point at another OpenAI-compatible provider, which helps you avoid vendor lock-in to a proprietary SDK. You keep control over prompt formatting, sampling parameters and context management; model-specific options such as reasoning controls may differ between providers.
Production limits for Kimi K3 on Lyceum
Lyceum's Serverless Inference documentation sets these limits and behaviors for moonshotai/kimi-k3 (checked 5 October 2026). Input, tool definitions, conversation history and output, including reasoning, all count against the 1M-token context, so leave room for the answer; our guide to LLM context length and GPU memory requirements covers the memory side, and our explainer on serverless GPU inference covers the billing model.
- Output: max_tokens is capped at 65,536 per request; larger values are rejected with a 400 error.
- Generation time: Up to 300 seconds per request. Stream long outputs so a client-side timeout does not cut them off first.
- Context window: 1M tokens. A prompt larger than the window is rejected with a 400 error.
- Reasoning: On by default, and the trace is billed as output tokens, so set max_tokens high enough for reasoning plus answer. Switch it off with reasoning_effort set to "none" or chat_template_kwargs {"thinking": false}. Kimi K3 also accepts low, medium, high and max, but Lyceum's docs advise treating reasoning as on or off, not as a depth dial.
- Tool calling: Kimi K3 is among the models Lyceum lists as honoring tool_choice set to "required" and a named function.
- Images and PDFs: Kimi K3 accepts images as OpenAI-style image_url content parts. It rejects PDF parts; Kimi K2.7 Code reads PDFs.
- Prompt caching: Automatic and best-effort. usage.prompt_tokens_details.cached_tokens reports the cached portion, and the field is absent on a cache miss.
- Multi-turn reasoning: Moonshot AI's model card says Kimi K3 expects the complete assistant message, including reasoning_content and tool_calls, to be passed back in later turns.