Run open models in production

For teams building AI products: call open models through one OpenAI‑compatible API, billed per token, or reserve a dedicated endpoint.

Serverless inference

Call any open model in our catalogue through one OpenAI‑compatible API. Billed per token, with no base fee.

Prompt caching

Repeated prompt starts still in GPU memory are billed at the cached input price, where listed.

Your first request
request.sh
BASE_URL=https://api.lyceum.technology/openai/v1

curl $BASE_URL/chat/completions \
  -H "Authorization: Bearer $LYCEUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-5.3",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Example usage: 1,000 input and 500 output tokens cost $0.0036.

  1. Base URL
  2. Runs in European data centres.
  3. OpenAI‑compatible, and Anthropic‑compatible for Claude Code
  4. Any of 23 open models
  5. Billed per token, no base feeGLM-5.3: $1.40 input, $4.40 output per million tokens

Dedicated endpoints

  • Reserved capacity

    GPUs reserved for your endpoint, planned around your production workload.

  • Model configuration

    Agreed before launch, so your model fits your workload.

  • Where it runs

    The data centre is agreed for your workload. Tell us if you need a specific EU location.

  • Commercial terms

    Commitment pricing, agreed per contract, so you know the cost before you launch.

Which option fits your workload

Per-token billing follows your traffic, so quiet hours cost nothing. A dedicated endpoint is priced for the capacity you reserve, so the busier you keep it, the less each token costs.

Need the whole machine?

Explore Compute

Per token costs lessTraffic comes and goes: stay on the API.

Dedicated costs lessSteady all day: ask us to price it against your token bill.

Bring the tools you already use

Serverless or dedicated, you call it through the same OpenAI‑compatible API. Create a key in your dashboard, then point Claude Code, OpenCode or your own app at it.

Claude Code

Point Claude Code at our Anthropic-compatible endpoint and choose a supported open model.

Terminal
export ANTHROPIC_BASE_URL="https://api.lyceum.technology/anthropic"
export ANTHROPIC_AUTH_TOKEN="$LYCEUM_API_KEY"
export ANTHROPIC_MODEL="z-ai/glm-5.3"

claude

Install Claude Code and set LYCEUM_API_KEY

my-repo
  • main.py2 KB
  • requirements.txt1 KB
  • README.md1 KB
  • .envnice try

4 objects

TODO.txt
  • [x] switch base URL
  • [ ] tell the team
  • [ ] stop refreshing the invoice

OpenCode

Add Lyceum as a custom OpenAI-compatible provider, with your endpoint and model.

opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "model": "lyceum/z-ai/glm-5.3",
  "provider": {
    "lyceum": {
      "npm": "@ai-sdk/openai-compatible",
      "options": {
        "baseURL": "https://api.lyceum.technology/openai/v1",
        "apiKey": "{env:LYCEUM_API_KEY}"
      },
      "models": {
        "z-ai/glm-5.3": { "name": "GLM-5.3" }
      }
    }
  }
}

Save in your project and set LYCEUM_API_KEY

backup_final_v2
  • backup_final_v3_REAL
  • backup_final_v3_REAL_2
  • use_this_one
  • opencode.json

Found it, 3 folders down

configs
Contains
opencode.json
Size
372 bytes
Fits on a floppy
3,963 times

1 object

Your application

Use any OpenAI-compatible client. Set the base URL, your API key and a model ID.

app.py
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LYCEUM_API_KEY"],
    base_url="https://api.lyceum.technology/openai/v1",
)

response = client.chat.completions.create(
    model="z-ai/glm-5.3",
    messages=[{"role": "user", "content": "Hello!"}],
)

print(response.choices[0].message.content)

Install openai and set LYCEUM_API_KEY

My GPUs

8 × H200, all yours.

One VM, billed per second.

See our GPUs

Fan noise: not your problem

README.txt
  1. 1. Create an API key.
  2. 2. Set the base URL.
  3. 3. Pick a model.

There is no step 4

Your data

Prompts and outputs are processed, not stored, and never used for training. GPUs run in European data centres.

How we handle your data

Your first request is one endpoint away