33 models, prompt caching, and more H100s
Lyceum

ISSUE #003 · JULY 2026

33 models, prompt caching, and more H100s

The catalog just grew to 33 models, including DeepSeek V4 Pro, GLM-5.2, and Kimi K2.7 Code. Prompt caching now cuts input costs by up to 83%, and more H100 nodes are available.

The Catalog Keeps Growing

When we launched Serverless Inference, the catalog started with models like Llama 3.3 70B, DeepSeek V3.2, Qwen3 235B, and Kimi K2. Since then, a new generation of open-source models has arrived, and we’ve been adding them as they drop.

Here is what is new on the platform:

DeepSeek V4 Pro: DeepSeek’s latest flagship for reasoning, coding, and long-horizon agent workflows, with a 1M token context window. Input at $1.75/1M, output at $3.50/1M.

GLM-5.2: ZAI’s newest flagship with strong bilingual reasoning, long-context understanding, and tool use. 1M token context, input at $1.50/1M, output at $4.50/1M.

Kimi K2.7 Code: Moonshot’s coding-focused agentic model with strong tool use and 256K context. Input at $1.25/1M, output at $4.50/1M.

MiniMax M3: 1M token context with reasoning and tool use, built for large-document and agentic workloads. Input at $0.40/1M, output at $2.00/1M.

Qwen3.5 397B and Qwen3.5 9B: Alibaba’s largest Qwen3.5 MoE model for complex reasoning, and a compact 9B model for fast, low-cost tasks starting at $0.15/1M input.

Also new: Kimi K2.5 and K2.6, GLM-5 and GLM-5.1, Qwen3 Coder 30B, MiniMax M2.5, NVIDIA’s Nemotron 3 family, and INTELLECT-3. All accessible through one OpenAI-compatible API.

Under the hood, smart routing directs every request to the best available capacity for your model automatically. You send the request, we handle where it runs.

Browse all models Get started

Prompt Caching Is Live

If your requests reuse the same context, a long system prompt, a document, a codebase, a tool definition, you have been paying full price for those tokens on every call. Not anymore.

Prompt caching is now live on selected models. Repeated input tokens are cached and billed at a fraction of the standard rate:

Model Input Cached input Savings
GLM-5.2 $1.50/1M $0.38/1M 75%
Kimi K2.7 Code $1.25/1M $0.31/1M 75%
MiniMax M3 $0.40/1M $0.10/1M 75%
Qwen3.5 9B $0.15/1M $0.04/1M 73%
Qwen3 Coder 30B $0.06/1M $0.01/1M 83%

Caching works automatically, no code changes required. The biggest wins go to agentic workloads, RAG pipelines, and coding assistants, where the same context gets sent over and over. On top of lower costs, cached requests also come back faster.

Read the docs

More H100 Nodes Available

Our GPU fleet already includes hundreds of H100s, and we’re expanding it further. Additional dedicated H100 nodes are now available on commitment: EU-based, with Kubernetes included. Part of the capacity is available immediately, further nodes come online at the beginning of August.

Demand across the fleet remains high. The B200 nodes we recently added were fully committed shortly after going live, and B300 capacity is planned for next quarter.

If you’re planning a training run, a product launch, or steady production workloads and need guaranteed H100 capacity, reply to this email or book a call below and we’ll walk you through pricing and availability.

I’m interested, let’s talk capacity

Not sure what GPU shape your workload needs?

Book a 30-minute validated-infrastructure assessment with our GPU solutions engineer. You walk away with a specific recommendation, not a pitch.

Book your assessment

Lyceum Technology, Berlin, Germany

Unsubscribe

Keep reading