When we launched Serverless Inference, the catalog started with models like Llama 3.3 70B, DeepSeek V3.2, Qwen3 235B, and Kimi K2. Since then, a new generation of open-source models has arrived, and we’ve been adding them as they drop.
Here is what is new on the platform:
DeepSeek V4 Pro: DeepSeek’s latest flagship for reasoning, coding, and long-horizon agent workflows, with a 1M token context window. Input at $1.75/1M, output at $3.50/1M.
GLM-5.2: ZAI’s newest flagship with strong bilingual reasoning, long-context understanding, and tool use. 1M token context, input at $1.50/1M, output at $4.50/1M.
Kimi K2.7 Code: Moonshot’s coding-focused agentic model with strong tool use and 256K context. Input at $1.25/1M, output at $4.50/1M.
MiniMax M3: 1M token context with reasoning and tool use, built for large-document and agentic workloads. Input at $0.40/1M, output at $2.00/1M.
Qwen3.5 397B and Qwen3.5 9B: Alibaba’s largest Qwen3.5 MoE model for complex reasoning, and a compact 9B model for fast, low-cost tasks starting at $0.15/1M input.
Also new: Kimi K2.5 and K2.6, GLM-5 and GLM-5.1, Qwen3 Coder 30B, MiniMax M2.5, NVIDIA’s Nemotron 3 family, and INTELLECT-3. All accessible through one OpenAI-compatible API.
Under the hood, smart routing directs every request to the best available capacity for your model automatically. You send the request, we handle where it runs.