Lyceum5 October 2026
Max Niroomand

Max Niroomand

CTO and co-founder, Lyceum

New inference engine for GLM-5.3

We cut GLM-5.3 costs by 20%

The new version of our inference engine is so much faster that we can cut the cost of GLM-5.3 by 20%.

GLM-5.3 is one of our most used models, thanks to its incredible price-performance. We built custom modifications on top of NVIDIA Dynamo, NVIDIA's inference framework.

The result: serverless inference with the KV-cache hit rate you would normally only see on dedicated machines.

Figure 1: three turns of one agent session go through a cache-aware router to GPU node 2, which already holds the session's KV cache.

Fig. 1 Every turn of a session goes back to the GPU that already holds its KV cache, so the context isn't computed twice.

What changes for GLM-5.3USD per million tokens, excluding VAT
 BeforeNowChange
Input$1.75$1.40−20%
Cached input$0.44$0.26−41%
Output$4.50$4.40−2%
Try GLM-5.3See all models and prices →

Shoutout

A huge congratulations to our rockstar team around Mina Tawfik, Aurelien Bloch and Timo Nicolai, who put in one of the most insane shifts I've ever seen to get this live.

Mina Tawfik, inference engineer at Lyceum

Mina Tawfik

Inference engineer

Aurelien Bloch, inference engineer at Lyceum

Aurelien Bloch

Inference engineer

Timo Nicolai, member of technical staff at Lyceum

Timo Nicolai

Member of technical staff

One more thing

Go check it out on our brand-new website (best website I've ever seen, insane job Mahir). We have a new logo too, though I'm still bitter that my ChatGPT-generated logo didn't make the cut.

The old Lyceum logo, made in ChatGPT

2025: one hour in ChatGPT

→The new Lyceum wordmark

2026: the new Lyceum

     lyceum.technology
The new Lyceum homepage, Run open models in Europe, with its pixel ground moving

Have a look at the new website →

Cheers,

Max

CTO and co-founder, Lyceum