DeepSeek’s latest mixture-of-experts model with vision capabilities is available now. It uses a new encoder-decoder design with 8B active parameters for input and 16B for output. DeepSeek reports roughly 1/4th of V4 Flash’s KV cache memory per token. DeepSeek’s recommendation, and ours: If you’re running workloads on DeepSeek V4 Pro, consider switching to V4.1 Flash. DeepSeek reports better performance on most benchmarks at a lower price. |