systemsJul 16, 2026 · 20 min
Fixing ollama's MoE memory estimate — and a 3× kernel speedup that wasn't
Auditing Ollama and llama.cpp GPU memory estimation and quantized kernel execution. Shipped PR #17201 fixes a -20.6% VRAM under-estimation in MoE models. Includes analysis of MMQ kernel tile configurations and DeepSeek MLA execution paths.
mixture of expertscudamemory
−6.0%MoE VRAM error, from −20.6% low · shipped as PR #17201