Pick your GPU and a model, and we'll rank its quants from highest to lowest quality — with the fit verdict, size, an estimated speed, and a plain-English note on what you give up at each step down. No account needed.
Search for your card above (Apple Silicon routes to unified memory automatically). Without a GPU we'll still show sizes and the quality guidance below.
Used to judge partial CPU offload when a quant overflows VRAM.
Save your rig for one-click answers.
Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.
Save my rigLower precision shrinks the KV cache (fits more context).
Full 16-bit (bfloat16) weights: no quantization loss, but about twice an 8-bit quant's size.
Note: Official Moonshot BF16 safetensors — full Kimi-VL-A3B-Instruct package (16.408B-A3B VL MoE + MoonViT vision encoder, tokenizer + modeling code): source-format artifact for fine-tuning/conversion. Full-capability package. MIT.
Effectively lossless in community experience — the closest you get to the original model, at the largest practical GGUF size.
Note: mradermacher Q8_0 quant of Kimi-VL-A3B-Instruct (16.408B-A3B vision-language MoE) — high-quality option; bundled mmproj-f16 required for vision. MIT.
The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.
Note: mradermacher Q4_K_M quant of Kimi-VL-A3B-Instruct (16.408B-A3B vision-language MoE) — recommended balanced pick; bundled mmproj-Q8_0 required for vision. MIT.
An i-quant that packs 4-bit weights tighter than Q4_K_S — similar quality in community reports and a smaller file, though slightly heavier to decode on some runtimes.
Note: mradermacher IQ4_XS quant of Kimi-VL-A3B-Instruct (16.408B-A3B vision-language MoE) — compact low-bit option; bundled mmproj-Q8_0 required for vision. MIT.
Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise). Quality notes reflect community experience, not measured benchmarks.