Pick your GPU and a model, and we'll rank its quants from highest to lowest quality —
with the fit verdict, size, an estimated speed, and a plain-English note on what you give up
at each step down. No account needed.
1 · Your GPU
No matching GPU in the list — you can type the model manually.
Search for your card above (Apple Silicon routes to unified memory automatically).
Without a GPU we'll still show sizes and the quality guidance below.
Used to judge partial CPU offload when a quant overflows VRAM.
Save your rig for one-click answers.
Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.
Pick your GPU above to see which of these fit, plus estimated speeds. Sizes and the quality guidance below don't need a rig.
GGUF · Q8_0
7.95 GB
Effectively lossless in community experience — the closest you get to the original model, at the largest practical GGUF size.
Note: Near-lossless; larger footprint.
GGUF · Q4_K_MCommunity default
4.58 GB
The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.
Note: Balanced size/quality — good default.
Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise).
Quality notes reflect community experience, not measured benchmarks.