Pick your GPU and a model, and we'll rank its quants from highest to lowest quality —
with the fit verdict, size, an estimated speed, and a plain-English note on what you give up
at each step down. No account needed.
1 · Your GPU
No matching GPU in the list — you can type the model manually.
Search for your card above (Apple Silicon routes to unified memory automatically).
Without a GPU we'll still show sizes and the quality guidance below.
Used to judge partial CPU offload when a quant overflows VRAM.
Save your rig for one-click answers.
Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.
Note: Google official QAT Q4_0 quant of Gemma-4-E4B-it (7.996B) — compact low-bit option. Apache-2.0.
GGUF · Q4_K_MCommunity default
4.64 GB
The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.
Note: Balanced size/quality — good default.
Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise).
Quality notes reflect community experience, not measured benchmarks.