Pick your GPU and a model, and we'll rank its quants from highest to lowest quality —
with the fit verdict, size, an estimated speed, and a plain-English note on what you give up
at each step down. No account needed.
1 · Your GPU
No matching GPU in the list — you can type the model manually.
Search for your card above (Apple Silicon routes to unified memory automatically).
Without a GPU we'll still show sizes and the quality guidance below.
Used to judge partial CPU offload when a quant overflows VRAM.
Save your rig for one-click answers.
Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.
Pick your GPU above to see which of these fit, plus estimated speeds. Sizes and the quality guidance below don't need a rig.
GGUF · Q4_K_MCommunity default
201.59 GB
The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.
Note: Unsloth Q4_K_M quant of the GLM-4.7 358B-A32B MoE giant — preservation copy. MIT (Z.ai).
GGUF · IQ4_XS
178.41 GB
An i-quant that packs 4-bit weights tighter than Q4_K_S — similar quality in community reports and a smaller file, though slightly heavier to decode on some runtimes.
Note: Unsloth IQ4_XS (imatrix 4-bit) quant of the GLM-4.7 358B-A32B MoE giant — preservation copy. MIT (Z.ai).
Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise).
Quality notes reflect community experience, not measured benchmarks.