Which quant should I pick?

Pick your GPU and a model, and we'll rank its quants from highest to lowest quality — with the fit verdict, size, an estimated speed, and a plain-English note on what you give up at each step down. No account needed.

1 · Your GPU

  • No matching GPU in the list — you can type the model manually.

Search for your card above (Apple Silicon routes to unified memory automatically). Without a GPU we'll still show sizes and the quality guidance below.

Used to judge partial CPU offload when a quant overflows VRAM.

Save your rig for one-click answers.

Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.

Save my rig

2 · Model & context

Lower precision shrinks the KV cache (fits more context).

Pick your GPU above to see which of these fit, plus estimated speeds. Sizes and the quality guidance below don't need a rig.
  1. GGUF · Q8_0
    14.51 GB

    Effectively lossless in community experience — the closest you get to the original model, at the largest practical GGUF size.

    Note: Bartowski Q8_0 quant of Phi-4 (14.66B dense) — high-quality option. MIT.

  2. GGUF · Q4_K_M Community default
    8.28 GB

    The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.

    Note: Balanced size/quality — good default.

  3. GGUF · IQ4_XS
    7.40 GB

    An i-quant that packs 4-bit weights tighter than Q4_K_S — similar quality in community reports and a smaller file, though slightly heavier to decode on some runtimes.

    Note: Bartowski IQ4_XS quant of Phi-4 (14.66B dense) — compact low-bit option. MIT.

Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise). Quality notes reflect community experience, not measured benchmarks.