Describe your hardware, pick a model, then drag the context slider — verdicts update live. No account needed.
No discrete GPU — CPU-only or Apple unified memory below.
Set this for Macs; it replaces the discrete-GPU pool.
Save this rig — get an email when a new model fits it.
Your rig is saved in this browser. Create a free account to keep it and get alerts.
Save my rigLower precision shrinks the KV cache (fits more context).
Fits on GPU. Max context on this rig: .
Partial GPU offload — set n_gpu_layers = .
Runs in Apple unified memory. CPU-only — runs in system RAM (no usable dGPU).
Won't fit — exceeds VRAM + RAM even at minimum context.
Multi-GPU: requires tensor-split across your cards.
Estimates use weights × 1.05 + KV cache + 512 MiB framework overhead. Real usage varies by runtime.
Paste a Hugging Face model URL (or org/repo id) to check whether it fits the hardware above — even models we don't carry.
That doesn't look like a Hugging Face model URL or org/repo id.
Too many lookups from your connection — wait a minute and try again.
Looking up…
Computing a fit profile for from Hugging Face metadata — this page will check again in about 20 seconds.
Still computing — try again shortly.
We can't compute a fit for this repo.
We carry this model — it has been selected above, and the normal per-quant verdicts are shown in the results section.
is license-gated on Hugging Face.
Fit estimate only: to use this model you must accept the license on Hugging Face. We don't host or link its weights.
Fit for
Estimated from Hugging Face metadataQuant sizes from (labelled estimate).
Quant sizes estimated from parameter count — no measured files. Every figure below is an estimate.
Speed estimate: tokens/s (estimate — real speed varies by runtime).
We don't carry this model — figures come from Hugging Face metadata, not measured downloads.