Pick a GPU to see exactly which models run on it, the sweet-spot quant for each, and estimated speeds. Cards are grouped by generation; memory bandwidth (which drives tokens/sec) is shown on every one. Counts below are models that run fully on GPU at a 8,192-token context.
selected to compare · pick at least 2