Community-submitted & scraped throughput measurements. Sortable by generation speed.
← Back to model| Rig / GPU | Quant | Runtime | Context | Prompt tok/s | Gen tok/s | VRAM (GB) | Source | Reporter |
|---|---|---|---|---|---|---|---|---|
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 4,096 | 1,155.80 | 34.70 | — | scraped | anonymous |
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 16,384 | 913.20 | 33.50 | — | scraped | anonymous |
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 32,768 | 723.70 | 31.40 | — | scraped | anonymous |
selected to compare · pick at least 2