Community-submitted & scraped throughput measurements. Sortable by generation speed.
← Back to model| Rig / GPU | Quant | Runtime | Context | Prompt tok/s | Gen tok/s | VRAM (GB) | Source | Reporter |
|---|---|---|---|---|---|---|---|---|
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 4,096 | 3,625.60 | 119.40 | — | scraped | anonymous |
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 16,384 | 3,068.90 | 115.00 | — | scraped | anonymous |
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 32,768 | 2,453.40 | 107.50 | — | scraped | anonymous |
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 65,536 | 1,765.10 | 98.90 | — | scraped | anonymous |
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 131,072 | 1,147.10 | 83.00 | — | scraped | anonymous |
| RTX 3090 | GGUF · Q4_K_M | llama.cpp | 262,144 | 671.40 | 64.40 | — | scraped | anonymous |
selected to compare · pick at least 2