openai
Community-submitted & scraped throughput measurements. Sortable by generation speed.
← Back to model| Rig / GPU | Quant | Runtime | Context | Prompt tok/s | Gen tok/s | VRAM (GB) | Source | Reporter |
|---|---|---|---|---|---|---|---|---|
| Apple M2 Ultra (192 GB) | GGUF · F16 | llama.cpp | 2,048 | 1,244.57 | 79.68 | — | scraped | anonymous |
| AMD Ryzen AI Max+ 395 (128 GB) | GGUF · F16 | llama.cpp | 2,048 | — | 55.57 | — | scraped | anonymous |
| NVIDIA DGX Spark (128 GB) | GGUF · F16 | llama.cpp | 2,048 | 1,737.17 | 45.87 | — | scraped | anonymous |
| NVIDIA DGX Spark (128 GB) | GGUF · F16 | llama.cpp | 4,096 | 1,777.81 | 43.41 | — | scraped | anonymous |
| NVIDIA DGX Spark (128 GB) | GGUF · F16 | llama.cpp | 8,192 | 1,720.17 | 41.52 | — | scraped | anonymous |
| NVIDIA DGX Spark (128 GB) | GGUF · F16 | llama.cpp | 16,384 | 1,512.23 | 38.39 | — | scraped | anonymous |
| NVIDIA DGX Spark (128 GB) | GGUF · F16 | llama.cpp | 32,768 | 1,231.86 | 34.29 | — | scraped | anonymous |
selected to compare · pick at least 2