openai

gpt-oss-120b — performance reports

Community-submitted & scraped throughput measurements. Sortable by generation speed.

← Back to model
Ran this model? Log in to submit your own performance result.
Rig / GPU Quant Runtime Context Prompt tok/s Gen tok/s VRAM (GB) Source Reporter
Apple M2 Ultra (192 GB) GGUF · F16 llama.cpp 2,048 1,244.57 79.68 scraped anonymous
AMD Ryzen AI Max+ 395 (128 GB) GGUF · F16 llama.cpp 2,048 55.57 scraped anonymous
NVIDIA DGX Spark (128 GB) GGUF · F16 llama.cpp 2,048 1,737.17 45.87 scraped anonymous
NVIDIA DGX Spark (128 GB) GGUF · F16 llama.cpp 4,096 1,777.81 43.41 scraped anonymous
NVIDIA DGX Spark (128 GB) GGUF · F16 llama.cpp 8,192 1,720.17 41.52 scraped anonymous
NVIDIA DGX Spark (128 GB) GGUF · F16 llama.cpp 16,384 1,512.23 38.39 scraped anonymous
NVIDIA DGX Spark (128 GB) GGUF · F16 llama.cpp 32,768 1,231.86 34.29 scraped anonymous

selected to compare · pick at least 2