Qwen

Qwen3-Coder-30B-A3B-Instruct — performance reports

Community-submitted & scraped throughput measurements. Sortable by generation speed.

← Back to model
Ran this model? Log in to submit your own performance result.
Rig / GPU Quant Runtime Context Prompt tok/s Gen tok/s VRAM (GB) Source Reporter
NVIDIA DGX Spark (128 GB) GGUF · Q8_0 llama.cpp 2,048 1,654.25 44.26 scraped anonymous
NVIDIA DGX Spark (128 GB) GGUF · Q8_0 llama.cpp 32,768 686.45 26.92 scraped anonymous

selected to compare · pick at least 2