RTX Pro 6000 Blackwell 96GB

NVIDIA RTX Pro (Blackwell)

VRAM
96 GB
Memory bandwidth
1,792 GB/s
fp16 compute
503.8 TFLOPS
Class
Workstation

Figures assume this GPU plus 32 GB of system RAM (a typical desktop pairing). "What runs on it" is judged at a 8,192-token context. Speeds are estimates, not measurements.

What runs on it

Model Sweet-spot quant Est. speed Community
DeepSeek-R1-Distill-Llama-8B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 112–149 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-1.5B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 505–673 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-14B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 62–83 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-32B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 29–39 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-7B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 125–167 tok/s (estimate) no community data
Devstral-Small-2-24B-Instruct-2512 mistralai Q8_0 Runs fully on GPU @ 8K ctx 41–54 tok/s (estimate) no community data
GLM-4.7-Flash zai-org Q8_0 Runs fully on GPU @ 8K ctx no community data
Hunyuan-A13B-Instruct tencent Q8_0 Runs fully on GPU @ 8K ctx no community data
Kimi-Dev-72B moonshotai UD-Q5_K_XL Runs fully on GPU @ 8K ctx 19–25 tok/s (estimate) no community data
Kimi-VL-A3B-Instruct moonshotai BF16 Runs fully on GPU @ 8K ctx no community data
Llama-3.1-8B-Instruct meta-llama Q8_0 Runs fully on GPU @ 8K ctx 112–149 tok/s (estimate) no community data
Llama-3.3-70B-Instruct meta-llama Q8_0 Runs fully on GPU @ 8K ctx 14–18 tok/s (estimate) no community data
Mistral-Small-3.2-24B-Instruct-2506 mistralai Q8_0 Runs fully on GPU @ 8K ctx 41–54 tok/s (estimate) no community data
Phi-4-mini-instruct microsoft Q8_0 Runs fully on GPU @ 8K ctx 208–278 tok/s (estimate) no community data
Phi-4-reasoning microsoft Q8_0 Runs fully on GPU @ 8K ctx 62–83 tok/s (estimate) no community data
Qwen2.5-7B-Instruct Qwen Q8_0 Runs fully on GPU @ 8K ctx 125–167 tok/s (estimate) no community data
Qwen2.5-Omni-7B Qwen BF16 Runs fully on GPU @ 8K ctx 47–63 tok/s (estimate) no community data
Qwen2.5-VL-32B-Instruct Qwen Q8_0 Runs fully on GPU @ 8K ctx 29–39 tok/s (estimate) no community data
Qwen3-14B Qwen Q8_0 Runs fully on GPU @ 8K ctx 63–84 tok/s (estimate) no community data
Qwen3-235B-A22B-Instruct-2507 Qwen UD-Q2_K_XL Runs fully on GPU @ 8K ctx no community data
Qwen3-30B-A3B-Instruct-2507 Qwen Q8_0 Runs fully on GPU @ 8K ctx no community data
Qwen3-32B Qwen Q8_0 Runs fully on GPU @ 8K ctx 29–39 tok/s (estimate) no community data
Qwen3-8B Qwen Q8_0 Runs fully on GPU @ 8K ctx 108–145 tok/s (estimate) no community data
Qwen3-Coder-30B-A3B-Instruct Qwen Q8_0 Runs fully on GPU @ 8K ctx no community data
Qwen3-Embedding-0.6B Qwen Q8_0 Runs fully on GPU @ 8K ctx 681–908 tok/s (estimate) no community data
Qwen3-Embedding-4B Qwen Q4_K_M Runs fully on GPU @ 8K ctx 290–387 tok/s (estimate) no community data
Qwen3-Embedding-8B Qwen Q4_K_M Runs fully on GPU @ 8K ctx 183–244 tok/s (estimate) no community data
Qwen3-Omni-30B-A3B-Instruct Qwen Q4_K_M Runs fully on GPU @ 8K ctx no community data
Qwen3-Reranker-0.6B Qwen BF16 Runs fully on GPU @ 8K ctx 505–673 tok/s (estimate) no community data
Qwen3-Reranker-4B Qwen BF16 Runs fully on GPU @ 8K ctx 116–155 tok/s (estimate) no community data
Qwen3-Reranker-8B Qwen BF16 Runs fully on GPU @ 8K ctx 61–82 tok/s (estimate) no community data
Qwen3.6-27B Qwen Q8_0 Runs fully on GPU @ 8K ctx 35–47 tok/s (estimate) no community data
Qwen3.6-35B-A3B Qwen Q8_0 Runs fully on GPU @ 8K ctx no community data
SmolLM3-3B HuggingFaceTB BF16 Runs fully on GPU @ 8K ctx 159–212 tok/s (estimate) no community data
dots.ocr rednote-hilab BF16 Runs fully on GPU @ 8K ctx 170–227 tok/s (estimate) no community data
gemma-4-12B-it google Q8_0 Runs fully on GPU @ 8K ctx 79–106 tok/s (estimate) no community data
gemma-4-26B-A4B-it google Q8_0 Runs fully on GPU @ 8K ctx no community data
gemma-4-31B-it google Q8_0 Runs fully on GPU @ 8K ctx 31–41 tok/s (estimate) no community data
gemma-4-E2B-it google Q8_0 Runs fully on GPU @ 8K ctx 210–280 tok/s (estimate) no community data
gemma-4-E4B-it google Q8_0 Runs fully on GPU @ 8K ctx 129–172 tok/s (estimate) no community data
gpt-oss-120b openai F16 Runs fully on GPU @ 8K ctx no community data
gpt-oss-20b openai F16 Runs fully on GPU @ 8K ctx no community data
phi-4 microsoft Q8_0 Runs fully on GPU @ 8K ctx 62–83 tok/s (estimate) no community data

"Est. speed" is a modelled range labelled estimate (D8) for generation (decode) throughput. "Community" shows the median of approved user-submitted reports on this GPU class only where enough exist — never an estimate. "pp" is measured prompt-processing (ingestion) throughput from approved community reports; rows without a measurement show none.

What can I run on a RTX Pro 6000 Blackwell 96GB?

On this GPU, 12 catalog models run fully on the GPU at an 8,192-token context. The most capable is Qwen3-235B-A22B-Instruct-2507 at UD-Q2_K_XL (needs ~93.1 GiB). Pick a smaller model or a lower quant for more headroom.

Biggest model: Qwen3-235B-A22B-Instruct-2507 at UD-Q2_K_XL

llama-server -m Qwen3-235B-A22B-Instruct-2507-UD-Q2_K_XL-00001-of-00002.gguf -c 8192 -ngl 999

Derived from the fit engine at an 8,192-token context. See more answer packs.

Add a second RTX Pro 6000 Blackwell 96GB?

A second RTX Pro 6000 Blackwell 96GB pools VRAM: 2 × 96 GB = 192 GB combined. Bigger models can then load because their weights split across both cards — but a second card does not make generation proportionally faster (see the reality check below).

A second RTX Pro 6000 Blackwell 96GB adds 96 GB of VRAM for about $11,830 — ≈$123.23/GB of added VRAM (new price as of 2026-07-18 — Newegg).

3 more catalog models could newly fit fully in the combined 192 GB at a 8,192-token context — for example:

Estimate — assumes the model's weights split across both cards (a layer split, as llama.cpp does by default). This is a combined-VRAM projection, not a measured or verdict-chipped result: the fit engine treats two cards as one summed memory pool and does not model the link between them. Check your exact model and context in the calculator.

The honest reality of a second card

  • Bandwidth doesn't add. Two cards give more VRAM, not more memory bandwidth per token. Token generation is bandwidth-bound, so a layer-split model decodes at roughly one card's speed — not double.
  • Layer-split vs tensor-parallel. The common desktop setup (llama.cpp) splits layers across cards and runs them in sequence, so one GPU works at a time. True tensor-parallel serving (e.g. vLLM) can use both at once, but wants matched cards and a fast interconnect.
  • pp vs tg. Prompt processing (pp) can gain more from a second card than token generation (tg); raw decode throughput barely moves. Don't expect a 2× tok/s jump.
  • PCIe / NUMA. Cards talk over PCIe (or NVLink where supported), far slower than on-card VRAM. A layer split crosses it about once per token so the hit is small; tensor-parallel crosses it constantly, and cards on different CPU sockets (NUMA) add latency.
  • Power & PSU. A second card roughly doubles GPU power draw — check PSU headroom, connectors, slot spacing and airflow before buying.

Read: Multi-GPU for local LLMs — when a second card is worth it →

Derived from the fit engine at a 8,192-token context, comparing a single 96 GB card against a summed 192 GB two-card pool.

Price history

Prices are point-in-time observations, not live quotes.

new

  • Jul 2026 · $11,830 new — Newegg (medium confidence)

Not enough history yet — trends appear once at least three dated observations are recorded.

selected to compare · pick at least 2