meta-llama

Llama-3.3-70B-Instruct — performance reports

Community-submitted & scraped throughput measurements. Sortable by generation speed.

← Back to model
Ran this model? Log in to submit your own performance result.
Rig / GPU Quant Runtime Context Prompt tok/s Gen tok/s VRAM (GB) Source Reporter
Apple M3 Ultra (512 GB) GGUF · Q4_K_M llama.cpp 2,048 162.30 14.41 scraped anonymous

selected to compare · pick at least 2