Llama 3.1 8B
Measured at 73.74 tok/s with 5.4 GB peak VRAM.
ollama run llama3.1:8bMeasured GPU guide / RTX 3070
An 8GB RTX 3070 can run useful local models without CPU offload. These are owned-hardware measurements at Q4_K_M and 4K context, not VRAM estimates.
Measured results
Both runs used Ollama and a code-generation workload. Generation speed, prompt speed, and peak VRAM come directly from the recorded run.
| Model | Generation | Prompt | Peak VRAM | Fit |
|---|---|---|---|---|
| Llama 3.1 8BQ4_K_M / 4K | 73.74 tok/s | 625.23 tok/s | 5.4 GB | Full GPU |
| Qwen3.5 9BQ4_K_M / 4K | 59.39 tok/s | 18.70 tok/s | 6.8 GB | Full GPU |
Measured at 73.74 tok/s with 5.4 GB peak VRAM.
ollama run llama3.1:8bMeasured at 59.39 tok/s with 6.8 GB peak VRAM.
ollama run qwen3.5:9bContext matters
The measured rows above are exact only at 4K context. A larger context cache consumes more VRAM and can move model layers to the CPU. That reduces speed. The sizing desk recomputes the answer for the context and workload you choose.
Size my exact RTX 3070 setup