# RTX 3070 8GB runs Llama 3.1 8B

- canonical_url: https://bmdpat.com/receipts/rtx-3070/llama-3-1-8b
- markdown_url: https://bmdpat.com/receipts/rtx-3070/llama-3-1-8b.md
- last_updated: 2026-07-26
- answer_kind: measured
- confidence: measured
- measured_on: 2026-07-26
- source: Reports/Sizing-Desk/fluarmn-tier-8gb-2026-07-25.jsonl
- gpu: RTX 3070 8GB (rtx-3070; 8 GB VRAM)
- model: Llama 3.1 8B (llama-3-1-8b)

This is one reproducible result from owned hardware. It does not prove identical performance on every driver, runtime, thermal condition, prompt, or context length.

## Headline run
- quantization: Q4_K_M
- context: 4K context
- workload: Agent code task · 512 tokens
- generation: 73.74 tok/s
- prompt_eval: 625.23 tok/s
- peak_vram: 5.4 GB
- model_tag: llama3.1:8b
- ollama_command: ollama run llama3.1:8b

## Recorded runs
- measured_on: 2026-07-26; context: 4K context; workload: Agent code task · 512 tokens; quantization: Q4_K_M; generation: 73.74 tok/s; prompt_eval: 625.23 tok/s; peak_vram: 5.4 GB; source: Reports/Sizing-Desk/fluarmn-tier-8gb-2026-07-25.jsonl

- personalized_path: https://bmdpat.com/desk?gpu=rtx-3070&vram=8&ctx=4096&use=code&priority=speed&utm_source=fit-receipt&utm_medium=artifact&utm_campaign=saved-rig
- correction_path: https://bmdpat.com/desk