# RTX 5090 32GB runs Llama 3.1 8B

- canonical_url: https://bmdpat.com/receipts/rtx-5090/llama-3-1-8b
- markdown_url: https://bmdpat.com/receipts/rtx-5090/llama-3-1-8b.md
- last_updated: 2026-07-09
- answer_kind: measured
- confidence: measured
- measured_on: 2026-07-09
- source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md
- gpu: RTX 5090 32GB (rtx-5090; 32 GB VRAM)
- model: Llama 3.1 8B (llama-3-1-8b)

This is one reproducible result from owned hardware. It does not prove identical performance on every driver, runtime, thermal condition, prompt, or context length.

## Headline run
- quantization: Q4_K_M
- context: 4K context
- workload: Short generation · 256 tokens
- generation: 228.9 tok/s
- prompt_eval: 679 tok/s
- peak_vram: 7.2 GB
- model_tag: llama3.1:8b
- ollama_command: ollama run llama3.1:8b

## Recorded runs
- measured_on: 2026-07-09; context: 4K context; workload: Short generation · 256 tokens; quantization: Q4_K_M; generation: 228.9 tok/s; prompt_eval: 679 tok/s; peak_vram: 7.2 GB; source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md
- measured_on: 2026-07-09; context: 4K context; workload: Agent code task · 512 tokens; quantization: Q4_K_M; generation: 227.8 tok/s; prompt_eval: 1,028 tok/s; peak_vram: 7.8 GB; source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md
- measured_on: 2026-07-09; context: 8K context; workload: Long-context summarize; quantization: Q4_K_M; generation: 206.7 tok/s; prompt_eval: 12,109 tok/s; peak_vram: 7.8 GB; source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md

- personalized_path: https://bmdpat.com/desk?gpu=rtx-5090&vram=32&ctx=4096&use=code&priority=speed&utm_source=fit-receipt&utm_medium=artifact&utm_campaign=saved-rig
- correction_path: https://bmdpat.com/desk