§ 004 / MEASURED FIT RECEIPT
RTX 5090 32GB runs Llama 3.1 8B.
One reproducible answer from owned hardware. The number below is tied to the exact quant, context, workload, and runtime setup shown.
GPU
RTX 5090 32GB
Model tag
llama3.1:8b
Generation
228.9 tok/s
Measured
2026-07-09
measured on real hardware
Headline run
Workload
Short generation · 256 tokens
Prompt eval
679 tok/s
Peak VRAM
7.2 GB
Source date
2026-07-09
Run command
ollama run llama3.1:8bAll recorded runs
Context changes the memory budget and speed. A result measured at 4K is not a measured claim at 8K.
| Context | Workload | Quant | Generation | Prompt | Peak VRAM |
|---|---|---|---|---|---|
| 4K context | Short generation · 256 tokens | Q4_K_M | 228.9 tok/s | 679 | 7.2 GB |
| 4K context | Agent code task · 512 tokens | Q4_K_M | 227.8 tok/s | 1,028 | 7.8 GB |
| 8K context | Long-context summarize | Q4_K_M | 206.7 tok/s | 12,109 | 7.8 GB |
Evidence and next step
This receipt is a measured data point, not a promise that every runtime or driver behaves identically. The source record is kept below so the claim can be audited and extended with the next run.
- Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md