§ 004 / MEASURED FIT RECEIPT
RTX 3070 8GB runs Llama 3.1 8B.
One reproducible answer from owned hardware. The number below is tied to the exact quant, context, workload, and runtime setup shown.
GPU
RTX 3070 8GB
Model tag
llama3.1:8b
Generation
73.74 tok/s
Measured
2026-07-26
measured on real hardware
Headline run
Workload
Agent code task · 512 tokens
Prompt eval
625.23 tok/s
Peak VRAM
5.4 GB
Source date
2026-07-26
Run command
ollama run llama3.1:8bAll recorded runs
Context changes the memory budget and speed. A result measured at 4K is not a measured claim at 8K.
| Context | Workload | Quant | Generation | Prompt | Peak VRAM |
|---|---|---|---|---|---|
| 4K context | Agent code task · 512 tokens | Q4_K_M | 73.74 tok/s | 625.23 | 5.4 GB |
Evidence and next step
This receipt is a measured data point, not a promise that every runtime or driver behaves identically. The source record is kept below so the claim can be audited and extended with the next run.
- Reports/Sizing-Desk/fluarmn-tier-8gb-2026-07-25.jsonl