§ 004 / MEASURED FIT RECEIPT
RTX 3070 8GB runs Qwen3.5 9B.
One reproducible answer from owned hardware. The number below is tied to the exact quant, context, workload, and runtime setup shown.
GPU
RTX 3070 8GB
Model tag
qwen3.5:9b
Generation
59.39 tok/s
Measured
2026-07-26
measured on real hardware
Headline run
Workload
Agent code task · 512 tokens
Prompt eval
18.70 tok/s
Peak VRAM
6.8 GB
Source date
2026-07-26
Run command
ollama run qwen3.5:9bAll recorded runs
Context changes the memory budget and speed. A result measured at 4K is not a measured claim at 8K.
| Context | Workload | Quant | Generation | Prompt | Peak VRAM |
|---|---|---|---|---|---|
| 4K context | Agent code task · 512 tokens | Q4_K_M | 59.39 tok/s | 18.70 | 6.8 GB |
Evidence and next step
This receipt is a measured data point, not a promise that every runtime or driver behaves identically. The source record is kept below so the claim can be audited and extended with the next run.
- Reports/Sizing-Desk/fluarmn-tier-8gb-2026-07-25.jsonl