[bmdpat]

§ 004 / MEASURED FIT RECEIPT

RTX 3070 8GB runs Qwen3.5 9B.

One reproducible answer from owned hardware. The number below is tied to the exact quant, context, workload, and runtime setup shown.

GPU

RTX 3070 8GB

Model tag

qwen3.5:9b

Generation

59.39 tok/s

Measured

2026-07-26

measured on real hardware

Headline run

Q4_K_M · 4K context

Workload

Agent code task · 512 tokens

Prompt eval

18.70 tok/s

Peak VRAM

6.8 GB

Source date

2026-07-26

Run command

ollama run qwen3.5:9b

All recorded runs

Context changes the memory budget and speed. A result measured at 4K is not a measured claim at 8K.

Measured runs for RTX 3070 8GB and Qwen3.5 9B
ContextWorkloadQuantGenerationPromptPeak VRAM
4K contextAgent code task · 512 tokensQ4_K_M59.39 tok/s18.706.8 GB

Evidence and next step

This receipt is a measured data point, not a promise that every runtime or driver behaves identically. The source record is kept below so the claim can be audited and extended with the next run.

  • Reports/Sizing-Desk/fluarmn-tier-8gb-2026-07-25.jsonl
Enter your GPU and save your rig for future model drops →