[bmdpat]

§ 004 / FIT RECEIPTS

What your hardware actually runs.

These are measured answers, not VRAM guesses: exact GPU, model, quant, context, workload, and observed speed.

Size my GPU4 measured GPU/model receiptsMachine-readable index

measured fit receipt

RTX 3070 8GB runs Llama 3.1 8B

speed

73.74 tok/s

setup

Q4_K_M · 4K context

Agent code task · 512 tokens · 1 recorded run

Open the receipt

measured fit receipt

RTX 3070 8GB runs Qwen3.5 9B

speed

59.39 tok/s

setup

Q4_K_M · 4K context

Agent code task · 512 tokens · 1 recorded run

Open the receipt

measured fit receipt

RTX 5090 32GB runs gemma4:26b

speed

198.8 tok/s

setup

Q4_K_M · 4K context

Short generation · 256 tokens · 3 recorded runs

Open the receipt

measured fit receipt

RTX 5090 32GB runs Llama 3.1 8B

speed

228.9 tok/s

setup

Q4_K_M · 4K context

Short generation · 256 tokens · 3 recorded runs

Open the receipt

Why these exist

A calculator can tell you whether weights fit. A receipt tells you what happened on a real machine. Use one as a starting point, then open the desk with your own GPU and save the rig for future model drops.

Get a fit for my own card →