[bmdpat]

§ 004 / MEASURED FIT RECEIPT

RTX 5090 32GB runs gemma4:26b.

One reproducible answer from owned hardware. The number below is tied to the exact quant, context, workload, and runtime setup shown.

GPU

RTX 5090 32GB

Model tag

gemma4:26b

Generation

198.8 tok/s

Measured

2026-07-09

measured on real hardware

Headline run

Q4_K_M · 4K context

Workload

Short generation · 256 tokens

Prompt eval

15 (cold) tok/s

Peak VRAM

19.9 GB

Source date

2026-07-09

Run command

This measured model does not have a verified Ollama tag in the catalog yet. Use the model label and source receipt to reproduce the run; no command is invented here.

All recorded runs

Context changes the memory budget and speed. A result measured at 4K is not a measured claim at 8K.

Measured runs for RTX 5090 32GB and gemma4:26b
ContextWorkloadQuantGenerationPromptPeak VRAM
4K contextShort generation · 256 tokensQ4_K_M198.8 tok/s15 (cold)19.9 GB
4K contextAgent code task · 512 tokensQ4_K_M207.4 tok/s19720.2 GB
8K contextLong-context summarizeQ4_K_M180.2 tok/s6,17920.2 GB

Evidence and next step

This receipt is a measured data point, not a promise that every runtime or driver behaves identically. The source record is kept below so the claim can be audited and extended with the next run.

  • Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md
Enter your GPU and save your rig for future model drops →