§ 004 / MEASURED FIT RECEIPT
RTX 5090 32GB runs gemma4:26b.
One reproducible answer from owned hardware. The number below is tied to the exact quant, context, workload, and runtime setup shown.
GPU
RTX 5090 32GB
Model tag
gemma4:26b
Generation
198.8 tok/s
Measured
2026-07-09
measured on real hardware
Headline run
Workload
Short generation · 256 tokens
Prompt eval
15 (cold) tok/s
Peak VRAM
19.9 GB
Source date
2026-07-09
Run command
This measured model does not have a verified Ollama tag in the catalog yet. Use the model label and source receipt to reproduce the run; no command is invented here.
All recorded runs
Context changes the memory budget and speed. A result measured at 4K is not a measured claim at 8K.
| Context | Workload | Quant | Generation | Prompt | Peak VRAM |
|---|---|---|---|---|---|
| 4K context | Short generation · 256 tokens | Q4_K_M | 198.8 tok/s | 15 (cold) | 19.9 GB |
| 4K context | Agent code task · 512 tokens | Q4_K_M | 207.4 tok/s | 197 | 20.2 GB |
| 8K context | Long-context summarize | Q4_K_M | 180.2 tok/s | 6,179 | 20.2 GB |
Evidence and next step
This receipt is a measured data point, not a promise that every runtime or driver behaves identically. The source record is kept below so the claim can be audited and extended with the next run.
- Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md