# RTX 5090 32GB runs gemma4:26b

- canonical_url: https://bmdpat.com/receipts/rtx-5090/gemma4-26b
- markdown_url: https://bmdpat.com/receipts/rtx-5090/gemma4-26b.md
- last_updated: 2026-07-09
- answer_kind: measured
- confidence: measured
- measured_on: 2026-07-09
- source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md
- gpu: RTX 5090 32GB (rtx-5090; 32 GB VRAM)
- model: gemma4:26b (gemma4-26b)

This is one reproducible result from owned hardware. It does not prove identical performance on every driver, runtime, thermal condition, prompt, or context length.

## Headline run
- quantization: Q4_K_M
- context: 4K context
- workload: Short generation · 256 tokens
- generation: 198.8 tok/s
- prompt_eval: 15 (cold) tok/s
- peak_vram: 19.9 GB
- model_tag: gemma4:26b
- ollama_command: not verified

## Recorded runs
- measured_on: 2026-07-09; context: 4K context; workload: Short generation · 256 tokens; quantization: Q4_K_M; generation: 198.8 tok/s; prompt_eval: 15 (cold) tok/s; peak_vram: 19.9 GB; source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md
- measured_on: 2026-07-09; context: 4K context; workload: Agent code task · 512 tokens; quantization: Q4_K_M; generation: 207.4 tok/s; prompt_eval: 197 tok/s; peak_vram: 20.2 GB; source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md
- measured_on: 2026-07-09; context: 8K context; workload: Long-context summarize; quantization: Q4_K_M; generation: 180.2 tok/s; prompt_eval: 6,179 tok/s; peak_vram: 20.2 GB; source: Reports/5090/benchmarks/2026-07-09-5090-sweep-llama8b-gemma26b.csv via /5090-reports/latest.md

- personalized_path: https://bmdpat.com/desk?gpu=rtx-5090&vram=32&ctx=4096&use=code&priority=speed&utm_source=fit-receipt&utm_medium=artifact&utm_campaign=saved-rig
- correction_path: https://bmdpat.com/desk