---
date: 2026-08-02
generated_at: 2026-08-02 01:54
artifact_type: 5090-report
status: benchmark-data-present
gpu_name: "NVIDIA GeForce RTX 5090"
driver_version: "610.88"
memory_total_mib: "32607"
power_limit_w: "575.00"
benchmark_rows: 1
benchmark_failures: 0
source: Reports/5090/benchmarks/2026-08-02-ollama-v2.csv
---

# The 5090 Reports - 2026-08-02

Production AI agents on hardware I own. This dated artifact records one bounded local agent-workload benchmark with phase timings on an RTX 5090.

## Hardware Snapshot

| GPU | Driver | Memory used / total MiB | Power draw / limit W | Temp C |
|---|---:|---:|---:|---:|
| NVIDIA GeForce RTX 5090 | 610.88 | 17651 / 32607 | 245.96 / 575.00 | 63 |

## Fresh Instrumented Benchmark Row

| Model | Captured at | Workload | Tok/sec | Watts avg | Load ms | Prompt ms | Eval ms | Total ms |
|---|---|---|---:|---:|---:|---:|---:|---:|
| qwen3.5:9b | 2026-08-02 01:53:55 CT | single-shot-local-agent-decision | 84.94 | 300.81 | 6105.25 | 151.18 | 753.47 | 7013.40 |

## Method

One bounded /api/generate workload, num_ctx=4096, num_predict=64. The runner records capture time, load duration, prompt evaluation duration, output evaluation duration, total duration, and GPU power. Residency state is still explicitly unrecorded, so this artifact does not pretend to classify the row as cold or warm. The source CSV and full rollup remain in the private vault.

## Next Measurement

Hold the model and context fixed, then change one workload variable and compare the full row set.
