Writing · Tag
#rtx-3070
2 posts tagged #rtx-3070.
- 5 min read
A 27B Model Fit on an 8 GB GPU. It Was Slow.
Qwen3.8-27B Q3_K_S loaded on an RTX 3070. VRAM used 7,435 of 8,192 MiB. Decode ran 2.07 tok/s. Fit on 8 GB is not a usable rate.
#local-llm#ollama#qwen#vram - 5 min read
Unload Local LLMs After Every Test
A local model test is not over when text appears. I unload the model, read idle VRAM, and record the result before I start another run.
#local-llm#ollama#rtx-3070#5090-reports
The AI agent build notes
Real costs, real tools, no fluff. One evidence-backed note on Friday when there is something worth sharing.
Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.