Topic
Quantization
6 posts on quantization, guides and lab notes from real runs on hardware we own. New posts land here automatically. Start anywhere, or grab the copy-paste prompts that ship with them.
- 5 min read
3 Tests Before a GGUF Quant Runs Your Coding Agent
A GGUF file fitting in VRAM does not prove it can run your coding agent. Use this local acceptance test for tools, patches, and repeat runs.
- 5 min read
49W Average Hid a 338W Burst on Gemma 26B
Gemma 4 26B Q4_K_M averaged 49 W on a long RTX 5090 run and peaked at 338 W. Keep both watt numbers before you compute energy per token.
- 5 min read
A 7 GB 27B Model Lost to My 17 GB Default
A 6.66 GiB ternary 27B model fit my RTX 5090, but it lost the default slot. File size is only the first local model-selection gate.
- 5 min read
Q4_K_M vs Q5_K_M (q4km vs q5km): Which to Download? (2026)
Stuck between q4km and q5km? Verdict: start with Q4_K_M on 8–24GB cards; move to Q5_K_M only if it still fits with KV-cache headroom. The 5-minute VRAM check inside.
- 5 min read
Will That Local Model Fit? Do the VRAM Math First
A local LLM needs about half a gigabyte of VRAM per billion parameters at Q4, then KV cache and context stack on top. Here is how to know a model fits before you download 40 GB.
- 6 min read
Q4_K_M vs Q5_K_M vs Q8_0 vs IQ4_XS: GGUF Quants Explained (2026)
Which GGUF file do you download? Start with Q4_K_M when VRAM is tight, test Q5_K_M or Q8_0 on your real task, keep the smallest that passes. What each quant means, IQ4_XS vs K-quants, a test to copy.
The AI agent build notes
Real costs, real tools, no fluff. One evidence-backed note on Friday when there is something worth sharing.
Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.