Topic
Llama CPP
11 posts on llama cpp, guides and lab notes from real runs on hardware we own. New posts land here automatically. Start anywhere, or grab the copy-paste prompts that ship with them.
- 5 min read
Why Your Local Model Fits and Still Fails at Long Context
A local model can load and still run out of memory at longer context. Compare two controlled loads, inspect cache logs, and test the real workload.
- 5 min read
Prove llama.cpp Tensor Split Used Every GPU
A tensor-split flag is only a request. Pin the split, watch every GPU, and save one repeatable llama.cpp receipt before trusting the result.
- 5 min read
Does Ollama Include That New llama.cpp Feature?
A llama.cpp release note does not prove that Ollama can use the feature. I trace the active runtime path, pin versions, and test the same workload.
- 5 min read
Q4_K_M vs Q5_K_M (q4km vs q5km): Which to Download? (2026)
Stuck between q4km and q5km? Verdict: start with Q4_K_M on 8–24GB cards; move to Q5_K_M only if it still fits with KV-cache headroom. The 5-minute VRAM check inside.
- 5 min read
Will That Local Model Fit? Do the VRAM Math First
A local LLM needs about half a gigabyte of VRAM per billion parameters at Q4, then KV cache and context stack on top. Here is how to know a model fits before you download 40 GB.
- 5 min read
How to Run Local LLM Verifier Loops on Owned Hardware
A local LLM workflow needs more than a model prompt. It needs a verifier loop that proves the file, command, URL, or report changed before the agent claims done.
- 7 min read
llama.cpp Multi-GPU Guide: --tensor-split & --split-mode (2026)
Running a 70B across 2 GPUs and hitting OOM? The llama.cpp multi-GPU reference: --tensor-split ratios, --split-mode layer vs row, --main-gpu, and the VRAM math from a real two-card rig.
- 6 min read
llama.cpp -ngl Flag Explained: 5 Fixes When It Stays on CPU
Set -ngl 99 in llama-cli and the GPU still sits idle? The flag isn't the bug. What -ngl (--n-gpu-layers) does, the 30-second load-log check, and the 5 real causes ranked by how often they bite.
- 6 min read
Q4_K_M vs Q5_K_M vs Q8_0 vs IQ4_XS: GGUF Quants Explained (2026)
Which GGUF file do you download? Start with Q4_K_M when VRAM is tight, test Q5_K_M or Q8_0 on your real task, keep the smallest that passes. What each quant means, IQ4_XS vs K-quants, a test to copy.
- 7 min read
llama.cpp --n-gpu-layers: -1, 0, Partial GPU Offload (2026)
Not sure what to set --n-gpu-layers to? -1 offloads all layers, 0 keeps it on CPU, a number splits the model. VRAM headroom rules, examples, and the CPU-fallback fix. (2026)
- 7 min read
Local LLM on Consumer GPUs: 50 req/s, $0/Call [Benchmarks 2026]
Cloud LLM bills hit $2K/month fast. An RTX 5070 Ti serves Llama 3.1 at 50 req/s for $0 per call, we benchmarked 4 consumer GPUs and built the exact production setup.
The AI agent build notes
Real costs, real tools, no fluff. One evidence-backed note on Friday when there is something worth sharing.
Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.