§ 001 / LOCAL AI MAP
What local AI hardware and models can you actually run?
A single map of the sizing desk, VRAM math, quant choice, and measured RTX 5090 runs. 40 models and 58 GPUs. No sales call.
§ 002 / ANSWER MACHINES
Start at the desk. Use a calculator when you need the math.
§ 003 / GUIDES
Deep dives that feed the same question
- llama.cpp n-gpu-layers explainedHow layer offload decides what actually fits.
- GGUF Q4 vs Q5 vs Q8Pick a quant without guessing from Reddit threads.
- Local LLM inference on consumer GPUsWhat a single gaming card can host in practice.
- Pick a GGUF quant for a VRAM budgetA budget-first path into the same math.
- GPU scarcity and consumer hardwareWhy local fit still matters when cloud GPUs are tight.
§ 004 / FAQ
Short answers
- Where do I check if a local model fits my GPU?
- Open the sizing desk at /desk. Paste the GPU you own. It returns ranked models, quants, expected speed, and a run command.
- What if I know the model I want, not the GPU?
- Use "I want this model" on the sizing desk. It ranks catalog GPUs that can run that model, and marks the smallest card that still fits fully.
- Where are the measured numbers from?
- The 5090 Reports publish owned-hardware runs. The desk shows a measured badge only when a matching GPU and model row exists.
- Do the old calculators still work?
- Yes. The VRAM calculator, model picker, and quant compare stay live for deep math. The desk is the one-question front door.
Get the local AI lab notes
New benchmark rows, VRAM fit checks, and model-fit notes from measured runs on owned hardware. One evidence-backed note on Friday when there is something worth sharing.
Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.