Skip to content
[bmdpat]

Local LLM Toolkit

Pick the local model that fits your GPU.

Tell it your GPU, workload, and tradeoff. It ranks 40 local models across 6 workloads and 3 priorities, then prefills VRAM checks before you download a 1590GB+ weight file.

Picker

Use case
What matters most?

5/5 free runs left today

Current pick

gpt-oss 20B

native MXFP4 reasoning model that stays practical on 16GB+ cards.

Quant

MXFP4 native

Speed

51.4-98.8 tok/s

Ranked recommendations

6 local models for this GPU

4K context

Your gpt-oss 20B on RTX 4090 24GB config report

The full report for the top pick: recommended quant, the complete VRAM tradeoff table, and the llama.cpp launch command with the right --n-gpu-layers. Sent once, immediately.

Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.

Default context

4K tokens

Scope

40 models / 6 workloads / 3 priorities

Recent usage

0 tracked runs / 30d

FAQ

Which local LLM should I run on a 24GB GPU?
On RTX 4090 24GB, the current Code Generation + Quality profile ranks Qwen3 Coder 30B-A3B (3B active / 30.5B MoE, Q8_0, Partial offload) first. The largest current full-GPU recommendation in that profile is Codestral 22B (22B dense, Q8_0, Full GPU). The nearest 70B-class recommendation is Llama 3.3 70B (70B dense, Q4_K_M, Partial offload), so treat that as a quality/offload pick instead of the default for fast loops.
What is the best quantization for local LLMs?
The picker does not use one universal best quantization. On the same RTX 4090 24GB reference setup, Code Generation + Quality currently chooses Q8_0 for Qwen3 Coder 30B-A3B, Chat/Assistant + Speed chooses IQ3_XXS for Qwen3.5 35B-A3B, and Chat/Assistant + VRAM Efficiency chooses IQ4_XS for Llama 3.2 3B. Those claims come from the ranked catalog plus the VRAM fit estimate at 4,096 tokens.
Should I pick the fastest model or the highest quality model?
Use Speed when tokens per second matters; the current reference winner is Qwen3.5 35B-A3B (3B active / 35B MoE, IQ3_XXS, Full GPU) at 67.6-129.9 tok/s. Use Quality when model strength matters more; the current reference winner is Qwen3 Coder 30B-A3B (3B active / 30.5B MoE, Q8_0, Partial offload) at 32.4-62.3 tok/s. Use VRAM Efficiency when keeping the model on one GPU matters; the current reference winner is Llama 3.2 3B (3B dense, IQ4_XS, Full GPU) at 60.2-115.8 tok/s.