gpt-oss 20B
3.6B active / 20B MoEnative MXFP4 reasoning model that stays practical on 16GB+ cards.
Quant
MXFP4 native
Speed
51.4-98.8
GPU layers
24/24
Score
92
Essential cookies keep the site working. Optional analytics are off unless you allow them.
Review cookie policyLocal LLM Toolkit
Tell it your GPU, workload, and tradeoff. It ranks 40 local models across 6 workloads and 3 priorities, then prefills VRAM checks before you download a 1590GB+ weight file.
5/5 free runs left today
Current pick
native MXFP4 reasoning model that stays practical on 16GB+ cards.
Quant
MXFP4 native
Speed
51.4-98.8 tok/s
Ranked recommendations
native MXFP4 reasoning model that stays practical on 16GB+ cards.
Quant
MXFP4 native
Speed
51.4-98.8
GPU layers
24/24
Score
92
Qwen3.5 MoE quality with a small active path.
Quant
Q8_0
Speed
29.2-56.2
GPU layers
26/40
Score
91
multimodal MoE quality with a compact active path.
Quant
Q8_0
Speed
31.2-60
GPU layers
26/30
Score
91
current Qwen3.5 quality for 24GB-class cards.
Quant
Q8_0
Speed
11.4-22
GPU layers
53/64
Score
90
new Qwen3.8 multimodal model for 24GB-class cards.
Quant
Q8_0
Speed
11.4-22
GPU layers
53/64
Score
90
largest current dense Gemma 4 option.
Quant
Q8_0
Speed
9.9-19
GPU layers
44/60
Score
90
The full report for the top pick: recommended quant, the complete VRAM tradeoff table, and the llama.cpp launch command with the right --n-gpu-layers. Sent once, immediately.
Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.
Default context
4K tokens
Scope
40 models / 6 workloads / 3 priorities
Recent usage
0 tracked runs / 30d