§ 001 / FIT INDEX
Which local models fit on which GPU?
One table. 5 cards from 8 GB to 32 GB. 17 models. Each cell shows the best quant that fits fully in VRAM at 8K context.
Catalog verified
The fit comes from the same calculator as the sizing desk. Measured values come from runs on the RTX 5090 32GB and the RTX 3070 8GB.
§ 002 / THE TABLE
Best full fit at 8K context
Smallest model first. Pick a card to open the desk with that card filled in.
RTX 3070 8GBOpen in the desk
- Llama 3.2 3B3B denseQ8_0near-lossless · 4.6 GB est.
- Qwen3.5 4B4B denseQ8_0near-lossless · 5.5 GB est.
- Gemma 4 E4B4B dense / 8B totalQ6_Khigh · 6.8 GB est.
- Llama 3.1 8B8B denseQ6_Khigh · 7.7 GB est.Measured 73.74 tok/sQ4_K_M · 4K context
- Ministral 3 8B8B denseQ5_K_Mbalanced+ · 7.5 GB est.
- Qwen3.5 9B9.7B denseQ5_K_Mbalanced+ · 7.3 GB est.Measured 59.39 tok/sQ4_K_M · 4K context
- Gemma 4 12B12B denseQ4_K_Mbalanced · 7.9 GB est.
- DeepSeek-R1 Distill Qwen 14B14.8B denseIQ3_XXScompressed · 7.9 GB est.
- gpt-oss 20B3.6B active / 20B MoEdoes not fit
- Gemma 4 26B-A4B4B active / 26B MoEdoes not fit
- Qwen3.5 27B27B densedoes not fit
- Qwen3.8 27B27.8B densedoes not fit
- Gemma 4 31B31B densedoes not fit
- Qwen3 32B32.8B densedoes not fit
- Qwen3.5 35B-A3B3B active / 35B MoEdoes not fit
- Llama 3.3 70B70B densedoes not fit
- gpt-oss 120B5.1B active / 120B MoEdoes not fit
RTX 3060 12GBOpen in the desk
- Llama 3.2 3B3B denseQ8_0near-lossless · 4.6 GB est.
- Qwen3.5 4B4B denseQ8_0near-lossless · 5.5 GB est.
- Gemma 4 E4B4B dense / 8B totalQ8_0near-lossless · 8.6 GB est.
- Llama 3.1 8B8B denseQ8_0near-lossless · 9.5 GB est.
- Ministral 3 8B8B denseQ8_0near-lossless · 10.5 GB est.
- Qwen3.5 9B9.7B denseQ8_0near-lossless · 10.5 GB est.
- Gemma 4 12B12B denseQ6_Khigh · 10.3 GB est.
- DeepSeek-R1 Distill Qwen 14B14.8B denseQ5_K_Mbalanced+ · 11.9 GB est.
- gpt-oss 20B3.6B active / 20B MoEdoes not fit
- Gemma 4 26B-A4B4B active / 26B MoEIQ3_XXScompressed · 11.4 GB est.
- Qwen3.5 27B27B denseIQ2_XXSmax compression · 8.9 GB est.
- Qwen3.8 27B27.8B denseIQ2_XXSmax compression · 8.9 GB est.
- Gemma 4 31B31B denseIQ2_XXSmax compression · 10.7 GB est.
- Qwen3 32B32.8B denseIQ2_XXSmax compression · 11.9 GB est.
- Qwen3.5 35B-A3B3B active / 35B MoEIQ2_XXSmax compression · 10.9 GB est.
- Llama 3.3 70B70B densedoes not fit
- gpt-oss 120B5.1B active / 120B MoEdoes not fit
RTX 5080 16GBOpen in the desk
- Llama 3.2 3B3B denseQ8_0near-lossless · 4.6 GB est.
- Qwen3.5 4B4B denseQ8_0near-lossless · 5.5 GB est.
- Gemma 4 E4B4B dense / 8B totalQ8_0near-lossless · 8.6 GB est.
- Llama 3.1 8B8B denseQ8_0near-lossless · 9.5 GB est.
- Ministral 3 8B8B denseQ8_0near-lossless · 10.5 GB est.
- Qwen3.5 9B9.7B denseQ8_0near-lossless · 10.5 GB est.
- Gemma 4 12B12B denseQ8_0near-lossless · 13 GB est.
- DeepSeek-R1 Distill Qwen 14B14.8B denseQ6_Khigh · 13.4 GB est.
- gpt-oss 20B3.6B active / 20B MoEMXFP4 nativenative · 13.5 GB est.
- Gemma 4 26B-A4B4B active / 26B MoEQ4_K_Mbalanced · 16 GB est.
- Qwen3.5 27B27B denseIQ4_XSsmall+ · 15.3 GB est.
- Qwen3.8 27B27.8B denseIQ4_XSsmall+ · 15.3 GB est.
- Gemma 4 31B31B denseIQ3_XXScompressed · 14.3 GB est.
- Qwen3 32B32.8B denseIQ3_XXScompressed · 15.6 GB est.
- Qwen3.5 35B-A3B3B active / 35B MoEIQ3_XXScompressed · 15.1 GB est.
- Llama 3.3 70B70B densedoes not fit
- gpt-oss 120B5.1B active / 120B MoEdoes not fit
RTX 4090 24GBOpen in the desk
- Llama 3.2 3B3B denseQ8_0near-lossless · 4.6 GB est.
- Qwen3.5 4B4B denseQ8_0near-lossless · 5.5 GB est.
- Gemma 4 E4B4B dense / 8B totalQ8_0near-lossless · 8.6 GB est.
- Llama 3.1 8B8B denseQ8_0near-lossless · 9.5 GB est.
- Ministral 3 8B8B denseQ8_0near-lossless · 10.5 GB est.
- Qwen3.5 9B9.7B denseQ8_0near-lossless · 10.5 GB est.
- Gemma 4 12B12B denseQ8_0near-lossless · 13 GB est.
- DeepSeek-R1 Distill Qwen 14B14.8B denseQ8_0near-lossless · 16.8 GB est.
- gpt-oss 20B3.6B active / 20B MoEMXFP4 nativenative · 13.5 GB est.
- Gemma 4 26B-A4B4B active / 26B MoEQ6_Khigh · 21.3 GB est.
- Qwen3.5 27B27B denseQ6_Khigh · 22.4 GB est.
- Qwen3.8 27B27.8B denseQ6_Khigh · 22.4 GB est.
- Gemma 4 31B31B denseQ5_K_Mbalanced+ · 22.8 GB est.
- Qwen3 32B32.8B denseQ4_K_Mbalanced · 21.2 GB est.
- Qwen3.5 35B-A3B3B active / 35B MoEQ4_K_Mbalanced · 21.2 GB est.
- Llama 3.3 70B70B denseIQ2_XXSmax compression · 23.2 GB est.
- gpt-oss 120B5.1B active / 120B MoEdoes not fit
RTX 5090 32GBOpen in the desk
- Llama 3.2 3B3B denseQ8_0near-lossless · 4.6 GB est.
- Qwen3.5 4B4B denseQ8_0near-lossless · 5.5 GB est.
- Gemma 4 E4B4B dense / 8B totalQ8_0near-lossless · 8.6 GB est.
- Llama 3.1 8B8B denseQ8_0near-lossless · 9.5 GB est.Measured 206.7 tok/sQ4_K_M · 8K context
- Ministral 3 8B8B denseQ8_0near-lossless · 10.5 GB est.
- Qwen3.5 9B9.7B denseQ8_0near-lossless · 10.5 GB est.
- Gemma 4 12B12B denseQ8_0near-lossless · 13 GB est.
- DeepSeek-R1 Distill Qwen 14B14.8B denseQ8_0near-lossless · 16.8 GB est.
- gpt-oss 20B3.6B active / 20B MoEMXFP4 nativenative · 13.5 GB est.
- Gemma 4 26B-A4B4B active / 26B MoEQ8_0near-lossless · 27.3 GB est.Measured 180.2 tok/sQ4_K_M · 8K context
- Qwen3.5 27B27B denseQ8_0near-lossless · 28.8 GB est.
- Qwen3.8 27B27.8B denseQ8_0near-lossless · 28.8 GB est.
- Gemma 4 31B31B denseQ6_Khigh · 25.9 GB est.
- Qwen3 32B32.8B denseQ6_Khigh · 27.8 GB est.
- Qwen3.5 35B-A3B3B active / 35B MoEQ6_Khigh · 28.4 GB est.
- Llama 3.3 70B70B denseIQ3_XXScompressed · 31.2 GB est.
- gpt-oss 120B5.1B active / 120B MoEdoes not fit
§ 003 / HOW TO READ IT
What a cell means
- A quant name means the model fits fully in VRAM at that quant. The table tries the desk's list in order and shows the first that fits: Q8_0, Q6_K, Q5_K_M, Q4_K_M, IQ4_XS, IQ3_XXS, IQ2_XXS. A model that ships in one weight format uses only that format.
- The GB figure is the calculator's estimate for weights, KV cache and runtime overhead. It is an estimate, not a measurement.
- "does not fit" means no quant on that list fits fully in VRAM. The desk can still show a slower run that puts some layers in system RAM.
- "Measured" means we ran that model on that card. The line under it names the quant and context of that run. They can differ from the fit above it.
- The fit reads only the VRAM size. Any card with the same VRAM gets the same answer.
§ 004 / HOW WE PICKED
The cards and the models
- Cards: one discrete card for each size. We take a card with measured runs first, then a desk quick pick, then the first card of that size in the GPU catalog.
- Models: every model the desk ranks in its top 8 for any of these cards, at the desk defaults (use: Chat/Assistant, priority: Quality, context: 8K). We add any model we measured on these cards.