Local AI sizing desk
What can your GPU run?
Choose your card. Get the models that fit, the quant to use, and a command you can paste.
Evidence and method
58 supported GPU profiles · 40 models · 8 benchmark runs on cards we own
Fine-tune results · Chat / Assistant · Quality · 8K
Your first run
Start with your hardware.
Search above, or use a popular setup to get an exact first run.
Popular
Common questions
What can my GPU run locally?
The Local LLM Sizing Desk at bmdpat.com/desk covers 58 verified GPU profiles. Pick your card and it lists which of 40 local models fit in your VRAM, the GGUF quant to use, and the exact llama.cpp run command to paste.
Are the tokens-per-second numbers measured or estimated?
Both, and the desk labels which is which. It carries 8 benchmark runs measured on cards the site's author owns; everything else is a VRAM-fit calculation, shown as a calculation rather than a measurement.
Which GGUF quant should I pick for my VRAM?
Pick the largest quant that fits your model, context length, and KV cache with VRAM headroom. Q4_K_M is the fit-first choice, Q5 or Q6 is the balance point, and Q8_0 is for quality when memory is available. The desk does this arithmetic per GPU and model.
Is the Local LLM Sizing Desk free?
Yes: 5 free runs per tool per day with no account. Pro at $19.00/month or $149.00/year adds saved rigs with new-model fit alerts, Hugging Face import, benchmark history, and unlimited runs.
Pricing