[bmdpat]

Local AI sizing desk

What can your GPU run?

Choose your card. Get the models that fit, the quant to use, and a command you can paste.

Evidence and method

58 supported GPU profiles · 40 models · 8 benchmark runs on cards we own

Fine-tune results · Chat / Assistant · Quality · 8K
Workload
Optimize for
Context length

Your first run

Start with your hardware.

Search above, or use a popular setup to get an exact first run.

Popular

Common questions

What can my GPU run locally?

The Local LLM Sizing Desk at bmdpat.com/desk covers 58 verified GPU profiles. Pick your card and it lists which of 40 local models fit in your VRAM, the GGUF quant to use, and the exact llama.cpp run command to paste.

Are the tokens-per-second numbers measured or estimated?

Both, and the desk labels which is which. It carries 8 benchmark runs measured on cards the site's author owns; everything else is a VRAM-fit calculation, shown as a calculation rather than a measurement.

Which GGUF quant should I pick for my VRAM?

Pick the largest quant that fits your model, context length, and KV cache with VRAM headroom. Q4_K_M is the fit-first choice, Q5 or Q6 is the balance point, and Q8_0 is for quality when memory is available. The desk does this arithmetic per GPU and model.

Is the Local LLM Sizing Desk free?

Yes: 5 free runs per tool per day with no account. Pro at $19.00/month or $149.00/year adds saved rigs with new-model fit alerts, Hugging Face import, benchmark history, and unlimited runs.

Pricing

Save rigs and keep fit alerts in one place.

See Pro
Free · $0Monthly · $19/moYearly · $149/year