Skip to content
[bmdpat]

Local LLM Toolkit

VRAM calculator for local LLMs.

Pick a model, quant, context window, and GPU. The default GGUF check estimates 45/80 GPU layers for Llama 3.3 70B at Q4_K_M on RTX 4090 24GB, then gives you --n-gpu-layers 45 before trial and error.

Calculator

5/5 free runs left today

CPU offload

45/80 layers on GPU

42.1GB required24GB available

Full GPU load needs about 18.1GB more VRAM.

GPU VRAM used

23.9GB

Model weights

40.3GB

KV cache

1.3GB

Speed estimate

8.2-15.7 tok/s

llama.cpp flag
--n-gpu-layers 45

The model can run with CPU offload. Expect lower throughput and more system RAM pressure.

Needs more VRAM

Some links here are affiliate links. If you buy or rent through them I may earn a commission at no extra cost to you. As an Amazon Associate I earn from qualifying purchases.

Your Llama 3.3 70B on RTX 4090 24GB config report

The full report for this setup: recommended quant, the complete VRAM tradeoff table, and the llama.cpp launch command with the right --n-gpu-layers. Sent once, immediately.

Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.

Default target

llama.cpp

Scope

36 model presets / 58 GPUs / 9 quant levels

Recent usage

2 tracked runs / 30d

FAQ

How much VRAM do I need for Llama 3 70B?
Llama 3.3 70B at Q4_K_M weighs about 40.3GB before runtime overhead and KV cache. With the calculator's 4K tokens default context, the full estimate is 42.1GB. A 24GB card can run it with partial GPU offload, but full GPU residency usually needs a 48GB class card, unified-memory machine, or multiple GPUs.
What does --n-gpu-layers do in llama.cpp?
`--n-gpu-layers` controls how many transformer layers llama.cpp keeps on the GPU. Higher values are faster when they fit in VRAM. Lower values spill more work to CPU and system RAM.
Can I run a 70B model on 24GB VRAM?
Yes, but usually not fully in VRAM. On RTX 4090 24GB, the current calculator estimate offloads 45/80 layers and emits `--n-gpu-layers 45` at the default 4K tokens. Use Q4_K_M or smaller, offload as many layers as fit, and expect CPU offload to reduce tokens per second.