Skip to main content
Back to Developer Tools

LLM VRAM & Hardware Fit Calculator

Check if your GPU or Mac can run local LLMs — VRAM estimates by model size, quantization and context length.

Check Your Hardware
Select your hardware, quantization, and context length. Results update instantly — everything runs locally in your browser.
Available memory
16.0 GB
Memory bandwidth
288 GB/s
Largest model @ 85% pool
≈ 23B · Q4_K_M
Model Compatibility
25 of 30 models can run on this hardware
ModelMemoryVerdictEst. speed (tok/s)Pool used
DeepSeek Coder V2 LiteMoEMid-size · Coding
Runs on GPU
9.41 GBPerfect fit132
59%
Gemma 3 12BMid-size · General
Runs on GPU
7.25 GBPerfect fit26.4
45%
Mistral Nemo 12BMid-size · General
Runs on GPU
7.36 GBPerfect fit26.0
46%
Phi-4 14BMid-size · Reasoning
Runs on GPU
8.33 GBPerfect fit22.6
52%
Qwen3 14BMid-size · General
Runs on GPU
8.76 GBPerfect fit21.4
55%
Mistral 7B v0.3Light · General
Runs on GPU
4.62 GBPerfect fit44.0
29%
Qwen2.5 7BLight · General
Runs on GPU
4.84 GBPerfect fit41.7
30%
Qwen2.5 Coder 7BLight · Coding
Runs on GPU
4.84 GBPerfect fit41.7
30%
DeepSeek-R1 Distill 7BLight · Reasoning
Runs on GPU
4.84 GBPerfect fit41.7
30%
Llama 3.1 8BLight · General
Runs on GPU
5.06 GBPerfect fit39.6
32%
Ministral 8BLight · General
Runs on GPU
5.06 GBPerfect fit39.6
32%
Qwen3 8BLight · General
Runs on GPU
5.17 GBPerfect fit38.6
32%
GLM-4 9BLight · General
Runs on GPU
5.61 GBPerfect fit35.2
35%
Gemma 2 9BLight · General
Runs on GPU
5.72 GBPerfect fit34.4
36%
Qwen3 1.7BTiny · General
Runs on GPU
1.54 GBPerfect fit186
10%
Llama 3.2 3BTiny · General
Runs on GPU
2.40 GBPerfect fit99.0
15%
Phi-4-miniTiny · General
Runs on GPU
2.73 GBPerfect fit83.4
17%
Gemma 3 4BTiny · General
Runs on GPU
2.85 GBPerfect fit79.2
18%
Qwen3 4BTiny · General
Runs on GPU
2.85 GBPerfect fit79.2
18%
Qwen3 30B A3BMoEPro · General
MoE expert offload
17.2 GBRuns well76.8
49%
Mixtral 8x7BMoEPro · General
MoE expert offload
25.8 GBRuns well19.6
73%
Qwen3 32BPro · General
Partial CPU+GPU offload
18.4 GBRuns well2.90
52%
DeepSeek-R1 Distill 32BPro · Reasoning
Partial CPU+GPU offload
18.4 GBRuns well2.90
52%
Gemma 3 27BPro · General
Runs on GPU
15.3 GBTight fit11.7
96%
Mistral Small 3.1 24BMid-size · General
Runs on GPU
13.7 GBTight fit13.2
86%
Llama 3.3 70BFlagship · General
Runs on GPU
38.4 GBWon't fit—
240%
Qwen2.5 72BFlagship · General
Runs on GPU
39.5 GBWon't fit—
247%
Llama 4 ScoutMoEFlagship · General
Runs on GPU
58.5 GBWon't fit—
366%
Qwen3 235B A22BMoEFlagship · General
Runs on GPU
124 GBWon't fit—
775%
DeepSeek-R1MoEFlagship · Reasoning
Runs on GPU
348 GBWon't fit—
2178%

All figures are formula-based estimates anchored on the open-source llmfit project. Real throughput varies by runtime (Ollama, llama.cpp, LM Studio, MLX), prompt length, and kernel quality. For measured numbers on your machine, run llmfit bench.

Approximate VRAM by model size

Rough planning estimates for ~8k context, Q4 weights plus KV cache. Real usage varies by runtime, context length and MoE offload.

Model classQuantApprox. VRAMTypical hardware
7–8B denseQ4_K_M5–6 GB8 GB GPU or Mac 8–16 GB
7–8B denseQ8_09–10 GB12–16 GB GPU
13–14B denseQ4_K_M9–11 GB12–16 GB GPU
27–32B denseQ4_K_M16–20 GB24 GB GPU (e.g. RTX 4090)
70–72B denseQ4_K_M35–42 GBMulti-GPU or Mac 48–64 GB+
MoE (~8×7B active)Q4_K_M20–28 GB*24–32 GB+; experts may offload

LLM hardware FAQ

How much VRAM do I need to run a 7B LLM?

Most 7–8B models at Q4 quantization need about 5–6 GB for weights plus KV cache at a short context. An 8 GB GPU can often run them; longer contexts or Q8 need closer to 10 GB.

Can I run a 70B model on a single GPU?

Usually not with full weights on one consumer card. Q4 70B-class models often need ~35–42 GB. You typically need multi-GPU, heavy offload to system RAM, or a Mac with large unified memory.

What quantization is best for a local LLM?

Q4_K_M is the common default: good quality-to-size trade-off. Use Q5/Q6 when you have spare VRAM; Q8 when quality matters more than speed or when the model is small enough to fit.

Does Mac unified memory work for LLM inference?

Yes. Apple Silicon can run many models because CPU and GPU share memory. A 32–64 GB Mac can often host mid-size models that would need multi-GPU on discrete setups, with bandwidth-limited speed.

Does longer context length use more VRAM?

Yes. The KV cache grows with context length and batch size. A model that “fits” at 4k tokens may OOM at 32k — always leave headroom beyond weights alone.