LLM VRAM & Hardware Fit Calculator
Check if your GPU or Mac can run local LLMs — VRAM estimates by model size, quantization and context length.
| Model | Memory | Verdict | Est. speed (tok/s) | Pool used |
|---|---|---|---|---|
Runs on GPU | 9.41 GB | Perfect fit | 132 | 59% |
Gemma 3 12BMid-size · General Runs on GPU | 7.25 GB | Perfect fit | 26.4 | 45% |
Mistral Nemo 12BMid-size · General Runs on GPU | 7.36 GB | Perfect fit | 26.0 | 46% |
Phi-4 14BMid-size · Reasoning Runs on GPU | 8.33 GB | Perfect fit | 22.6 | 52% |
Qwen3 14BMid-size · General Runs on GPU | 8.76 GB | Perfect fit | 21.4 | 55% |
Mistral 7B v0.3Light · General Runs on GPU | 4.62 GB | Perfect fit | 44.0 | 29% |
Qwen2.5 7BLight · General Runs on GPU | 4.84 GB | Perfect fit | 41.7 | 30% |
Qwen2.5 Coder 7BLight · Coding Runs on GPU | 4.84 GB | Perfect fit | 41.7 | 30% |
DeepSeek-R1 Distill 7BLight · Reasoning Runs on GPU | 4.84 GB | Perfect fit | 41.7 | 30% |
Llama 3.1 8BLight · General Runs on GPU | 5.06 GB | Perfect fit | 39.6 | 32% |
Ministral 8BLight · General Runs on GPU | 5.06 GB | Perfect fit | 39.6 | 32% |
Qwen3 8BLight · General Runs on GPU | 5.17 GB | Perfect fit | 38.6 | 32% |
GLM-4 9BLight · General Runs on GPU | 5.61 GB | Perfect fit | 35.2 | 35% |
Gemma 2 9BLight · General Runs on GPU | 5.72 GB | Perfect fit | 34.4 | 36% |
Qwen3 1.7BTiny · General Runs on GPU | 1.54 GB | Perfect fit | 186 | 10% |
Llama 3.2 3BTiny · General Runs on GPU | 2.40 GB | Perfect fit | 99.0 | 15% |
Phi-4-miniTiny · General Runs on GPU | 2.73 GB | Perfect fit | 83.4 | 17% |
Gemma 3 4BTiny · General Runs on GPU | 2.85 GB | Perfect fit | 79.2 | 18% |
Qwen3 4BTiny · General Runs on GPU | 2.85 GB | Perfect fit | 79.2 | 18% |
MoE expert offload | 17.2 GB | Runs well | 76.8 | 49% |
MoE expert offload | 25.8 GB | Runs well | 19.6 | 73% |
Qwen3 32BPro · General Partial CPU+GPU offload | 18.4 GB | Runs well | 2.90 | 52% |
DeepSeek-R1 Distill 32BPro · Reasoning Partial CPU+GPU offload | 18.4 GB | Runs well | 2.90 | 52% |
Gemma 3 27BPro · General Runs on GPU | 15.3 GB | Tight fit | 11.7 | 96% |
Mistral Small 3.1 24BMid-size · General Runs on GPU | 13.7 GB | Tight fit | 13.2 | 86% |
Llama 3.3 70BFlagship · General Runs on GPU | 38.4 GB | Won't fit | — | 240% |
Qwen2.5 72BFlagship · General Runs on GPU | 39.5 GB | Won't fit | — | 247% |
Runs on GPU | 58.5 GB | Won't fit | — | 366% |
Runs on GPU | 124 GB | Won't fit | — | 775% |
Runs on GPU | 348 GB | Won't fit | — | 2178% |
All figures are formula-based estimates anchored on the open-source llmfit project. Real throughput varies by runtime (Ollama, llama.cpp, LM Studio, MLX), prompt length, and kernel quality. For measured numbers on your machine, run llmfit bench.
Approximate VRAM by model size
Rough planning estimates for ~8k context, Q4 weights plus KV cache. Real usage varies by runtime, context length and MoE offload.
| Model class | Quant | Approx. VRAM | Typical hardware |
|---|---|---|---|
| 7–8B dense | Q4_K_M | 5–6 GB | 8 GB GPU or Mac 8–16 GB |
| 7–8B dense | Q8_0 | 9–10 GB | 12–16 GB GPU |
| 13–14B dense | Q4_K_M | 9–11 GB | 12–16 GB GPU |
| 27–32B dense | Q4_K_M | 16–20 GB | 24 GB GPU (e.g. RTX 4090) |
| 70–72B dense | Q4_K_M | 35–42 GB | Multi-GPU or Mac 48–64 GB+ |
| MoE (~8×7B active) | Q4_K_M | 20–28 GB* | 24–32 GB+; experts may offload |
LLM hardware FAQ
How much VRAM do I need to run a 7B LLM?
Most 7–8B models at Q4 quantization need about 5–6 GB for weights plus KV cache at a short context. An 8 GB GPU can often run them; longer contexts or Q8 need closer to 10 GB.
Can I run a 70B model on a single GPU?
Usually not with full weights on one consumer card. Q4 70B-class models often need ~35–42 GB. You typically need multi-GPU, heavy offload to system RAM, or a Mac with large unified memory.
What quantization is best for a local LLM?
Q4_K_M is the common default: good quality-to-size trade-off. Use Q5/Q6 when you have spare VRAM; Q8 when quality matters more than speed or when the model is small enough to fit.
Does Mac unified memory work for LLM inference?
Yes. Apple Silicon can run many models because CPU and GPU share memory. A 32–64 GB Mac can often host mid-size models that would need multi-GPU on discrete setups, with bandwidth-limited speed.
Does longer context length use more VRAM?
Yes. The KV cache grows with context length and batch size. A model that “fits” at 4k tokens may OOM at 32k — always leave headroom beyond weights alone.