How much memory do you need to run AI models locally?
Quick answer
At the usual 4-bit compression, a model needs about 0.6 GB per billion parameters, plus memory for the conversation, plus about 1 GB for the software. That works out to an 8 GB graphics card for models up to about 9 billion parameters, 16 GB for gpt-oss 20B, 24 GB for the popular 27B to 35B models, about 64 GB of unified memory for a 70B model, and 96 GB or more for 120B-class models such as gpt-oss 120B.
The question sounds simple, but most answers online are rules of thumb that ignore two things: how much memory the conversation itself takes, and how differently new models are built. So we read the numbers straight from the model files. Every figure on this page comes from each model’s actual download size and its layer layout, and our calculator uses the same data.
The formula
Memory needed = model file + conversation memory + about 1 GB of overhead.
- Model file. At the common 4-bit compression (called Q4_K_M or MXFP4 in download names), a model takes about 0.6 GB per billion parameters. An 8B model is about 5 GB; a 70B model about 42.5 GB.
- Conversation memory (the “KV cache”). Every token of the conversation so far is kept in memory. How much each token costs varies more than tenfold between models.
- Overhead. The software needs working space; we allow 0.8 GB, plus whatever your operating system and display use.
How much each model needs
For a long chat of about 8,000 tokens (roughly 6,000 words):
| Model (4-bit) | Memory needed | Smallest graphics card | Smallest unified memory |
|---|---|---|---|
| Gemma 4 E4B 4.59 GB file | 5.5 GB | 8 GB | 16 GB |
| Llama 3.1 8B 4.92 GB file | 6.8 GB | 8 GB | 16 GB |
| Qwen3.5 9B 5.68 GB file | 6.7 GB | 8 GB | 16 GB |
| Gemma 4 12B 7.12 GB file | 8.4 GB | 12 GB | 16 GB |
| gpt-oss 20B 12.11 GB file, mixture of experts | 13 GB | 16 GB | 24 GB |
| Qwen3.8 27B 16.46 GB file | 18 GB | 24 GB | 24 GB |
| Qwen3 Coder 30B-A3B 18.56 GB file, mixture of experts | 20 GB | 24 GB | 32 GB |
| Qwen3.6 35B-A3B 22.13 GB file, mixture of experts | 23 GB | 24 GB | 32 GB |
| Llama 3.3 70B 42.52 GB file | 46 GB | None (split with RAM) | 64 GB |
| gpt-oss 120B 63.39 GB file, mixture of experts | 64 GB | None (split with RAM) | 96 GB |
“Smallest graphics card” assumes the card’s memory minus 0.6 GB is usable. “Smallest unified memory” assumes AI can use 75% of it, the usual default on Macs. For a computer without a graphics card, subtract about 4 GB for the operating system from your RAM.
Why long conversations need more memory
Here the gap between older and newer models is dramatic. With a 32K-token conversation (32,768 tokens: a long document, or a long working session):
| Model | Conversation memory per token | For 32K tokens |
|---|---|---|
| Llama 3.3 70B | 320 KB | 10.7 GB |
| Llama 3.1 8B | 128 KB | 4.3 GB |
| Qwen3 Coder 30B-A3B | 96 KB | 3.2 GB |
| Qwen3.5 9B | 32 KB | 1.1 GB |
| gpt-oss 20B | 24 KB | 0.8 GB |
| Qwen3.6 35B-A3B | 20 KB | 0.7 GB |
Newer designs keep full conversation memory in only some of their layers (Qwen3.5 and Qwen3.6 in one layer out of four), or let many layers look only at recent tokens (half the layers in gpt-oss, five in six in Gemma 4). That is why Qwen3.6 35B-A3B needs about a sixth of the conversation memory of Llama 3.1 8B, a much smaller model from 2024.
Graphics card, unified memory or system RAM?
- Graphics card: the fastest option, but you are limited to the card’s video memory, 8 to 32 GB on consumer cards. Models that don’t fit can spill into system RAM at a large speed cost, except for mixture-of-experts models, which cope well. See what an 8 GB card can run.
- Unified memory (Apple Macs, AMD Ryzen AI Max, NVIDIA DGX Spark): the graphics processor shares up to 128 GB or more with the processor. Slower than a big graphics card, but the only affordable way to run 70B and 120B-class models.
- System RAM only: the cheapest way to start. Fine for small and mixture-of-experts models. See running AI without a graphics card.
Other ways to fit a bigger model
- Stronger compression. 3-bit and 2-bit versions of models are smaller, at a growing cost in answer quality. 4-bit is the usual sweet spot; 8-bit (about 1.06 GB per billion parameters) is close to the original quality but nearly doubles the memory needed.
- Compressing the conversation memory. Many apps can store the conversation at 8 bits instead of 16, halving that part.
- A shorter conversation. Most apps let you cap the context length. A smaller cap leaves more room for the model.
What to buy
Use our calculator with the model you actually want. As a starting point: 16 GB graphics cards for up to 20B-class models, 24 to 32 GB cards for the 27B to 35B generation, and 64 to 128 GB of unified memory for 70B and larger. Our guides to mini PCs, laptops and desktops cover the options.
Questions people ask
Is 16 GB of RAM enough for local AI?
For small models, yes. A computer with 16 GB of system RAM and no graphics card has about 12 GB free for AI, enough for models up to about 12 billion parameters at 4 bits. A 16 GB graphics card can hold more, including gpt-oss 20B, because it doesn't share its memory with the operating system.
Can I combine video memory and system RAM?
Yes. llama.cpp and the apps built on it can keep part of a model in system RAM when it doesn't fit on the graphics card. The part in system RAM runs at system-memory speed, so it works well for mixture-of-experts models and poorly for large dense models. Our RTX 4060 tests show both cases.
Does the conversation length really matter?
It can. Each token of conversation is stored in memory. For Llama 3.3 70B that is 320 KB per token, so a 32K-token conversation (32,768 tokens) adds about 10.7 GB. Newer models are far more efficient: Qwen3.6 35B-A3B needs 20 KB per token, or about 0.7 GB for the same conversation.
How much of a Mac's unified memory can AI use?
By default macOS lets the graphics processor use roughly two thirds to three quarters of unified memory, and the limit can be raised in software. We assume 75% in our calculator, so a 64 GB Mac gives about 48 GB to a model.
What changed
- Sep 25, 2026: First published. Figures calculated from each model's file on Hugging Face.