How much memory each model needs
Model file size plus conversation memory plus about 0.8 GB of runtime overhead. Newer models store far less conversation memory per token, which is why a long chat costs Llama 3.3 70B gigabytes but Qwen3.6 35B-A3B only a fraction of that.
| Model | File | 8K chat | 32K | 128K | Per 1K tokens |
|---|---|---|---|---|---|
| Gemma 4 E4B Q4_0 | 4.59 GB | 5.5 GB | 5.9 GB | 7.6 GB | 17 MB |
| Llama 3.1 8B Q4_K_M | 4.92 GB | 6.8 GB | 10 GB | 23 GB | 134 MB |
| Qwen3.5 9B Q4_K_M | 5.68 GB | 6.7 GB | 7.6 GB | 11 GB | 34 MB |
| Gemma 4 12B Q4_K_M | 7.12 GB | 8.4 GB | 8.8 GB | 10 GB | 17 MB |
| gpt-oss 20B MXFP4 | 12.11 GB | 13 GB | 14 GB | 16 GB | 25 MB |
| Qwen3.8 27B Q4_K_M | 16.46 GB | 18 GB | 19 GB | 26 GB | 67 MB |
| Qwen3 Coder 30B-A3B Q4_K_M | 18.56 GB | 20 GB | 23 GB | 32 GB | 101 MB |
| Qwen3.6 35B-A3B Q4_K_M | 22.13 GB | 23 GB | 24 GB | 26 GB | 21 MB |
| Llama 3.3 70B Q4_K_M | 42.52 GB | 46 GB | 54 GB | 86 GB | 336 MB |
| gpt-oss 120B MXFP4 | 63.39 GB | 64 GB | 65 GB | 69 GB | 38 MB |
How the calculator works
- Memory: we read each model's exact file size and its layer layout straight from the model file on Hugging Face, so conversation memory reflects how that model actually stores context (many new models only keep it in some layers).
- Available memory: a graphics card's video memory minus 0.6 GB; 75% of unified memory (the usual default limit on Macs, which you can often raise); system RAM minus 4 GB for the operating system.
- Speed: writing speed is limited by memory bandwidth, because each new token reads the model's active weights once. We use the share of bandwidth our own tests achieve: 90% on a graphics card and 55% on a desktop processor, measured on our RTX 4060 and Ryzen 9 7900X. For unified-memory machines we assume 70% until we have tested one. Very fast memory is harder to keep busy, so above about 300 GB/s (unified) or 1,000 GB/s (graphics cards) we scale the share down, in line with published results for Apple Ultra chips and the RTX 5090, and add about 1 ms of fixed time per token on graphics cards. AMD and Intel cards get a brand correction from published measurements, because their software currently gets less out of the same hardware.
- Split across graphics card and RAM: the overflow runs at system-memory speed, with a further 20% penalty measured on our desktop. Mixture-of-experts models keep their shared weights on the graphics card, which is why they cope with this far better than dense models.
- Accuracy: on our desktop every measured result falls inside the estimated range. Our slower ACEMAGIC K1 mini PC ran 15% to 36% faster than estimated, so for machines with slow memory our figures are cautious. Treat other machines' figures as good estimates, not guarantees. See our methodology and raw benchmarks.