1. Memory size decides which models fit
A model has to fit in memory to run at a usable speed. At the common 4-bit compression, a model needs roughly 0.6 GB per billion parameters, plus room for the conversation. In practice:
- 16 GB runs small and mid-size models (up to about 9 billion parameters) with room to spare.
- 32 GB adds mixture-of-experts models such as gpt-oss 20B (a 12 GB file) and the popular 30B-class models.
- 64 to 128 GB of unified memory is where 70B-class models and 120B-class mixture-of-experts models such as gpt-oss 120B (a 63 GB file) become possible.
2. Memory bandwidth decides speed
Every word the model writes requires reading its active weights from memory once, so speed scales with memory bandwidth. This is why two mini PCs with the same amount of memory can differ by five times in speed. Typical theoretical peaks: about 51 GB/s for dual-channel DDR4-3200, about 90 to 100 GB/s for dual-channel DDR5, and 256 GB/s for AMD's Ryzen AI Max+ 395 with its 256-bit LPDDR5X memory.
3. Can you upgrade it?
Budget and mainstream mini PCs usually take standard SO-DIMM memory you can upgrade later. The fastest AI mini PCs solder their memory next to the chip to reach high bandwidth, so buy the capacity you will need on day one.
What about the NPU?
The NPU is a small AI accelerator. It powers Windows Copilot+ features, which require an NPU of 40 TOPS or more, but most apps for running language models locally use the graphics chip and memory instead. For local AI, treat the NPU as a bonus, not the deciding spec.