Skip to content
AIPCs

Can it run?

Pick a model and a computer. See whether the model fits in memory and roughly how fast it will write.

A token is about three quarters of a word.

Result

Yes. Qwen3.5 9B fits on this machine with a 8K-token conversation.

Estimated writing speed

38 to 49 tokens/s

Fast. We measured 46 tokens/s on this exact setup (T0002).

Model weights
5.7 GB
Conversation memory (8K tokens)
0.3 GB
Runtime overhead
0.8 GB
Total needed
6.7 GB
Available on this machine
7.4 GB

What runs where, at a glance

Estimated writing speed in tokens per second with an 8K-token conversation. "No" means it does not fit; an asterisk means part of the model runs from system RAM (graphics cards assume 32 GB of DDR5 system memory).

Estimated tokens per second by model and computer, 8K context
Computer Gemma 4 E4BQwen3.5 9Bgpt-oss 20BQwen3.8 27BQwen3.6 35B-A3BLlama 3.3 70Bgpt-oss 120B
NVIDIA RTX 4060 8 GB + 32 GB DDR5-6400 (our desktop) 644435*4*40*NoNo
NVIDIA RTX 5060 Ti 16 GB 99701289*63*NoNo
NVIDIA RTX 4090 24 GB 188144246522442*No
NVIDIA RTX 5090 32 GB 230182303683002*No
Apple M5 Pro, 64 GB (Mac mini 2026) 61407713765No
AMD Ryzen AI Max+ 395, 128 GB 5233651164447
NVIDIA DGX Spark (GB10), 128 GB 5536701269450
Apple M5 Max 40-core GPU, 128 GB (Mac Studio, MacBook Pro 2026) 875611018108778
Ryzen AI 9 HX 370 mini PC, 96 GB DDR5-5600 14918318113
32 GB dual-channel DDR5-5600 14918318NoNo

How much memory each model needs

Model file size plus conversation memory plus about 0.8 GB of runtime overhead. Newer models store far less conversation memory per token, which is why a long chat costs Llama 3.3 70B gigabytes but Qwen3.6 35B-A3B only a fraction of that.

ModelFile8K chat32K128KPer 1K tokens
Gemma 4 E4B Q4_0 4.59 GB 5.5 GB 5.9 GB 7.6 GB 17 MB
Llama 3.1 8B Q4_K_M 4.92 GB 6.8 GB 10 GB 23 GB 134 MB
Qwen3.5 9B Q4_K_M 5.68 GB 6.7 GB 7.6 GB 11 GB 34 MB
Gemma 4 12B Q4_K_M 7.12 GB 8.4 GB 8.8 GB 10 GB 17 MB
gpt-oss 20B MXFP4 12.11 GB 13 GB 14 GB 16 GB 25 MB
Qwen3.8 27B Q4_K_M 16.46 GB 18 GB 19 GB 26 GB 67 MB
Qwen3 Coder 30B-A3B Q4_K_M 18.56 GB 20 GB 23 GB 32 GB 101 MB
Qwen3.6 35B-A3B Q4_K_M 22.13 GB 23 GB 24 GB 26 GB 21 MB
Llama 3.3 70B Q4_K_M 42.52 GB 46 GB 54 GB 86 GB 336 MB
gpt-oss 120B MXFP4 63.39 GB 64 GB 65 GB 69 GB 38 MB

How the calculator works

  • Memory: we read each model's exact file size and its layer layout straight from the model file on Hugging Face, so conversation memory reflects how that model actually stores context (many new models only keep it in some layers).
  • Available memory: a graphics card's video memory minus 0.6 GB; 75% of unified memory (the usual default limit on Macs, which you can often raise); system RAM minus 4 GB for the operating system.
  • Speed: writing speed is limited by memory bandwidth, because each new token reads the model's active weights once. We use the share of bandwidth our own tests achieve: 90% on a graphics card and 55% on a desktop processor, measured on our RTX 4060 and Ryzen 9 7900X. For unified-memory machines we assume 70% until we have tested one. Very fast memory is harder to keep busy, so above about 300 GB/s (unified) or 1,000 GB/s (graphics cards) we scale the share down, in line with published results for Apple Ultra chips and the RTX 5090, and add about 1 ms of fixed time per token on graphics cards. AMD and Intel cards get a brand correction from published measurements, because their software currently gets less out of the same hardware.
  • Split across graphics card and RAM: the overflow runs at system-memory speed, with a further 20% penalty measured on our desktop. Mixture-of-experts models keep their shared weights on the graphics card, which is why they cope with this far better than dense models.
  • Accuracy: on our desktop every measured result falls inside the estimated range. Our slower ACEMAGIC K1 mini PC ran 15% to 36% faster than estimated, so for machines with slow memory our figures are cautious. Treat other machines' figures as good estimates, not guarantees. See our methodology and raw benchmarks.