Memory bandwidth predicts local AI speed: 23 chips, one model
Quick answer
How fast a computer writes with a local AI model is set mostly by its memory bandwidth. Across 23 processors and graphics chips running the same 7-billion-parameter model, most reached 55% to 85% of the speed their bandwidth allows; our RTX 4060 (272 GB/s) wrote 65.5 tokens per second, our Ryzen 9 7900X alone (102 GB/s) 15.5, and a $369 mini PC with one memory stick (21 GB/s) 3.9. A good estimate for any machine: tokens per second is about 0.7 times the bandwidth in GB/s, divided by the model size in GB.
Every time a local AI model writes a word, it reads its weights from memory. For a 7-billion-parameter model compressed to 4 bits, that is about 3.75 GB of data per token. So a simple rule should hold: the faster the memory, the faster the answer. We tested that rule on four chips in our own two machines and checked it against published results for 19 more.
The rule, and why it works
Divide a machine’s memory bandwidth by the amount of data the model reads per token and you get the fastest it could possibly write. Our RTX 4060 has 272 GB/s of memory bandwidth, so with Llama 2 7B (3.75 GB read per token) the limit is 272 / 3.75 = about 72 tokens per second. We measured 65.5, or 90% of the limit.
The same arithmetic works for a processor with no graphics card. Our Ryzen 9 7900X has dual-channel DDR5-6400 memory, a theoretical 102.4 GB/s, which allows about 27 tokens per second. We measured 15.5, or 57% of the limit. Desktop processors reach a smaller share because they cannot keep their memory as busy as a graphics card can.
| Model | Where it ran | Reading the prompttokens/s | Writing the answertokens/s | Test |
|---|---|---|---|---|
| Llama 2 7B Q4_0 | NVIDIA GeForce RTX 4060 8 GB Whole model on GPU | 2,716 | 65.5 | T0002 |
| Llama 2 7B Q4_0 | AMD Ryzen 9 7900X (CPU only) CPU only | 134 | 15.5 | T0001 |
23 chips, one model
We chose Llama 2 7B at 4 bits (Q4_0) for our reference test because it is the model the llama.cpp community has used to benchmark hardware for years. That means we can place our results next to published ones measured the same way, from an M1 MacBook to an RTX 5090.
- Our tests
- Apple Silicon (published)
- PC graphics and AI chips (published)
Hover a point for its value. Every point is listed in the table below.
The points hug a straight line: across an 80-fold range of bandwidth, writing speed rises almost in step with it. Most reach between 55% and 85% of their theoretical limit.
| Machine | Bandwidth (GB/s) | Writing speed (tokens/s) | Share of limit | Source |
|---|---|---|---|---|
| AMD Ryzen 3 4300U (CPU only) | 21.3 | 3.9 | 69% | Our test T0003 |
| AMD Radeon Graphics, integrated (Ryzen 3 4300U) | 21.3 | 4.8 | 85% | Our test T0004 |
| Apple M1 (8-core GPU) | 68 | 14.2 | 78% | llama.cpp #4167 |
| Apple M2 (10-core GPU) | 100 | 21.9 | 82% | llama.cpp #4167 |
| AMD Ryzen 9 7900X (CPU only) | 102.4 | 15.5 | 57% | Our test T0001 |
| Apple M4 (10-core GPU) | 120 | 24.1 | 75% | llama.cpp #4167 |
| Apple M3 Pro (18-core GPU) | 150 | 30.7 | 77% | llama.cpp #4167 |
| Apple M2 Pro (19-core GPU) | 200 | 38.9 | 73% | llama.cpp #4167 |
| AMD Ryzen AI Max+ 395 (Radeon 8060S), ROCm | 256 | 48.0 | 70% | llama.cpp #15021 |
| NVIDIA GeForce RTX 4060 8 GB | 272 | 65.5 | 90% | Our test T0002 |
| Apple M4 Pro (20-core GPU) | 273 | 50.7 | 70% | llama.cpp #4167 |
| NVIDIA GB10 (DGX Spark), Vulkan | 273 | 58.5 | 80% | llama.cpp #10879 |
| Apple M5 Pro (20-core GPU) | 307 | 66.3 | 81% | llama.cpp #4167 |
| Apple M3 Max (40-core GPU) | 400 | 66.3 | 62% | llama.cpp #4167 |
| Apple M4 Max (40-core GPU) | 546 | 83.1 | 57% | llama.cpp #4167 |
| Intel Arc A770 16 GB, Vulkan | 560 | 45.2 | 30% | llama.cpp #10879 |
| Apple M5 Max (40-core GPU) | 614 | 119.9 | 73% | llama.cpp #4167 |
| Apple M2 Ultra (76-core GPU) | 800 | 94.3 | 44% | llama.cpp #4167 |
| Apple M3 Ultra (80-core GPU) | 800 | 92.1 | 43% | llama.cpp #4167 |
| NVIDIA RTX 3090, Vulkan | 936 | 164.1 | 66% | llama.cpp #10879 |
| AMD RX 7900 XTX, Vulkan | 960 | 182.6 | 71% | llama.cpp #10879 |
| NVIDIA RTX 4090, Vulkan | 1008 | 188.0 | 70% | llama.cpp #10879 |
| NVIDIA RTX 5090, Vulkan | 1792 | 263.6 | 55% | llama.cpp #10879 |
What the exceptions tell you
- The very fastest chips fall further behind. An RTX 5090 (1,792 GB/s) writes at 264 tokens per second, about 55% of its limit, and Apple’s Ultra chips reach under 50%. With this small model, they finish reading the weights so quickly that computation and software overheads start to dominate. With larger models they get closer to their limit.
- Software matters. The Intel Arc A770 has twice the bandwidth of our RTX 4060 but wrote slower in its published result (45 versus 65.5 tokens per second). Drivers and how well a chip’s backend is optimized can cost more than half the potential speed.
- Unified-memory machines land in the middle. AMD’s Ryzen AI Max+ 395 (256 GB/s) wrote at 48 tokens per second and NVIDIA’s DGX Spark (273 GB/s) at 58.5, 70% and 80% of their limits. What they offer is not speed but capacity: up to 128 GB, enough for models no consumer graphics card can hold.
The rule of thumb
For buying decisions, this is accurate enough:
Tokens per second ≈ 0.7 × memory bandwidth (GB/s) ÷ model size (GB)
Use about 0.85 for a modern NVIDIA graphics card and about 0.55 for a desktop processor on its own. Slow memory is easier to keep busy: our budget mini PC reached about 70% on its processor and 85% on its integrated graphics. Two refinements make it work for current models:
- For mixture-of-experts models, use the active size, not the file size. gpt-oss 20B is a 12 GB file, but only about 2.6 GB is read per token, which is why our Ryzen 9 7900X alone writes it at 21.8 tokens per second, faster than the much smaller Qwen3.5 9B (10.4).
- Long conversations add reading. The conversation memory (KV cache) is read for every token too. On our RTX 4060, Llama 2 7B slowed from 65.5 to 43.3 tokens per second with 4,096 tokens of conversation, because that older model stores a lot per token. Newer models store far less and barely slow down.
Our Can it run? calculator applies all of this for you, using each model’s real file size, active size and conversation memory, read directly from the model files.
What this means when you buy
- Memory capacity decides which models you can run; bandwidth decides how fast. Check both.
- For models up to about 9 billion parameters, an 8 GB graphics card like the RTX 4060 is already fast (45 to 68 tokens per second in our tests).
- For bigger models, you are choosing between a graphics card with more video memory (fast, but typically 8 to 32 GB on consumer cards) and a unified-memory machine with 64 to 128 GB (slower, but it fits much larger models).
- Spec sheets that lead with NPU “TOPS” are measuring something else. For local language models, look for the memory bandwidth figure.
The slow end: a $369 mini PC
Our ACEMAGIC K1 has a 4-core Ryzen 3 4300U and ships with a single 16 GB stick of DDR4-2666. With one stick, memory runs on one channel, so it has only 21.3 GB/s, the lowest bandwidth in this study. It lands where the rule predicts: 3.9 tokens per second on its processor (69% of its limit) and 4.8 on its integrated graphics (85%).
| Model | Where it ran | Reading the prompttokens/s | Writing the answertokens/s | Test |
|---|---|---|---|---|
| Llama 2 7B Q4_0 | AMD Ryzen 3 4300U (CPU only) CPU only | 20.5 | 3.94 | T0003 |
| Llama 2 7B Q4_0 | AMD Radeon Graphics, integrated (Ryzen 3 4300U) Whole model on GPU | 38.9 | 4.82 | T0004 |
The lesson for budget machines: check that memory runs in dual channel. The K1 has a second, empty slot, and a matching stick would double its bandwidth to 42.7 GB/s.
Questions people ask
Does the NPU matter for running language models locally?
Not much today. Writing speed is limited by how fast memory can be read, not by how many operations the chip can do, and most local AI apps run models on the graphics processor or CPU rather than the NPU. The NPU matters for Windows Copilot+ features.
Why is reading the prompt so much faster than writing the answer?
When a model reads your prompt it can process hundreds of tokens in one pass over its weights, so the chip's compute power is the limit. When it writes, it produces one token at a time and must read all its active weights for each one, so memory bandwidth is the limit. On our RTX 4060, Llama 2 7B reads at about 2,700 tokens per second but writes at 65.
How do I find my computer's memory bandwidth?
Graphics cards list it on the maker's spec sheet. For system memory, multiply the number of memory channels by 8 bytes and by the memory speed: dual-channel DDR5-6400 is 2 x 8 x 6400 = 102.4 GB/s. Apple lists unified memory bandwidth in each Mac's specifications; for other machines, check the chip maker's spec page.
Why do the fastest chips fall further below the limit?
At very high bandwidth, reading the weights stops being the only bottleneck and the chip's compute, software overheads and synchronization start to matter. That is why an RTX 5090 reaches about 55% of its theoretical speed on this small model while an RTX 4060 reaches 90%.
Sources
- llama.cpp discussion #4167: Performance of llama.cpp on Apple Silicon M-series, checked Sep 25, 2026
- llama.cpp discussion #10879: Performance of llama.cpp with Vulkan, checked Sep 25, 2026
- llama.cpp discussion #15021: Performance of llama.cpp on AMD ROCm (HIP), checked Sep 25, 2026
What changed
- Sep 26, 2026: Added our ACEMAGIC K1 mini PC results (T0003, T0004).
- Sep 25, 2026: First published with our desktop results (T0001, T0002) and 19 published results.