Skip to content
AIPCs

Memory bandwidth predicts local AI speed: 23 chips, one model

By AIPCs Editor Updated Tested by us

Quick answer

How fast a computer writes with a local AI model is set mostly by its memory bandwidth. Across 23 processors and graphics chips running the same 7-billion-parameter model, most reached 55% to 85% of the speed their bandwidth allows; our RTX 4060 (272 GB/s) wrote 65.5 tokens per second, our Ryzen 9 7900X alone (102 GB/s) 15.5, and a $369 mini PC with one memory stick (21 GB/s) 3.9. A good estimate for any machine: tokens per second is about 0.7 times the bandwidth in GB/s, divided by the model size in GB.

Every time a local AI model writes a word, it reads its weights from memory. For a 7-billion-parameter model compressed to 4 bits, that is about 3.75 GB of data per token. So a simple rule should hold: the faster the memory, the faster the answer. We tested that rule on four chips in our own two machines and checked it against published results for 19 more.

The rule, and why it works

Divide a machine’s memory bandwidth by the amount of data the model reads per token and you get the fastest it could possibly write. Our RTX 4060 has 272 GB/s of memory bandwidth, so with Llama 2 7B (3.75 GB read per token) the limit is 272 / 3.75 = about 72 tokens per second. We measured 65.5, or 90% of the limit.

The same arithmetic works for a processor with no graphics card. Our Ryzen 9 7900X has dual-channel DDR5-6400 memory, a theoretical 102.4 GB/s, which allows about 27 tokens per second. We measured 15.5, or 57% of the limit. Desktop processors reach a smaller share because they cannot keep their memory as busy as a graphics card can.

Llama 2 7B on our desktop, with and without the graphics card
Model Where it ran Reading the prompttokens/s Writing the answertokens/s Test
Llama 2 7B Q4_0 NVIDIA GeForce RTX 4060 8 GB Whole model on GPU 2,716 65.5 T0002
Llama 2 7B Q4_0 AMD Ryzen 9 7900X (CPU only) CPU only 134 15.5 T0001
Llama 2 7B on our desktop, with and without the graphics card. Measured with llama.cpp's llama-bench (build 11191 (4b1a27fa0)): prompt speed over 512 tokens, answer speed over 128 tokens, average of 5 runs. Higher is better.

23 chips, one model

We chose Llama 2 7B at 4 bits (Q4_0) for our reference test because it is the model the llama.cpp community has used to benchmark hardware for years. That means we can place our results next to published ones measured the same way, from an M1 MacBook to an RTX 5090.

Faster memory, faster answers Writing speed for the same model (Llama 2 7B, 4-bit) on 23 machines, against memory bandwidth. Both axes are logarithmic.
  • Our tests
  • Apple Silicon (published)
  • PC graphics and AI chips (published)
20 50 100 200 500 1000 2000 2 5 10 20 50 100 200 500 Memory bandwidth, GB/s Tokens per second Theoretical limit 70% of limit Apple M1 (8-core GPU): 14.15 tokens/s at 68 GB/s (78% of the theoretical limit) Apple M1 Apple M2 (10-core GPU): 21.91 tokens/s at 100 GB/s (82% of the theoretical limit) Apple M4 (10-core GPU): 24.11 tokens/s at 120 GB/s (75% of the theoretical limit) Apple M3 Pro (18-core GPU): 30.74 tokens/s at 150 GB/s (77% of the theoretical limit) Apple M2 Pro (19-core GPU): 38.86 tokens/s at 200 GB/s (73% of the theoretical limit) Apple M4 Pro (20-core GPU): 50.74 tokens/s at 273 GB/s (70% of the theoretical limit) Apple M5 Pro (20-core GPU): 66.33 tokens/s at 307 GB/s (81% of the theoretical limit) Apple M3 Max (40-core GPU): 66.31 tokens/s at 400 GB/s (62% of the theoretical limit) Apple M4 Max (40-core GPU): 83.06 tokens/s at 546 GB/s (57% of the theoretical limit) Apple M5 Max (40-core GPU): 119.92 tokens/s at 614 GB/s (73% of the theoretical limit) Apple M5 Max Apple M2 Ultra (76-core GPU): 94.27 tokens/s at 800 GB/s (44% of the theoretical limit) Apple M3 Ultra (80-core GPU): 92.14 tokens/s at 800 GB/s (43% of the theoretical limit) AMD Ryzen AI Max+ 395 (Radeon 8060S), ROCm: 47.97 tokens/s at 256 GB/s (70% of the theoretical limit) Ryzen AI Max+ 395 NVIDIA GB10 (DGX Spark), Vulkan: 58.51 tokens/s at 273 GB/s (80% of the theoretical limit) DGX Spark Intel Arc A770 16 GB, Vulkan: 45.22 tokens/s at 560 GB/s (30% of the theoretical limit) Arc A770 NVIDIA RTX 3090, Vulkan: 164.05 tokens/s at 936 GB/s (66% of the theoretical limit) AMD RX 7900 XTX, Vulkan: 182.63 tokens/s at 960 GB/s (71% of the theoretical limit) NVIDIA RTX 4090, Vulkan: 187.97 tokens/s at 1008 GB/s (70% of the theoretical limit) NVIDIA RTX 5090, Vulkan: 263.63 tokens/s at 1792 GB/s (55% of the theoretical limit) RTX 5090 AMD Ryzen 9 7900X (CPU only) (our test T0001): 15.5 tokens/s at 102.4 GB/s (57% of the theoretical limit) Ryzen 9 7900X, CPU only (T0001) NVIDIA GeForce RTX 4060 8 GB (our test T0002): 65.5 tokens/s at 272 GB/s (90% of the theoretical limit) RTX 4060 (T0002) AMD Ryzen 3 4300U (CPU only) (our test T0003): 3.94 tokens/s at 21.3 GB/s (69% of the theoretical limit) K1, CPU only (T0003) AMD Radeon Graphics, integrated (Ryzen 3 4300U) (our test T0004): 4.82 tokens/s at 21.3 GB/s (85% of the theoretical limit) K1, integrated graphics (T0004)

Hover a point for its value. Every point is listed in the table below.

The points hug a straight line: across an 80-fold range of bandwidth, writing speed rises almost in step with it. Most reach between 55% and 85% of their theoretical limit.

Llama 2 7B Q4_0 writing speed by machine
Machine Bandwidth (GB/s) Writing speed (tokens/s) Share of limit Source
AMD Ryzen 3 4300U (CPU only) 21.3 3.9 69% Our test T0003
AMD Radeon Graphics, integrated (Ryzen 3 4300U) 21.3 4.8 85% Our test T0004
Apple M1 (8-core GPU) 68 14.2 78% llama.cpp #4167
Apple M2 (10-core GPU) 100 21.9 82% llama.cpp #4167
AMD Ryzen 9 7900X (CPU only) 102.4 15.5 57% Our test T0001
Apple M4 (10-core GPU) 120 24.1 75% llama.cpp #4167
Apple M3 Pro (18-core GPU) 150 30.7 77% llama.cpp #4167
Apple M2 Pro (19-core GPU) 200 38.9 73% llama.cpp #4167
AMD Ryzen AI Max+ 395 (Radeon 8060S), ROCm 256 48.0 70% llama.cpp #15021
NVIDIA GeForce RTX 4060 8 GB 272 65.5 90% Our test T0002
Apple M4 Pro (20-core GPU) 273 50.7 70% llama.cpp #4167
NVIDIA GB10 (DGX Spark), Vulkan 273 58.5 80% llama.cpp #10879
Apple M5 Pro (20-core GPU) 307 66.3 81% llama.cpp #4167
Apple M3 Max (40-core GPU) 400 66.3 62% llama.cpp #4167
Apple M4 Max (40-core GPU) 546 83.1 57% llama.cpp #4167
Intel Arc A770 16 GB, Vulkan 560 45.2 30% llama.cpp #10879
Apple M5 Max (40-core GPU) 614 119.9 73% llama.cpp #4167
Apple M2 Ultra (76-core GPU) 800 94.3 44% llama.cpp #4167
Apple M3 Ultra (80-core GPU) 800 92.1 43% llama.cpp #4167
NVIDIA RTX 3090, Vulkan 936 164.1 66% llama.cpp #10879
AMD RX 7900 XTX, Vulkan 960 182.6 71% llama.cpp #10879
NVIDIA RTX 4090, Vulkan 1008 188.0 70% llama.cpp #10879
NVIDIA RTX 5090, Vulkan 1792 263.6 55% llama.cpp #10879

What the exceptions tell you

  • The very fastest chips fall further behind. An RTX 5090 (1,792 GB/s) writes at 264 tokens per second, about 55% of its limit, and Apple’s Ultra chips reach under 50%. With this small model, they finish reading the weights so quickly that computation and software overheads start to dominate. With larger models they get closer to their limit.
  • Software matters. The Intel Arc A770 has twice the bandwidth of our RTX 4060 but wrote slower in its published result (45 versus 65.5 tokens per second). Drivers and how well a chip’s backend is optimized can cost more than half the potential speed.
  • Unified-memory machines land in the middle. AMD’s Ryzen AI Max+ 395 (256 GB/s) wrote at 48 tokens per second and NVIDIA’s DGX Spark (273 GB/s) at 58.5, 70% and 80% of their limits. What they offer is not speed but capacity: up to 128 GB, enough for models no consumer graphics card can hold.

The rule of thumb

For buying decisions, this is accurate enough:

Tokens per second ≈ 0.7 × memory bandwidth (GB/s) ÷ model size (GB)

Use about 0.85 for a modern NVIDIA graphics card and about 0.55 for a desktop processor on its own. Slow memory is easier to keep busy: our budget mini PC reached about 70% on its processor and 85% on its integrated graphics. Two refinements make it work for current models:

  1. For mixture-of-experts models, use the active size, not the file size. gpt-oss 20B is a 12 GB file, but only about 2.6 GB is read per token, which is why our Ryzen 9 7900X alone writes it at 21.8 tokens per second, faster than the much smaller Qwen3.5 9B (10.4).
  2. Long conversations add reading. The conversation memory (KV cache) is read for every token too. On our RTX 4060, Llama 2 7B slowed from 65.5 to 43.3 tokens per second with 4,096 tokens of conversation, because that older model stores a lot per token. Newer models store far less and barely slow down.

Our Can it run? calculator applies all of this for you, using each model’s real file size, active size and conversation memory, read directly from the model files.

What this means when you buy

  • Memory capacity decides which models you can run; bandwidth decides how fast. Check both.
  • For models up to about 9 billion parameters, an 8 GB graphics card like the RTX 4060 is already fast (45 to 68 tokens per second in our tests).
  • For bigger models, you are choosing between a graphics card with more video memory (fast, but typically 8 to 32 GB on consumer cards) and a unified-memory machine with 64 to 128 GB (slower, but it fits much larger models).
  • Spec sheets that lead with NPU “TOPS” are measuring something else. For local language models, look for the memory bandwidth figure.

The slow end: a $369 mini PC

Our ACEMAGIC K1 has a 4-core Ryzen 3 4300U and ships with a single 16 GB stick of DDR4-2666. With one stick, memory runs on one channel, so it has only 21.3 GB/s, the lowest bandwidth in this study. It lands where the rule predicts: 3.9 tokens per second on its processor (69% of its limit) and 4.8 on its integrated graphics (85%).

Llama 2 7B on the ACEMAGIC K1, processor and integrated graphics
Model Where it ran Reading the prompttokens/s Writing the answertokens/s Test
Llama 2 7B Q4_0 AMD Ryzen 3 4300U (CPU only) CPU only 20.5 3.94 T0003
Llama 2 7B Q4_0 AMD Radeon Graphics, integrated (Ryzen 3 4300U) Whole model on GPU 38.9 4.82 T0004
Llama 2 7B on the ACEMAGIC K1, processor and integrated graphics. Measured with llama.cpp's llama-bench (build 11191 (4b1a27fa0)): prompt speed over 512 tokens, answer speed over 128 tokens, average of 5 runs. Higher is better.

The lesson for budget machines: check that memory runs in dual channel. The K1 has a second, empty slot, and a matching stick would double its bandwidth to 42.7 GB/s.

Questions people ask

Does the NPU matter for running language models locally?

Not much today. Writing speed is limited by how fast memory can be read, not by how many operations the chip can do, and most local AI apps run models on the graphics processor or CPU rather than the NPU. The NPU matters for Windows Copilot+ features.

Why is reading the prompt so much faster than writing the answer?

When a model reads your prompt it can process hundreds of tokens in one pass over its weights, so the chip's compute power is the limit. When it writes, it produces one token at a time and must read all its active weights for each one, so memory bandwidth is the limit. On our RTX 4060, Llama 2 7B reads at about 2,700 tokens per second but writes at 65.

How do I find my computer's memory bandwidth?

Graphics cards list it on the maker's spec sheet. For system memory, multiply the number of memory channels by 8 bytes and by the memory speed: dual-channel DDR5-6400 is 2 x 8 x 6400 = 102.4 GB/s. Apple lists unified memory bandwidth in each Mac's specifications; for other machines, check the chip maker's spec page.

Why do the fastest chips fall further below the limit?

At very high bandwidth, reading the weights stops being the only bottleneck and the chip's compute, software overheads and synchronization start to matter. That is why an RTX 5090 reaches about 55% of its theoretical speed on this small model while an RTX 4060 reaches 90%.

Sources

  1. llama.cpp discussion #4167: Performance of llama.cpp on Apple Silicon M-series, checked Sep 25, 2026
  2. llama.cpp discussion #10879: Performance of llama.cpp with Vulkan, checked Sep 25, 2026
  3. llama.cpp discussion #15021: Performance of llama.cpp on AMD ROCm (HIP), checked Sep 25, 2026

What changed

  • Sep 26, 2026: Added our ACEMAGIC K1 mini PC results (T0003, T0004).
  • Sep 25, 2026: First published with our desktop results (T0001, T0002) and 19 published results.