What are TOPS? AI PC numbers explained, and why they don't predict AI speed
Quick answer
TOPS (trillions of operations per second) measures how much math a chip's AI engine can do. On Windows it matters for one thing: Copilot+ features need an NPU rated above 40 TOPS. It does not predict how fast a local AI chatbot writes. That is set by memory bandwidth: NVIDIA's RTX 5090 has 2.5 times the AI TOPS of the RTX 4090 but writes only about 1.4 times faster in published tests, in line with its 1.8 times more bandwidth. TOPS matter more for reading long prompts.
AI PC spec sheets lead with one number: TOPS. A laptop “delivers 80 TOPS”; a mini PC offers “up to 126 TOPS”. It sounds like a measure of how good the computer is at AI. It is, but only for some kinds of AI work, and not for the one many buyers care about most: running a chatbot-style model on the computer itself.
What TOPS means
TOPS stands for trillions of operations per second: the peak amount of simple math a chip’s AI hardware can do. Three kinds of chips quote it:
| Chip | Example | AI TOPS |
|---|---|---|
| NPU, a small AI engine inside the processor | Qualcomm Snapdragon X2 Elite Extreme (X2E-94) | 80 |
| NPU | AMD Ryzen AI Max+ 395 | 50 (126 counting processor and graphics too) |
| Graphics card | NVIDIA RTX 4060 | 242 |
| Graphics card | NVIDIA RTX 5090 | 3,352 |
Two cautions when comparing them. Some makers quote the NPU alone and others add the processor and graphics, as AMD’s 126 “overall” TOPS does. And TOPS can be counted at different numeric precisions, so figures from different makers, or even different generations, don’t compare directly.
Where TOPS matter: Windows AI features
Microsoft’s Copilot+ PC label requires an NPU that can do more than 40 TOPS. That unlocks Windows features that run on the NPU, such as improved Windows search, and it lets apps written for Windows’ AI tools use the NPU. The NPU’s advantage is efficiency: it can run AI tasks for hours without draining a laptop battery. If you want these features, check for the Copilot+ label; beyond the 40 TOPS bar, a higher NPU number rarely changes what you can do.
Where TOPS don’t matter: how fast a local model writes
When an AI model writes an answer, it produces one word-piece (token) at a time, and for each one it reads all of its active weights from memory. The chip spends most of its time waiting for memory, not doing math. So writing speed is set by memory bandwidth, and extra TOPS go unused.
NVIDIA’s own cards show it. In published llama.cpp results for the same small model, the RTX 5090 has 2.5 times the AI TOPS of the RTX 4090 (3,352 versus 1,321) but writes only 1.4 times faster (264 versus 188 tokens per second), close to its 1.8 times more memory bandwidth. Our study of 23 chips shows the same pattern from a $369 mini PC to the RTX 5090: speed follows bandwidth.
The NPU usually doesn’t take part at all. Popular local AI apps such as LM Studio and Ollama mostly run models on the graphics chip or the processor, so a 50 or 80 TOPS NPU sits idle while they work.
Where TOPS do help: reading long prompts
Reading your prompt is different. The model processes hundreds of tokens in one pass over its weights, so the chip’s raw math power becomes the limit. That is where a graphics card’s TOPS pay off: our RTX 4060 read prompts about 20 times faster than our 12-core Ryzen 9 7900X processor on its own, while writing about 4 times faster (T0001, T0002). If you work with long documents, compute matters too.
What to look at instead
- For AI features in Windows: the Copilot+ label (an NPU above 40 TOPS). See what an AI PC is.
- For running AI models locally: memory size (which models fit) and memory bandwidth (how fast they write). See how much memory local AI needs, or check any machine in our Can it run? calculator.
- For long documents: a graphics card or a chip with strong graphics, for fast prompt reading.
Questions people ask
What does TOPS stand for?
Trillions (tera) of operations per second: how many simple math operations, such as multiplications, a chip's AI hardware can do each second at its peak. It is a theoretical maximum, like a car's top speed.
How many TOPS do I need?
For Windows AI features, an NPU rated above 40 TOPS, which is what Microsoft requires for a Copilot+ PC. For running AI models such as chatbots on the computer itself, look at memory size and memory bandwidth instead; TOPS barely affect how fast they write.
Is an 80 TOPS NPU better than a 50 TOPS one?
For work that runs on the NPU, it has more headroom, but both clear the Copilot+ bar and run the same Windows features. Neither changes how fast popular local AI apps run, because those apps mostly use the graphics chip or processor.
Why do graphics cards have so many more TOPS than NPUs?
A graphics card is a big, power-hungry chip built for parallel math: NVIDIA rates even the RTX 4060 at 242 AI TOPS, against about 40 to 85 for laptop NPUs. NPUs are small and efficient, designed to run AI features for hours on battery rather than to be as fast as possible.
Sources
- Microsoft Learn: Copilot+ PCs developer guide (40 TOPS NPU), checked Sep 26, 2026
- NVIDIA GeForce graphics card comparison (AI TOPS), checked Sep 26, 2026
- Qualcomm Snapdragon X2 Elite product page (NPU TOPS), checked Sep 26, 2026
- AMD Ryzen AI Max+ 395 specifications, checked Sep 25, 2026
- llama.cpp discussion #10879: Performance of llama.cpp with Vulkan, checked Sep 25, 2026
What changed
- Sep 26, 2026: First published.