What "Tested" means
A product or result marked Tested by us was measured on hardware we had on our own desk,
with the procedure below. Each result carries a test ID (for example T0001) that links to the
benchmark database, where you can see the machine, the software build, the model file and
every setting.
What "Researched" means
We cannot test everything. A product marked Researched, not tested is one we recommend based on the manufacturer's specifications and published measurements from sources we trust, listed at the bottom of the page. We say so plainly, and we never describe hands-on experience we don't have.
The software
We measure with llama.cpp, the open-source engine that
LM Studio and many other local-AI apps build on. Its built-in llama-bench tool reports two speeds:
- Reading the prompt (prompt processing): how fast the model takes in 512 tokens of input. This decides how long you wait before an answer starts, especially with long documents.
- Writing the answer (token generation): how fast it produces 128 tokens of output. This is the speed you watch. About 10 tokens per second reads comfortably; below 4 feels slow.
Each test runs 5 times after a warm-up; we publish the average and the variation. We also repeat tests with 4,096 tokens already in the conversation, because long chats slow some models much more than others. We pin the llama.cpp build for a round of tests (currently build 11191) so results stay comparable.
The models
We use the same model files on every machine, verified by checksum:
- Llama 2 7B (Q4_0, 3.83 GB): The reference model most published llama.cpp benchmarks use, so our numbers compare directly with other hardware.
- Gemma 4 E4B (Q4_0, 4.59 GB): A popular small model from Google's Gemma 4 family.
- Qwen3.5 9B (Q4_K_M, 5.68 GB): A popular mid-size general model; about the largest dense model an 8 GB GPU holds with room for context.
- gpt-oss 20B (MXFP4, 12.11 GB): OpenAI's open-weight mixture-of-experts model: 21B parameters, about 3.6B active per token, widely benchmarked.
Our test machines
Reference desktop
- Processor: AMD Ryzen 9 7900X (12 cores / 24 threads, Zen 4)
- Memory: 32 GB DDR5-6400 (2 x 16 GB, dual channel) (theoretical peak 102.4 GB/s)
- Graphics: NVIDIA GeForce RTX 4060 8 GB (8 GB, 272 GB/s)
- Operating system: Windows 11 Home
The RTX 4060 runs on a PCIe 3.0 x1 link in this machine (the card supports PCIe 4.0 x8). That does not change results for models that fit entirely in its 8 GB, but it slows runs that split a model between GPU and CPU. We label those results.
ACEMAGIC K1
- Processor: AMD Ryzen 3 4300U (4 cores / 4 threads, Zen 2)
- Memory: 16 GB DDR4-2666 (one SO-DIMM as shipped, so single channel; the second slot is free) (theoretical peak 21.3 GB/s)
- Operating system: Windows 11 Pro
The integrated Radeon graphics (5 compute units) share system memory, so they run at the same 21.3 GB/s and have no separate VRAM figure. Our tests on it ran over a Remote Desktop session.
How we run a test
- Close other programs and confirm nothing else is using the processor or graphics card.
- Run our open-source script,
aipcs-bench, which records the hardware, drivers and software build, then runs the model ladder. - Import the raw output unchanged. The numbers on every page are generated from those files at build time, so a typo cannot creep in between the test and the page.
How our estimates work
The Can it run? calculator and our recommendations for machines we haven't tested use a simple, published model:
- Memory needed = the model file + the conversation memory (KV cache) for your context length + about 0.8 GB for the runtime. We read each model's layer layout from the file itself, because newer models store far less conversation memory per token than older ones.
- Writing speed is limited by memory bandwidth: each new token reads the model's active weights once. We multiply the theoretical speed by the share our own tests achieve (currently 90% on a graphics card and 55% on a desktop processor) and show a range. Machines with very fast memory keep a smaller share of it busy, as published results for Apple Ultra chips and the RTX 5090 show, so above about 300 GB/s (unified memory) or 1,000 GB/s (graphics cards) we scale that share down. On graphics cards we also add about 1 ms of fixed software time per token, which only matters for very fast cards.
- Other brands: AMD and Intel graphics cards currently get less speed from the same bandwidth than NVIDIA cards in local AI software. We apply a correction per brand, separately for dense and mixture-of-experts models, based on published measurements that each product page cites.
How accurate the estimates are
We check the estimate against every measurement we can find for our reference model (Llama 2 7B, 4-bit). Our own RTX 4060 and Ryzen 9 7900X results fall inside the estimated range, as do published results for the RTX 3090, RX 7900 XTX, RTX 4090, RTX 5090, Ryzen AI Max+ 395, Apple M4 Pro and M2 Ultra: 9 of 12 machines. The three misses (NVIDIA DGX Spark, Apple M4 Max and M5 Max) were all faster than we predicted, by 9% to 33%, so our figures for them err on the cautious side. All six of our desktop measurements with newer models also fall inside the estimated range. Our ACEMAGIC K1 mini PC, with slow single-channel memory (21.3 GB/s), ran 15% to 36% faster than estimated on both its processor and its integrated graphics: slow memory is easier to keep busy. It also ran gpt-oss 20B on its processor with a short prompt, which our memory rule counts as not fitting in 16 GB. So for slow machines our figures are cautious; we will adjust the formula once we have tested a second one.
Limits of our setup
- Two machines, one unit each. Individual units vary a little.
- Windows only so far.
- The RTX 4060 in our desktop sits on a PCIe 3.0 x1 connection. That doesn't affect models that fit fully in its 8 GB, but it slows tests that split a model between the graphics card and the processor. We label those results.
- We don't yet measure power draw at the wall or noise. Those will be added.
Corrections
Found a mistake? Email hello@aipcs.store. We fix errors quickly and note every change in the page's "What changed" section.