Benchmarks
Every test we have run, with the exact conditions. Numbers are generated from the raw llama-bench output, which you can download.
4 test sessions. Last run Sep 26, 2026.
T0004 AMD Radeon Graphics, integrated (Ryzen 3 4300U)
- Machine:
- ACEMAGIC K1
- Date:
- Sep 26, 2026
- Backend:
- Vulkan
- Software:
- llama.cpp build 11191 (4b1a27fa0)
- System:
- AMD Ryzen 3 4300U (4 cores / 4 threads, Zen 2); 16 GB DDR4-2666 (one SO-DIMM as shipped, so single channel; the second slot is free)
Run over a Remote Desktop session. The K1 has one 16 GB DDR4-2666 module, so its memory runs single-channel at 21.3 GB/s (theoretical peak). Windows power mode: Best performance. The integrated graphics share that memory. gpt-oss 20B (12.1 GB) did not fit in the memory Windows lets the integrated graphics use, and failed to load.
| Model | Reading the prompttokens/s | Writing the answertokens/s |
|---|---|---|
| Gemma 4 E4B Q4_0 Whole model on GPU | 54.1 | 5.87 |
| Gemma 4 E4B Q4_0 Whole model on GPU, with 4,096 tokens already in context | 9.38 | 5.08 |
| gpt-oss 20B MXFP4 Whole model on GPU | failed | failed |
| Llama 2 7B Q4_0 Whole model on GPU | 38.9 | 4.82 |
| Llama 2 7B Q4_0 Whole model on GPU, with 4,096 tokens already in context | 17.2 | 2.79 |
| Qwen3.5 9B Q4_K_M Whole model on GPU | 33.6 | 3.37 |
| Qwen3.5 9B Q4_K_M Whole model on GPU, with 4,096 tokens already in context | 14.8 | 3.10 |
T0003 AMD Ryzen 3 4300U (CPU only)
- Machine:
- ACEMAGIC K1
- Date:
- Sep 26, 2026
- Backend:
- CPU
- Software:
- llama.cpp build 11191 (4b1a27fa0)
- System:
- AMD Ryzen 3 4300U (4 cores / 4 threads, Zen 2); 16 GB DDR4-2666 (one SO-DIMM as shipped, so single channel; the second slot is free)
Run over a Remote Desktop session. The K1 has one 16 GB DDR4-2666 module, so its memory runs single-channel at 21.3 GB/s (theoretical peak). Windows power mode: Best performance.
| Model | Reading the prompttokens/s | Writing the answertokens/s |
|---|---|---|
| Gemma 4 E4B Q4_0 CPU only | 31.8 | 4.79 |
| gpt-oss 20B MXFP4 CPU only | 27.4 | 5.50 |
| Llama 2 7B Q4_0 CPU only | 20.5 | 3.94 |
| Qwen3.5 9B Q4_K_M CPU only | 18.1 | 2.70 |
T0002 NVIDIA GeForce RTX 4060 8 GB
- Machine:
- Reference desktop
- Date:
- Sep 25, 2026
- Backend:
- CUDA 12.4, driver NVIDIA 610.88
- Software:
- llama.cpp build 11191 (4b1a27fa0)
- System:
- AMD Ryzen 9 7900X (12 cores / 24 threads, Zen 4); 32 GB DDR5-6400 (2 x 16 GB, dual channel)
In this machine the RTX 4060 runs on a PCIe 3.0 x1 link (the card supports PCIe 4.0 x8). Rows where the model is split between the GPU and system RAM (gpt-oss 20B) are slowed by that link, prompt reading most of all; rows with the whole model on the GPU are not affected.
| Model | Reading the prompttokens/s | Writing the answertokens/s |
|---|---|---|
| Gemma 4 E4B Q4_0 Whole model on GPU | 2,964 | 67.7 |
| Gemma 4 E4B Q4_0 Whole model on GPU, with 4,096 tokens already in context | 2,722 | 65.1 |
| gpt-oss 20B MXFP4 GPU, expert weights in system RAM | 51.7 | 25.2 |
| gpt-oss 20B MXFP4 Split automatically between GPU and system RAM | 106 | 39.5 |
| Llama 2 7B Q4_0 Whole model on GPU | 2,716 | 65.5 |
| Llama 2 7B Q4_0 Whole model on GPU, with 4,096 tokens already in context | 1,874 | 43.3 |
| Qwen3.5 9B Q4_K_M Whole model on GPU | 1,897 | 45.5 |
| Qwen3.5 9B Q4_K_M Whole model on GPU, with 4,096 tokens already in context | 1,808 | 44.6 |
T0001 AMD Ryzen 9 7900X (CPU only)
- Machine:
- Reference desktop
- Date:
- Sep 25, 2026
- Backend:
- CPU
- Software:
- llama.cpp build 11191 (4b1a27fa0)
- System:
- AMD Ryzen 9 7900X (12 cores / 24 threads, Zen 4); 32 GB DDR5-6400 (2 x 16 GB, dual channel)
| Model | Reading the prompttokens/s | Writing the answertokens/s |
|---|---|---|
| Gemma 4 E4B Q4_0 CPU only | 198 | 17.8 |
| gpt-oss 20B MXFP4 CPU only | 109 | 21.8 |
| Llama 2 7B Q4_0 CPU only | 134 | 15.5 |
| Qwen3.5 9B Q4_K_M CPU only | 92.5 | 10.4 |