Skip to content
AIPCs

Can you run AI without a graphics card? We tested a desktop and a $369 mini PC

By AIPCs Editor Updated Tested by us

Quick answer

Yes, for small and mixture-of-experts models. Using only its processor and 32 GB of DDR5-6400, our Ryzen 9 7900X wrote at 10 to 22 tokens per second: comfortable with Gemma 4 E4B (17.8) and gpt-oss 20B (21.8), and just usable with Qwen3.5 9B (10.4). The weak spot is reading long prompts, at 90 to 200 tokens per second, so a 5,000-word document takes 30 seconds to over a minute to take in. On a $369 mini PC with slow, single-channel memory the same models wrote at only 3 to 6 tokens per second.

Plenty of people who want to try local AI already own a capable desktop without a gaming graphics card, or a mini PC, or a laptop with only integrated graphics. So we unplugged the question from the graphics card entirely and ran our model ladder on the processor alone: a Ryzen 9 7900X (12 cores) with 32 GB of dual-channel DDR5-6400.

The results

Processor only: Ryzen 9 7900X, 32 GB DDR5-6400, 12 threads
Model Reading the prompttokens/s Writing the answertokens/s
Gemma 4 E4B Q4_0 CPU only 198 17.8
gpt-oss 20B MXFP4 CPU only 109 21.8
Llama 2 7B Q4_0 CPU only 134 15.5
Qwen3.5 9B Q4_K_M CPU only 92.5 10.4
Processor only: Ryzen 9 7900X, 32 GB DDR5-6400, 12 threads. Test T0001 on our reference desktop, AMD Ryzen 9 7900X (CPU only). Measured with llama.cpp's llama-bench (build 11191 (4b1a27fa0)): prompt speed over 512 tokens, answer speed over 128 tokens, average of 5 runs. Higher is better.

About 10 tokens per second reads comfortably, so three of the four models are pleasant to use and the fourth is just about usable. None of this needed a graphics card or any special setup beyond llama.cpp.

The surprise: the biggest model was the fastest

gpt-oss 20B is the largest file we tested (12.1 GB) yet it wrote fastest, at 21.8 tokens per second. It is a mixture-of-experts model: for each token it uses only 4 of its 32 expert blocks, so it reads about 2.6 GB per token instead of the whole 12 GB. Qwen3.5 9B, a dense model, reads all of its 5.1 GB of working weights for every token, and was the slowest at 10.4.

On a processor, where memory is the bottleneck, mixture-of-experts models are the ones to pick. The same applies to mini PCs and laptops without a graphics card.

Why memory speed matters more than cores

Writing an answer means reading the model’s weights from memory once per token. Our processor’s memory can move at most 102.4 GB/s, and in practice llama.cpp achieved about 55% of that. Divide by the amount read per token and you get the speed. That is why adding cores barely helps writing speed, while faster memory helps almost in proportion. Our bandwidth study shows the same pattern on 23 processors and graphics chips.

Reading your prompt is different: the model can process many tokens in one pass, so the processor’s raw computing power becomes the limit. That is where a desktop processor falls far behind a graphics card.

The same model on the processor and on an 8 GB graphics card
Model Where it ran Reading the prompttokens/s Writing the answertokens/s Test
Qwen3.5 9B Q4_K_M AMD Ryzen 9 7900X (CPU only) CPU only 92.5 10.4 T0001
Qwen3.5 9B Q4_K_M NVIDIA GeForce RTX 4060 8 GB Whole model on GPU 1,897 45.5 T0002
The same model on the processor and on an 8 GB graphics card. Measured with llama.cpp's llama-bench (build 11191 (4b1a27fa0)): prompt speed over 512 tokens, answer speed over 128 tokens, average of 5 runs. Higher is better.

Reading the prompt was 20 times faster on the RTX 4060 (1,897 versus 92 tokens per second) and writing was 4.4 times faster.

What to use if you have no graphics card

  • Pick mixture-of-experts models, such as gpt-oss 20B or 30B-class “A3B” models, over dense models of similar file size.
  • Have enough RAM for the model plus 4 GB or so for your operating system and apps. See how much memory local AI needs.
  • Faster memory beats more cores. Dual-channel DDR5 is roughly twice as fast as dual-channel DDR4 for this work.
  • Keep prompts short, or be patient with long documents.

A $369 mini PC: one memory stick halves the speed

We also ran the same models on our ACEMAGIC K1, a budget mini PC with a 4-core Ryzen 3 4300U. It ships with a single 16 GB stick of DDR4-2666, so its memory works on one channel: 21.3 GB/s, about a fifth of our desktop’s.

Processor only: ACEMAGIC K1, Ryzen 3 4300U, one 16 GB DDR4-2666 stick, 4 threads
Model Reading the prompttokens/s Writing the answertokens/s
Gemma 4 E4B Q4_0 CPU only 31.8 4.79
gpt-oss 20B MXFP4 CPU only 27.4 5.50
Llama 2 7B Q4_0 CPU only 20.5 3.94
Qwen3.5 9B Q4_K_M CPU only 18.1 2.70
Processor only: ACEMAGIC K1, Ryzen 3 4300U, one 16 GB DDR4-2666 stick, 4 threads. Test T0003 on our acemagic k1, AMD Ryzen 3 4300U (CPU only). Measured with llama.cpp's llama-bench (build 11191 (4b1a27fa0)): prompt speed over 512 tokens, answer speed over 128 tokens, average of 5 runs. Higher is better.

Writing speed fell roughly in line with bandwidth: about 4 times slower than the desktop processor with 4.8 times less bandwidth. Llama 2 7B went from 15.5 to 3.9 tokens per second and Qwen3.5 9B from 10.4 to 2.7. gpt-oss 20B was again the fastest, at 5.5, even though its 12.1 GB file only just fits in 16 GB; it ran with our short test prompt, and our calculator, which leaves more room for Windows, counts it as not fitting.

Its built-in Radeon graphics share the same memory, and through llama.cpp’s Vulkan backend they read prompts almost twice as fast as the processor and wrote about a fifth faster (T0004). gpt-oss 20B did not fit in the memory Windows lets the graphics use. If your machine has integrated graphics, an app with a Vulkan option (such as LM Studio) is worth trying.

The K1 has two memory slots and ships with one filled. A second, matching stick would give it two channels, and by the rule above should roughly double its writing speed.

How we tested

llama.cpp build 11191 (CPU backend) on Windows 11, 12 threads (one per physical core), nothing else running. Each figure is the average of 5 llama-bench runs: prompt speed over 512 tokens and writing speed over 128 tokens. Details and raw output: test T0001. The K1 ran the same build with 4 threads (one per core) over a Remote Desktop session: T0003 and T0004.

Questions people ask

How much RAM do I need to run AI on a processor?

Enough for the model file, the conversation and your operating system. In practice 16 GB handles models up to about 12 billion parameters, 32 GB handles gpt-oss 20B and 30B-class mixture-of-experts models, and 120B-class models such as gpt-oss 120B (a 63 GB file) need 96 GB or more.

Do more CPU cores make it faster?

For writing answers, mostly no: speed is limited by how fast the memory can be read, not by the number of cores. For reading prompts, yes: that part is compute-bound, which is why a graphics card reads prompts 10 to 20 times faster.

Does faster RAM help?

Yes, almost in proportion. Writing speed tracks memory bandwidth, which is memory speed times the number of channels. Dual-channel DDR5-6400 (102.4 GB/s) should be roughly twice as fast as dual-channel DDR4-3200 (51.2 GB/s), and workstation processors with 4 or 8 memory channels are faster again. In our tests, a mini PC with a single stick of DDR4-2666 (21.3 GB/s) wrote about 4 times slower than our DDR5-6400 desktop.

What is the cheapest way to speed this up?

First check that your memory runs in dual channel: two memory sticks, not one. A single stick halves the bandwidth, and many budget mini PCs ship that way. After that, add a graphics card: even an 8 GB card wrote about 4 times faster than the processor alone in our tests, and read prompts 15 to 20 times faster.

What changed

  • Sep 26, 2026: Added our ACEMAGIC K1 mini PC results (T0003, T0004).
  • Sep 25, 2026: First published with test T0001.