What can an 8 GB graphics card run? We tested the RTX 4060
Quick answer
An 8 GB graphics card like the RTX 4060 runs 4-bit models up to about 9 billion parameters entirely on the card, and fast: 45 to 68 tokens per second in our tests. Larger mixture-of-experts models still work if the computer has 32 GB of system RAM; gpt-oss 20B wrote at 39.6 tokens per second with part of it in RAM. Large dense models such as a 27B slow to a crawl once they spill over, so for those you need more video memory.
Eight gigabytes is one of the most common amounts of video memory on graphics cards sold in the last few years, from the RTX 3060 Ti and RTX 4060 to the RTX 5060. So we started our testing there: an RTX 4060 in our desktop, with four popular open models, each compressed to about 4 bits per weight.
Models that fit on the card: fast
A model runs fastest when all of it sits in video memory. At 4 bits, that means models up to about 9 billion parameters on an 8 GB card, with room left for a conversation.
| Model | Reading the prompttokens/s | Writing the answertokens/s |
|---|---|---|
| Gemma 4 E4B Q4_0 Whole model on GPU | 2,964 | 67.7 |
| Gemma 4 E4B Q4_0 Whole model on GPU, with 4,096 tokens already in context | 2,722 | 65.1 |
| Llama 2 7B Q4_0 Whole model on GPU | 2,716 | 65.5 |
| Llama 2 7B Q4_0 Whole model on GPU, with 4,096 tokens already in context | 1,874 | 43.3 |
| Qwen3.5 9B Q4_K_M Whole model on GPU | 1,897 | 45.5 |
| Qwen3.5 9B Q4_K_M Whole model on GPU, with 4,096 tokens already in context | 1,808 | 44.6 |
All three write far faster than you can read: comfortable reading is about 10 tokens per second, and the slowest here is 43. Reading your prompt is faster still, 1,800 to 3,000 tokens per second, so even a long document is taken in within a few seconds.
Long conversations: newer models barely slow down
Every token of conversation is stored in the card’s memory and re-read for each new word, so long chats can slow a model down. How much depends on how the model is built. With 4,096 tokens of conversation already in memory:
- Llama 2 7B slowed by a third, from 65.5 to 43.3 tokens per second. It is an older design that stores 512 KB for every token.
- Qwen3.5 9B barely changed (45.5 to 44.6). It keeps full conversation memory in only 8 of its 32 layers, about 32 KB per token.
- Gemma 4 E4B also held up (67.7 to 65.1). Most of its layers only look at the most recent 512 tokens.
That efficiency also decides how long a conversation fits. By our calculation, an 8 GB card leaves room for about 28,000 tokens with Qwen3.5 9B, but only about 12,800 with Llama 3.1 8B.
Bigger than 8 GB: borrow system RAM
gpt-oss 20B is a 12.1 GB file, so it cannot fit on the card. llama.cpp can split it, keeping part of the model in system RAM. Because gpt-oss is a mixture-of-experts model, only about 2.6 GB of its weights are used for each token, which makes the split far less painful than you might expect.
| Model | Where it ran | Reading the prompttokens/s | Writing the answertokens/s | Test |
|---|---|---|---|---|
| gpt-oss 20B MXFP4 | NVIDIA GeForce RTX 4060 8 GB GPU, expert weights in system RAM | 51.7 | 25.2 | T0002 |
| gpt-oss 20B MXFP4 | NVIDIA GeForce RTX 4060 8 GB Split automatically between GPU and system RAM | 106 | 39.5 | T0002 |
| gpt-oss 20B MXFP4 | AMD Ryzen 9 7900X (CPU only) CPU only | 109 | 21.8 | T0001 |
- Automatic split (llama.cpp’s fit option): 39.6 tokens per second. As much of the model as fits goes on the card; the rest runs from system RAM.
- All expert weights in RAM: 25.2 tokens per second. Simpler, but slower.
- Processor only, for comparison: 21.8 tokens per second.
What about 27B and 35B models?
We have not tested these on the RTX 4060 yet, but our calculator, calibrated on the results above, gives a clear picture for this machine (8 GB card plus 32 GB of DDR5-6400):
- Qwen3.6 35B-A3B (mixture of experts, 22.1 GB): an estimated 32 to 44 tokens per second, split with system RAM. Very usable.
- Qwen3 Coder 30B-A3B (mixture of experts, 18.6 GB): an estimated 27 to 38 tokens per second.
- Qwen3.8 27B (dense, 16.5 GB): an estimated 3 to 4 tokens per second. Too slow for chat, because every token has to read about 10 GB from system RAM.
The lesson: on an 8 GB card, mixture-of-experts models are the way to run bigger AI. For large dense models you need a card with 16 to 24 GB of video memory, or a unified-memory computer.
Should you buy an 8 GB card for AI?
If you mainly want fast answers from 4B to 9B models, or you want to run mixture-of-experts models with plenty of system RAM, an 8 GB card does the job well. If you want 14B to 35B models to run entirely on the graphics card, 16 GB is the sensible minimum and 24 to 32 GB gives real headroom. Our graphics card guide compares the options at 2026 prices, and the calculator shows what fits on each size before you buy.
How we tested
We used llama.cpp build 11191 (CUDA 12.4) on Windows 11 with NVIDIA driver 610.88 and nothing else running on the card. Each figure is the average of 5 runs of llama-bench: prompt speed over 512 tokens, writing speed over 128 tokens. Full details and raw output are on our benchmarks page and methodology.
Questions people ask
Can an RTX 4060 run gpt-oss 20B?
Yes, if the computer has enough system RAM. The model is a 12.1 GB file, too big for 8 GB of video memory, but llama.cpp can keep part of it in system RAM. On our RTX 4060 with 32 GB of DDR5-6400 it wrote at 39.6 tokens per second (test T0002).
How long a conversation fits on an 8 GB card?
It depends on the model, because models store very different amounts of conversation memory per token. By our calculation, Qwen3.5 9B leaves room for about 28,000 tokens, Gemma 4 E4B for over 100,000, and Llama 3.1 8B for about 12,800, at standard 16-bit conversation memory.
Is an 8 GB graphics card worth buying for local AI in 2026?
For small models it is fast and inexpensive. If you want 14B to 35B models to run entirely on the card, a 16 GB or 24 GB card is the better buy; our calculator shows what each size can hold.
Does the PCIe slot matter?
Not for models that fit entirely on the card, which only use the slot while loading. It matters when a model is split between the card and system RAM. Our test card runs on a slow PCIe 3.0 x1 link, which particularly slowed prompt reading in our split tests; in a normal x8 or x16 slot those numbers would be higher.
What changed
- Sep 26, 2026: Updated calculator estimates after refining the model with published results for fast graphics cards.
- Sep 25, 2026: First published with tests T0001 and T0002.