Autor: Jakub Rusinowski · Ostatnia aktualizacja: 21 lipca 2026
Ranked for privacy-sensitive work: Work with data that must never leave your machine — legal, medical, financial or personal records.
Best overall: AMD Radeon RX 7900 XTX
24 GB VRAM at 960 GB/s. It runs 80 of the models that qualify for this workload; the strongest is Qwen 3.7 35B-A3B at an estimated 164.8 tokens/sec.
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | AMD Radeon RX 7900 XTX | 24 GB | $999 | 80 | Qwen 3.7 35B-A3B | ~164.8 tok/s |
| Best value | Intel Arc B580 | 12 GB | $249 | 55 | Qwen 3 14B | ~33 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 54 | Qwen 3 14B | ~27.8 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 91 | GPT-oss 120B | ~53.3 tok/s |
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | 80.4 | 24 GB | $999 | 80 | ~164.8 | 0.46 | $12 |
| NVIDIA GeForce RTX 3090 | 78.8 | 24 GB | $1,499 | 80 | ~162.2 | 0.46 | $19 |
| NVIDIA GeForce RTX 4090 | 78.2 | 24 GB | $1,599 | 80 | ~169.9 | 0.38 | $20 |
| NVIDIA GeForce RTX 5090 | 77.2 | 32 GB | $1,999 | 80 | ~233 | 0.41 | $25 |
| AMD Ryzen AI Max+ 395 | 69.5 | 96 GB | $1,999 | 91 | ~50.4 | 0.42 | $22 |
| NVIDIA DGX Spark | 67.1 | 128 GB | $4,699 | 91 | ~53.3 | 0.36 | $52 |
| Intel Arc B580 | 59.4 | 12 GB | $249 | 55 | ~33 | 0.17 | $5 |
| NVIDIA GeForce RTX 5060 | 57.4 | 8 GB | $299 | 45 | ~49.5 | 0.34 | $7 |
| NVIDIA GeForce RTX 5070 | 57.3 | 12 GB | $549 | 55 | ~46.9 | 0.19 | $10 |
| Intel Arc B570 | 56 | 10 GB | $219 | 54 | ~27.8 | 0.19 | $4 |
| NVIDIA GeForce RTX 3080 (10GB) | 54.9 | 10 GB | $699 | 54 | ~52.4 | 0.16 | $13 |
| NVIDIA GeForce RTX 5060 Ti 8GB | 53.3 | 8 GB | $379 | 45 | ~49.5 | 0.28 | $8 |
The AMD Radeon RX 7900 XTX — 24 GB of VRAM runs 80 qualifying models, the strongest being Qwen 3.7.
The Intel Arc B570 at $219, which runs 54 qualifying models.
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.