作者: Jakub Rusinowski · 最后更新: 2026年7月21日
Ranked for coding and software engineering: Writing, refactoring and debugging code in an editor or terminal, with the model reading real project files.
Best overall: AMD Radeon RX 7900 XTX
24 GB VRAM at 960 GB/s. It runs 87 of the models that qualify for this workload; the strongest is Qwen 3.7 35B-A3B at an estimated 164.8 tokens/sec.
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | AMD Radeon RX 7900 XTX | 24 GB | $999 | 87 | Qwen 3.7 35B-A3B | ~164.8 tok/s |
| Best value | Intel Arc B580 | 12 GB | $249 | 61 | Qwen 3 14B | ~33 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 60 | Qwen3-Coder 8B | ~47.2 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 98 | Devstral-2 123B | ~2.7 tok/s |
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | 80.8 | 24 GB | $999 | 87 | ~164.8 | 0.46 | $11 |
| NVIDIA GeForce RTX 3090 | 79.3 | 24 GB | $1,499 | 87 | ~162.2 | 0.46 | $17 |
| NVIDIA GeForce RTX 4090 | 78.6 | 24 GB | $1,599 | 87 | ~169.9 | 0.38 | $18 |
| NVIDIA GeForce RTX 5090 | 77.7 | 32 GB | $1,999 | 87 | ~233 | 0.41 | $23 |
| NVIDIA RTX 6000 Ada Generation | 77.1 | 48 GB | $6,799 | 89 | ~164.8 | 0.55 | $76 |
| NVIDIA L40S | 76.7 | 48 GB | $7,499 | 89 | ~154.1 | 0.44 | $84 |
| Intel Arc B580 | 67.4 | 12 GB | $249 | 61 | ~33 | 0.17 | $4 |
| NVIDIA GeForce RTX 5070 | 62.5 | 12 GB | $549 | 61 | ~46.9 | 0.19 | $9 |
| AMD Radeon RX 7900 XT | 61.9 | 20 GB | $899 | 78 | ~28.6 | 0.09 | $12 |
| NVIDIA GeForce RTX 4070 | 59.1 | 12 GB | $599 | 61 | ~36.2 | 0.18 | $10 |
| NVIDIA GeForce RTX 4070 Super | 58.8 | 12 GB | $599 | 61 | ~36.2 | 0.16 | $10 |
| NVIDIA GeForce RTX 3060 (12GB) | 57.2 | 12 GB | $329 | 61 | ~26.4 | 0.16 | $5 |
The AMD Radeon RX 7900 XTX — 24 GB of VRAM runs 87 qualifying models, the strongest being Qwen 3.7.
The Intel Arc B570 at $219, which runs 60 qualifying models.
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.