作者: Jakub Rusinowski · 最后更新: 2026年7月21日
Ranked for autonomous agents and tool use: Multi-step autonomous loops where the model calls tools, reads results and decides what to do next.
Best overall: AMD Radeon RX 7900 XTX
24 GB VRAM at 960 GB/s. It runs 24 of the models that qualify for this workload; the strongest is Ternary Bonsai 27B at an estimated 38.3 tokens/sec.
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | AMD Radeon RX 7900 XTX | 24 GB | $999 | 24 | Ternary Bonsai 27B | ~38.3 tok/s |
| Best value | Intel Arc B570 | 10 GB | $219 | 7 | GLM-6 9B | ~42.7 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 7 | GLM-6 9B | ~42.7 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 30 | Devstral-2 123B | ~2.7 tok/s |
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | 74.7 | 24 GB | $999 | 24 | ~38.3 | 0.11 | $42 |
| NVIDIA GeForce RTX 4090 | 73.7 | 24 GB | $1,599 | 24 | ~40.1 | 0.09 | $67 |
| NVIDIA GeForce RTX 5070 Ti | 72.9 | 16 GB | $749 | 12 | ~39.9 | 0.13 | $62 |
| NVIDIA GeForce RTX 5090 | 72.7 | 32 GB | $1,999 | 24 | ~66.6 | 0.12 | $83 |
| NVIDIA GeForce RTX 3090 | 72.2 | 24 GB | $1,499 | 24 | ~37.4 | 0.11 | $62 |
| NVIDIA GeForce RTX 5080 | 70.8 | 16 GB | $999 | 12 | ~42.5 | 0.12 | $83 |
| NVIDIA RTX 6000 Ada Generation | 70.6 | 48 GB | $6,799 | 24 | ~38.3 | 0.13 | $283 |
| AMD Radeon RX 7900 XT | 70.5 | 20 GB | $899 | 17 | ~32.4 | 0.1 | $53 |
| AMD Radeon RX 9070 XT | 69 | 16 GB | $599 | 12 | ~32.5 | 0.15 | $50 |
| NVIDIA L40S | 67.2 | 48 GB | $7,499 | 24 | ~34.8 | 0.1 | $312 |
| AMD Radeon RX 9070 | 66.9 | 16 GB | $549 | 12 | ~29.3 | 0.13 | $46 |
| AMD Radeon RX 7800 XT | 66.8 | 16 GB | $499 | 12 | ~28.6 | 0.11 | $42 |
The AMD Radeon RX 7900 XTX — 24 GB of VRAM runs 24 qualifying models, the strongest being Bonsai 27B.
The Intel Arc B570 at $219, which runs 7 qualifying models.
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.