Written by Jakub Rusinowski · Last updated July 21, 2026
Ranked for writing and editing: Drafting, rewriting and editing prose where voice and readability matter as much as correctness.
Best overall: NVIDIA GeForce RTX 5070 Ti
16 GB VRAM at 896 GB/s. It runs 70 of the models that qualify for this workload; the strongest is Mistral Small 3.1 24B at an estimated 40.9 tokens/sec.
| GPU | VRAM | MSRP | Models that fit | Best model it runs | Est. speed | |
|---|---|---|---|---|---|---|
| Best overall | NVIDIA GeForce RTX 5070 Ti | 16 GB | $749 | 70 | Mistral Small 3.1 24B | ~40.9 tok/s |
| Best value | NVIDIA GeForce RTX 5060 | 8 GB | $299 | 51 | GLM-4.7 9B | ~49.5 tok/s |
| Budget pick | Intel Arc B570 | 10 GB | $219 | 60 | Gemma 3 12B Instruct | ~30.2 tok/s |
| Most memory | NVIDIA DGX Spark | 128 GB | $4,699 | 98 | Llama 3.3 70B Instruct | ~4.6 tok/s |
| GPU | Score | VRAM | MSRP | Models fit | Est. speed | Tok/s per watt | Cost per model |
|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5070 Ti | 69.4 | 16 GB | $749 | 70 | ~40.9 | 0.14 | $11 |
| NVIDIA GeForce RTX 5080 | 67.4 | 16 GB | $999 | 70 | ~43.5 | 0.12 | $14 |
| AMD Radeon RX 9070 XT | 65.8 | 16 GB | $599 | 70 | ~33.3 | 0.15 | $9 |
| AMD Radeon RX 9070 | 63.4 | 16 GB | $549 | 70 | ~30 | 0.14 | $8 |
| AMD Radeon RX 7800 XT | 63.2 | 16 GB | $499 | 70 | ~29.3 | 0.11 | $7 |
| NVIDIA GeForce RTX 4080 Super | 62.4 | 16 GB | $999 | 70 | ~34.2 | 0.11 | $14 |
| NVIDIA GeForce RTX 4070 Ti Super | 61.3 | 16 GB | $799 | 70 | ~31.4 | 0.11 | $11 |
| NVIDIA GeForce RTX 4080 | 60.8 | 16 GB | $1,199 | 70 | ~33.3 | 0.1 | $17 |
| NVIDIA GeForce RTX 5060 | 59.6 | 8 GB | $299 | 51 | ~49.5 | 0.34 | $6 |
| NVIDIA GeForce RTX 5060 Ti 16GB | 58.2 | 16 GB | $429 | 70 | ~21.4 | 0.12 | $6 |
| NVIDIA GeForce RTX 5090 | 57.1 | 32 GB | $1,999 | 87 | ~57.2 | 0.1 | $23 |
| AMD Radeon RX 9060 XT 8GB | 56 | 8 GB | $299 | 51 | ~36.5 | 0.24 | $6 |
The NVIDIA GeForce RTX 5070 Ti — 16 GB of VRAM runs 70 qualifying models, the strongest being Mistral Small 3.1.
The Intel Arc B570 at $219, which runs 60 qualifying models.
8 GB is the entry point at which a model for this workload will run at all. More memory buys a stronger model, not just a faster one.