Autor: Jakub Rusinowski · Ostatnia aktualizacja: 15 lipca 2026
These are the strongest local models that fit entirely in 12 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.
| Qwen 2.5 Family — Qwen 2.5 14B Instruct | Q4_K_M · 8.4525 GB · ~26 tok/s on Intel Arc B580 |
| Qwen 3 — Qwen 3 14B | Q4_K_M · 8.935500000000001 GB · ~26 tok/s on Intel Arc B580 |
| DeepSeek R1 — DeepSeek R1 Distill Qwen 14B | Q4_K_M · 8.4525 GB · ~27 tok/s on Intel Arc B580 |
| Phi-4 Family — Phi-4 (14B) | Q4_K_M · 8.4525 GB · ~26 tok/s on Intel Arc B580 |
| Granite 3.0 — Granite 3.0 8B Instruct | Q4_K_M · 4.83 GB · ~41 tok/s on Intel Arc B580 |
| Qwen 3.5 (Legacy Listing — Unverified) — Qwen 3.5 14B | Q4_K_M · 8.4525 GB · ~27 tok/s on Intel Arc B580 |
| Bonsai 27B — Ternary Bonsai 27B | Ternary (1.58-bit, ~1.71 bpw) · 5.77125 GB · ~35 tok/s on Intel Arc B580 |
| Qwen 2.5 Family — Qwen 2.5 7B Instruct | Q4_K_M · 4.5885 GB · ~46 tok/s on Intel Arc B580 |
| Qwen 3 — Qwen 3 8B | Q4_K_M · 4.950749999999999 GB · ~41 tok/s on Intel Arc B580 |
| DeepSeek R1 — DeepSeek R1 Distill Llama 8B | Q4_K_M · 4.83 GB · ~42 tok/s on Intel Arc B580 |
| Qwen3-Coder — Qwen3-Coder 8B | Q4_K_M · 4.83 GB · ~42 tok/s on Intel Arc B580 |
| Mistral Family — Mistral NeMo 12B | Q4_K_M · 7.245 GB · ~30 tok/s on Intel Arc B580 |
| Qwen 3.5 (Legacy Listing — Unverified) — Qwen 3.5 14B | Q4_K_M · 8.4525 GB · ~27 tok/s on Intel Arc B580 |
| Gemma 4 (Legacy Listing — Unverified) — Gemma 4 12B | Q4_K_M · 7.245 GB · ~30 tok/s on Intel Arc B580 |
| Gemma 3 — Gemma 3 12B Instruct | Q4_K_M · 7.245 GB · ~28 tok/s on Intel Arc B580 |
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Qwen 2.5 Family, Qwen 3, DeepSeek R1, Phi-4 Family, Granite 3.0 all fit in 12 GB VRAM.
Intel Arc B580, NVIDIA GeForce RTX 3060 (12GB), NVIDIA GeForce RTX 5070, NVIDIA GeForce RTX 4070 Super.