作者: Jakub Rusinowski · 最后更新: 2026年6月26日
These are the strongest local models that fit entirely in 32 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.
| Qwen 2.5 Family — Qwen 2.5 Coder 32B | Q4_K_M · 19.32 GB · ~3 tok/s on Apple M4 |
| Llama 3.3 — Llama 3.3 70B Instruct | Q2_K_XS (Tight) · 20.212500000000002 GB · ~3 tok/s on Apple M4 |
| Qwen 3 — Qwen 3 32B | Q4_K_M · 19.802999999999997 GB · ~3 tok/s on Apple M4 |
| DeepSeek R1 — DeepSeek R1 Distill Qwen 32B | Q4_K_M · 19.32 GB · ~3 tok/s on Apple M4 |
| Qwen 2.5 Family — Qwen 2.5 14B Instruct | Q4_K_M · 8.4525 GB · ~7 tok/s on Apple M4 |
| Gemma 4 (Legacy Listing — Unverified) — Gemma 4 27B ⭐ | Q4_K_M · 16.30125 GB · ~4 tok/s on Apple M4 |
| Qwen 3 — Qwen 3 14B | Q4_K_M · 8.935500000000001 GB · ~7 tok/s on Apple M4 |
| Qwen 3.7 — Qwen 3.7 35B-A3B | Q4_K_M · 21.13125 GB · ~22 tok/s on Apple M4 |
| Codestral — Codestral 22B | Q4_K_M · 13.40325 GB · ~5 tok/s on Apple M4 |
| Qwen 3.6 — Qwen 3.6 35B-A3B | Q4_K_M · 21.13125 GB · ~22 tok/s on Apple M4 |
| DeepSeek R1 — DeepSeek R1 Distill Qwen 14B | Q4_K_M · 8.4525 GB · ~7 tok/s on Apple M4 |
| Phi-4 Family — Phi-4 (14B) | Q4_K_M · 8.4525 GB · ~7 tok/s on Apple M4 |
| Gemma 4 — Gemma 4 31B | Q4_K_M · 18.71625 GB · ~3 tok/s on Apple M4 |
| Qwen 3.5 (Legacy Listing — Unverified) — Qwen 3.5 32B | Q4_K_M · 19.32 GB · ~3 tok/s on Apple M4 |
| Gemma 3 — Gemma 3 27B Instruct | Q4_K_M · 16.30125 GB · ~3 tok/s on Apple M4 |
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Qwen 2.5 Family, Llama 3.3, Qwen 3, DeepSeek R1, Qwen 2.5 Family all fit in 32 GB VRAM.
Apple M4, Apple M5, NVIDIA GeForce RTX 5090, Apple M2 Pro.