Autor: Jakub Rusinowski · Ostatnia aktualizacja: 6 września 2026
Best all-round pick: Mistral Small 3.1 24B
The NVIDIA GeForce RTX 4080 has 16 GB of VRAM, of which about 16 GB is available to a model. The largest model it can hold is Mistral Small 3 (24B) (24B, 15.3 GB at Q4_K_M).
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| Mistral Small 3.1 24B | 95.9 | 15 GB | 125K | Apache 2.0 |
| Qwen 3.5 14B | 94.1 | 9.3 GB | 125K | Apache 2.0 |
| Qwen 3.5 14B | 93.2 | 9.3 GB | 125K | Apache 2.0 |
| Qwen 3 14B | 92.9 | 9.7 GB | 125K | Apache 2.0 |
| Gemma 3 12B Instruct | 92.2 | 8 GB | 128K | Gemma |
| Mistral Small 3.2 24B | 92.2 | 15 GB | 128K | Apache-2.0 |
| Workload | Recommended on this GPU |
|---|---|
| Coding | Devstral Small 2 24B (96.5) Qwen 3 14B (95.9) DeepSeek R1 Distill Qwen 14B (93.9) |
| General assistant | Mistral Small 3.1 24B (95.9) Qwen 3.5 14B (94.1) Qwen 3.5 14B (93.2) |
| Reasoning | Qwen 3 14B (96) DeepSeek R1 Distill Qwen 14B (94.2) Qwen 3.5 14B (93.8) |
| RAG | Devstral Small 2 24B (88.4) Qwen 3 14B (84.4) Mistral Small 3.1 24B (84) |
| Agents | Devstral Small 2 24B (92.6) GLM-6 9B (84.7) Qwen 3.5 14B (84.4) |
| Vision | Qwen 3.5 14B (95.4) Mistral Small 3.1 24B (95.2) Magistral Small 24B (92.1) |
These models are strong picks generally but exceed the 16 GB this card makes available.
| Model | Needs at Q4_K_M | Short by |
|---|---|---|
| Gemma 4 27B ⭐ | 17.2 GB | ~1.2 GB |
| Qwen3.8 27B | 17.6 GB | ~1.6 GB |
| Qwen 3.7 35B-A3B | 22 GB | ~6 GB |
| Gemma 4 31B | 19.6 GB | ~3.6 GB |
Mistral Small 3.1 24B is the strongest all-round pick that fits its 16 GB.
Mistral Small 3 (24B) — 24B parameters, needing 15.3 GB at Q4_K_M.
About 16 GB of its 16 GB, before the desktop and runtime overhead are accounted for.