Autor: Jakub Rusinowski · Ostatnia aktualizacja: 12 lipca 2026
Best all-round pick: GPT-oss 120B
The Apple M2 Max has 96 GB of unified memory, of which about 72 GB is available to a model. The largest model it can hold is GPT-oss 120B (117B, 71.3 GB at Q4_K_M).
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| GPT-oss 120B | 96.4 | 71.3 GB | 128K | Apache-2.0 |
| GLM-5.1 72B | 94.7 | 44.3 GB | 125K | MIT |
| Qwen 3.5 72B | 94.4 | 44.3 GB | 125K | Apache 2.0 |
| Llama 3.3 70B Instruct | 93.8 | 43.1 GB | 128K | Llama Community |
| Cogito v1 70B | 92.6 | 43.1 GB | 128K | Apache-2.0 |
| Command R+ (104B) | 92.6 | 63.6 GB | 125K | CC-BY-NC |
| Workload | Recommended on this GPU |
|---|---|
| Coding | Qwen3-Coder 80B-A3B (MoE) (98.6) Qwen 3.5 72B (96.6) Llama 3.3 70B Instruct (96.1) |
| General assistant | GPT-oss 120B (96.4) GLM-5.1 72B (94.7) Qwen 3.5 72B (94.4) |
| Reasoning | GLM-5.1 72B (100) Gemma 4 27B ⭐ (97.1) Qwen 3.5 72B (96.8) |
| RAG | Llama 4 Scout 17B (94.4) Gemma 4 31B (94.3) Qwen 3.6 27B (93.3) |
| Agents | Ternary Bonsai 27B (90.4) GLM-5.1 72B (88.4) Nex-N2.5 mini (87.5) |
| Vision | Qwen 3.5 72B (99.7) Llama 3.2 90B Vision Instruct (97.1) Qwen 3.7 35B-A3B (94.8) |
GPT-oss 120B is the strongest all-round pick that fits its 72 GB.
GPT-oss 120B — 117B parameters, needing 71.3 GB at Q4_K_M.
About 72 GB of its 96 GB, because macOS reserves a share of unified memory for the system.