Best Local LLMs for the NVIDIA DGX Spark

作者: Jakub Rusinowski · 最后更新: 2026年7月21日

Best all-round pick: GPT-oss 120B

The NVIDIA DGX Spark has 128 GB of VRAM, of which about 128 GB is available to a model. The largest model it can hold is Qwen3.8-Flash-Next (180B, 109.5 GB at Q4_K_M).

Best models overall on the NVIDIA DGX Spark

ModelScoreMemory at Q4_K_MContextLicence
GPT-oss 120B95.871.3 GB128KApache-2.0
Qwen 3.5 122B-A10B (MoE)94.274.5 GB125KApache 2.0
Qwen 3.5 122B-A10B92.574.5 GB128KApache-2.0
GLM-5.1 72B91.244.3 GB125KMIT
Command R+ (104B)91.263.6 GB125KCC-BY-NC
Llama 4 Scout 17B91.266.6 GB10MLlama Community

Best model by what you are doing

WorkloadRecommended on this GPU
CodingDevstral-2 123B (99)
Qwen 3.5 122B-A10B (MoE) (96.5)
Qwen3-Coder 80B-A3B (MoE) (96.2)
General assistantGPT-oss 120B (95.8)
Qwen 3.5 122B-A10B (MoE) (94.2)
Qwen 3.5 122B-A10B (92.5)
ReasoningGLM-5.1 72B (98.4)
Qwen 3.5 122B-A10B (MoE) (97)
Gemma 4 27B ⭐ (95.9)
RAGLlama 4 Scout 17B (93.6)
Qwen3.8-Flash-Next (93)
Gemma 4 31B (92.8)
AgentsDevstral-2 123B (95.2)
Qwen3.8-Flash-Next (93.9)
Ternary Bonsai 27B (89)
VisionQwen 3.5 72B (96.9)
Llama 3.2 90B Vision Instruct (95.2)
Qwen3.8-Flash-Next (93.7)

How these numbers are calculated

FAQ

What is the best LLM for the NVIDIA DGX Spark?

GPT-oss 120B is the strongest all-round pick that fits its 128 GB.

What is the largest model the NVIDIA DGX Spark can run?

Qwen3.8-Flash-Next — 180B parameters, needing 109.5 GB at Q4_K_M.

How much of the NVIDIA DGX Spark's memory can a model actually use?

About 128 GB of its 128 GB, before the desktop and runtime overhead are accounted for.

Compatibility Checks for the NVIDIA DGX Spark

By Workload

More

← All GPUs | NVIDIA DGX Spark specs