Best Local LLMs for the NVIDIA GeForce RTX 5090

Autor: Jakub Rusinowski · Ostatnia aktualizacja: 6 września 2026

Best all-round pick: Gemma 4 27B ⭐

The NVIDIA GeForce RTX 5090 has 32 GB of VRAM, of which about 32 GB is available to a model. The largest model it can hold is Command R (35B) (35B, 21.9 GB at Q4_K_M).

Best models overall on the NVIDIA GeForce RTX 5090

ModelScoreMemory at Q4_K_MContextLicence
Gemma 4 27B ⭐96.517.1 GB125KGemma License (commercial OK)
Qwen3.8 27B95.217.6 GB256KApache-2.0
Qwen 3.7 35B-A3B95.221.9 GB256KApache-2.0
Qwen 3 32B95.120.6 GB125KApache 2.0
Gemma 4 31B94.819.5 GB250KApache-2.0
Qwen 3.6 35B-A3B94.621.9 GB256KApache-2.0

Best model by what you are doing

WorkloadRecommended on this GPU
CodingQwen 3.7 35B-A3B (98.5)
Qwen 3 32B (98)
Qwen 3.6 35B-A3B (97.8)
General assistantGemma 4 27B ⭐ (96.5)
Qwen3.8 27B (95.2)
Qwen 3.7 35B-A3B (95.2)
ReasoningGemma 4 27B ⭐ (99.6)
GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (99.2)
Qwen 3 32B (98.3)
RAGGemma 4 31B (97.5)
Qwen 3.6 27B (96.2)
Qwen3.8 27B (95.4)
AgentsTernary Bonsai 27B (93.3)
Nemotron-Cascade 2 30B-A3B (92.8)
Devstral Small 2 24B (91.4)
VisionQwen 3.7 35B-A3B (98.1)
Gemma 4 31B (97.8)
Qwen 3.6 35B-A3B (97.4)

What the NVIDIA GeForce RTX 5090 cannot run

These models are strong picks generally but exceed the 32 GB this card makes available.

ModelNeeds at Q4_K_MShort by
Llama 3.3 70B Instruct43.1 GB~11.1 GB

How these numbers are calculated

FAQ

What is the best LLM for the NVIDIA GeForce RTX 5090?

Gemma 4 27B ⭐ is the strongest all-round pick that fits its 32 GB.

What is the largest model the NVIDIA GeForce RTX 5090 can run?

Command R (35B) — 35B parameters, needing 21.9 GB at Q4_K_M.

How much of the NVIDIA GeForce RTX 5090's memory can a model actually use?

About 32 GB of its 32 GB, before the desktop and runtime overhead are accounted for.

Compatibility Checks for the NVIDIA GeForce RTX 5090

Similar GPUs

By Workload

More

← All GPUs | NVIDIA GeForce RTX 5090 specs