Best Local LLMs for the NVIDIA L40S

Autor: Jakub Rusinowski · Ostatnia aktualizacja: 6 września 2026

Best all-round pick: GLM-5.1 72B

The NVIDIA L40S has 48 GB of VRAM, of which about 48 GB is available to a model. The largest model it can hold is Qwen 2.5 VL 72B Instruct (73B, 45.1 GB at Q4_K_M).

Best models overall on the NVIDIA L40S

ModelScoreMemory at Q4_K_MContextLicence
GLM-5.1 72B94.744.3 GB125KMIT
Qwen 3.5 72B94.444.3 GB125KApache 2.0
Gemma 4 27B ⭐9417.1 GB125KGemma License (commercial OK)
Llama 3.3 70B Instruct93.843.1 GB128KLlama Community
Qwen 3.7 35B-A3B93.221.9 GB256KApache-2.0
Qwen3.8 27B92.717.6 GB256KApache-2.0

Best model by what you are doing

WorkloadRecommended on this GPU
CodingQwen 3.7 35B-A3B (96.9)
Qwen 3.5 72B (96.6)
Qwen 3.6 35B-A3B (96.2)
General assistantGLM-5.1 72B (94.7)
Qwen 3.5 72B (94.4)
Gemma 4 27B ⭐ (94)
ReasoningGLM-5.1 72B (100)
Gemma 4 27B ⭐ (98.1)
GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (97.7)
RAGGemma 4 31B (95.6)
Qwen 3.6 27B (94.5)
Qwen 3.7 35B-A3B (93.8)
AgentsTernary Bonsai 27B (91.6)
Nemotron-Cascade 2 30B-A3B (91)
Devstral Small 2 24B (89.9)
VisionQwen 3.5 72B (99.7)
Qwen 3.7 35B-A3B (96.5)
Qwen 3.6 35B-A3B (95.8)

How these numbers are calculated

FAQ

What is the best LLM for the NVIDIA L40S?

GLM-5.1 72B is the strongest all-round pick that fits its 48 GB.

What is the largest model the NVIDIA L40S can run?

Qwen 2.5 VL 72B Instruct — 73B parameters, needing 45.1 GB at Q4_K_M.

How much of the NVIDIA L40S's memory can a model actually use?

About 48 GB of its 48 GB, before the desktop and runtime overhead are accounted for.

Compatibility Checks for the NVIDIA L40S

Similar GPUs

By Workload

More

← All GPUs | NVIDIA L40S specs