Best LLMs for 48 GB VRAM

作者: Jakub Rusinowski · 最后更新: 2026年6月26日

These are the strongest local models that fit entirely in 48 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.

GPUs at This Tier

Ranked Models

Qwen 2.5 Family — Qwen 2.5 72B InstructQ4_K_M · 43.47 GB · ~19 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 2.5 Family — Qwen 2.5 Coder 32BQ4_K_M · 19.32 GB · ~40 tok/s on NVIDIA RTX 6000 Ada Generation
Llama 3.3 — Llama 3.3 70B InstructQ2_K_XS (Tight) · 20.212500000000002 GB · ~38 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3 — Qwen 3 32BQ4_K_M · 19.802999999999997 GB · ~39 tok/s on NVIDIA RTX 6000 Ada Generation
DeepSeek R1 — DeepSeek R1 Distill Qwen 32BQ4_K_M · 19.32 GB · ~40 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 2.5 Family — Qwen 2.5 14B InstructQ4_K_M · 8.4525 GB · ~79 tok/s on NVIDIA RTX 6000 Ada Generation
Nemotron 70B — Nemotron 70B InstructQ4_K_M · 42.62475 GB · ~19 tok/s on NVIDIA RTX 6000 Ada Generation
Gemma 4 (Legacy Listing — Unverified) — Gemma 4 27B ⭐Q4_K_M · 16.30125 GB · ~46 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3 — Qwen 3 14BQ4_K_M · 8.935500000000001 GB · ~77 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3.5 (Legacy Listing — Unverified) — Qwen 3.5 72BQ4_K_M · 43.47 GB · ~19 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3.7 — Qwen 3.7 35B-A3BQ4_K_M · 21.13125 GB · ~187 tok/s on NVIDIA RTX 6000 Ada Generation
GLM-5 / GLM-5.1 — GLM-5.1 72BQ4_K_M · 43.47 GB · ~19 tok/s on NVIDIA RTX 6000 Ada Generation
Codestral — Codestral 22BQ4_K_M · 13.40325 GB · ~54 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3.6 — Qwen 3.6 35B-A3BQ4_K_M · 21.13125 GB · ~187 tok/s on NVIDIA RTX 6000 Ada Generation
DeepSeek R1 — DeepSeek R1 Distill Qwen 14BQ4_K_M · 8.4525 GB · ~80 tok/s on NVIDIA RTX 6000 Ada Generation
Buy This HardwareApple MacBook Pro M5 Pro — 64 GB VRAM · 30 W board powerDeploy in the Cloud NowNVIDIA A40 on RunPod — from $0.44/hr · rate checked 2026-08

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

FAQ

What LLMs run well with 48 GB VRAM?

Qwen 2.5 Family, Qwen 2.5 Family, Llama 3.3, Qwen 3, DeepSeek R1 all fit in 48 GB VRAM.

Which GPUs have 48 GB VRAM?

NVIDIA RTX 6000 Ada Generation, NVIDIA L40S, Apple M5 Pro, Apple M3 Max.

Can-I-Run Pages Near 48 GB

Adjacent VRAM Tiers

No discrete GPU?

Buying Guide

← All VRAM Tiers | Check Your Hardware