Best Local LLMs for Coding

作者: Jakub Rusinowski · 最后更新: 2026年9月11日

Writing, refactoring and debugging code in an editor or terminal, with the model reading real project files.

Top pick: Qwen3-Coder 8B

Scores 96.3/100 for coding and software engineering. 8B parameters, needing about 5.6 GB at Q4_K_M, 125K context, Apache-2.0.

Ranked for coding and software engineering

ModelScoreParamsContextLicenceQuality index
1. Qwen3-Coder 8B96.38B125KApache-2.0— (estimated)
2. Qwen 3.7 35B-A3B95.935B256KApache-2.0— (estimated)
3. Qwen 3 32B95.433B125KApache 2.0— (estimated)
4. Qwen 3.6 35B-A3B95.235B256KApache-2.0— (estimated)
5. Devstral Small 2 24B94.924B256KApache-2.0— (estimated)
6. DeepSeek R1 Distill Qwen 32B94.932B128KMIT87 (cited)

Best pick for your memory budget

The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.

MemoryTypical hardwareRecommended models
8 GBRTX 4060, RTX 3070, base MacBook AirQwen3-Coder 8B (96.3)
GLM-6 9B (93.2)
Qwen 3 8B (92.9)
12 GBRTX 3060 12 GB, RTX 5070Qwen 3 14B (95.9)
Qwen3-Coder 8B (94.9)
DeepSeek R1 Distill Qwen 14B (94.2)
16 GBRTX 5080, RTX 4080, RX 9070 XTDevstral Small 2 24B (96.5)
Qwen 3 14B (95.9)
DeepSeek R1 Distill Qwen 14B (93.9)
24 GBRTX 4090, RTX 3090, RX 7900 XTXQwen 3.7 35B-A3B (98.5)
Qwen 3 32B (98)
Qwen 3.6 35B-A3B (97.8)
48 GBRTX 6000 Ada, MacBook Pro M4 Max 48 GBQwen 3.7 35B-A3B (96.9)
Qwen 3.5 72B (96.6)
Qwen 3.6 35B-A3B (96.2)
128 GB+Mac Studio, DGX Spark, multi-GPUDevstral-2 123B (99)
Qwen 3.5 122B-A10B (MoE) (96.5)
Qwen3-Coder 80B-A3B (MoE) (96.2)

How this ranking works

Coding score carries 60% of the capability weight and reasoning the remaining 35%, because most real editor work is "understand this repo, then write correct code". Context is weighted heavily: below 16K tokens a model cannot hold a meaningful slice of a codebase, and 128K is treated as fully served. Latency matters (0.7) — a coding assistant slower than you type stops being used.

Worked example — Qwen3-Coder 8B: capability 88.4 × 0.419, quality 82.3 × 0.247, context 97.3 × 0.16, license 100 × 0.044, accessibility 100 × 0.129 + 6 tag bonus (coding, debugging, code-review).

Requirements applied: context floor 16,384 tokens (ideal 131,072), quality floor 55, licence weight 0.4, latency weight 0.7.

Running coding and software engineering locally

FAQ

What is the best local LLM for coding and software engineering?

Qwen3-Coder 8B, scoring 96.3/100 against this workload's published requirements. 124 models qualified.

What hardware do I need for coding and software engineering?

A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.

How were these models ranked?

Coding score carries 60% of the capability weight and reasoning the remaining 35%, because most real editor work is "understand this repo, then write correct code". Context is weighted heavily: below 16K tokens a model cannot hold a meaningful slice of a codebase, and 128K is treated as fully served. Latency matters (0.7) — a coding assistant slower than you type stops being used.

Hardware for This Workload

Related Workloads

Top Pick

Tools

← All workloads | Check your hardware