Answer four quick questions — or browse curated builds by price — and get exactly what to buy for local LLMs: DIY part lists, plug-and-play mini-PCs, and laptops, each with the models it can actually run.
Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
2026 GPU/DRAM prices are volatile — always check the live listing.
Under $500
Used RTX 3060 12GB Starter — $495
Prices as of Jul 6, 2026 · 12 GB VRAM
What it runs: Cosmos 3 16B · Ministral 3 14B · DeepSeek R1 14B — ~44 tok/s
The cheapest way into local AI that doesn't dead-end: 12GB VRAM runs every 7–8B model fast and stretches to 13–14B at tighter quantization.
What it runs: Qwen 3.5 122B (10B active) · GPT-OSS 120B · Qwen3-Coder 80B/3B active — ~18 tok/s
The unified-memory sweet spot: Strix Halo's 256 GB/s over 128GB runs 70B Q4 at usable speeds and giant MoE models no consumer GPU can hold — 30–40% cheaper than a comparable Mac Studio.
The most capable AI laptop you can buy: 48GB unified at 546 GB/s runs 32B models fast and fits 70B Q4 — capabilities no PC laptop matches at any wattage.
What it runs: Nemotron Cascade 2 70B · Cosmos 3 64B · Qwen 3.6 35B (3B active) — ~17 tok/s
The no-compromise workstation: dual RTX 4090s give 48GB of the fastest VRAM money reasonably buys, for 70B inference at interactive speeds and serious fine-tuning.
A usable starter build with a used RTX 3060 12GB lands under $500 and runs 7–8B models. The sweet spot is $1,000–1,300 — an RTX 4060 Ti 16GB build or a used RTX 3090 24GB build — which covers 14–32B models. Around $2,400–2,900 (RTX 4090 or dual RTX 3090) you can run 70B-class models at Q4.
What's the cheapest way to run a 70B model locally?
Two paths: a dual used RTX 3090 build (~$2,900, 48GB pooled VRAM, fastest) or a 128GB unified-memory mini-PC like the Beelink GTR9 Pro (~$1,899) — cheaper and silent, but roughly a quarter of the token speed because of its 256 GB/s memory bandwidth.
Mini-PC vs building your own for local AI — which is better?
DIY wins on speed per dollar: discrete GPUs have 3–7× the memory bandwidth of unified-memory boxes. Mini-PCs win on capacity per dollar, noise, and convenience — a 128GB Strix Halo box fits 70B+ models that no consumer GPU can hold. Buy DIY for interactive speed, mini-PC for big-model capacity or a quiet office.
Is a used RTX 3090 still worth it in 2026?
Yes — at ~$650–750 used it remains the best VRAM per dollar. 24GB runs 32B models at Q4 fully in VRAM, and its 936 GB/s bandwidth is close to an RTX 4090. Buy from a seller with returns and stress-test the memory on arrival.
Are Macs good for running local LLMs?
Apple Silicon is excellent for quiet, always-on setups: unified memory lets a Mac mini M4 Pro or MacBook Pro M4 Max load models a same-price GPU can't fit. Token speed is lower than a discrete NVIDIA card, and the M4 Max now tops out at 96GB unified memory — enough for 70B Q4, not the largest MoE models.