作者: Jakub Rusinowski · 最后更新: 2026年9月11日
Liquid AI's on-device line, built around their hybrid LIV convolution architecture rather than a stack of attention blocks. The 8B-A1B is the flagship: a sparse MoE small enough for phones, laptops and robots, reasoning-only (it emits an explicit chain of thought before answering), and trained on 38T tokens with large-scale RL. Liquid claims tool-calling quality comparable to models up to 4x its active size.
| Licence | What it permits | Applies to |
|---|---|---|
LFM Open License v1.0 | Commercial use permitted Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms. | LFM2.5-8B-A1B |
| LFM2.5-8B-A1B | Min 6 GB VRAM · Q4_K_M · 131,072 ctx · ollama run lfm2.5:8b |
The cheapest GPU that runs LFM2.5 locally (min 6 GB VRAM) is the Intel Arc B570 (10 GB).
Install Ollama then run: ollama run lfm2.5:8b
Minimum VRAM: 6 GB. For best results use Q4_K_M quantization.
LFM2.5 needs about 6 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: LFM2.5-8B-A1B (6 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Yes — LFM2.5 runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
Q4_K_M is the best balance of quality and VRAM for LFM2.5 in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
Install Ollama, then run: ollama run lfm2.5:8b. This downloads LFM2.5 and starts a local, OpenAI-compatible endpoint — no internet connection is needed after the initial download.