MarkTechPost
6/28/2026

Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference
Short summary
Liquid AI released LFM2.5-230M, a 230-parameter open-weight model optimized for on-device inference, achieving 213 tokens/second on Galaxy S25 Ultra and beating competitors like Qwen3.5-0.8B on instruction following. Supports llama.cpp, MLX, vLLM, SGLang, ONNX for flexible deployment.
- •230-parameter model runs at 213 tok/s on Galaxy S25 Ultra, 42 tok/s on Raspberry Pi
- •Outperforms larger models (Qwen3.5-0.8B, Gemma 3 1B) on instruction following
- •Supports multiple inference frameworks for on-device and edge deployment
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



