Back to feed
MarkTechPost
MarkTechPost
6/28/2026
Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference

Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference

Short summary

Liquid AI released LFM2.5-230M, a 230-parameter open-weight model optimized for on-device inference, achieving 213 tokens/second on Galaxy S25 Ultra and beating competitors like Qwen3.5-0.8B on instruction following. Supports llama.cpp, MLX, vLLM, SGLang, ONNX for flexible deployment.

  • 230-parameter model runs at 213 tok/s on Galaxy S25 Ultra, 42 tok/s on Raspberry Pi
  • Outperforms larger models (Qwen3.5-0.8B, Gemma 3 1B) on instruction following
  • Supports multiple inference frameworks for on-device and edge deployment

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more