Dev.to
8/4/2026

Self-Hosting AI Models on a Raspberry Pi 5: A Complete Guide to Free, Private, Local AI Inference
Short summary
A practical guide to running quantized LLMs on a Raspberry Pi 5 (8GB) using Ollama, eliminating API costs and keeping all data local. The author benchmarks five models (Qwen2.5-0.5B through Llama 3.1-8B) on tokens/sec and RAM usage, recommending Llama 3.2-3B as the sweet spot for daily use. Hardware requirements include NVMe SSD (not SD card), active cooling, and the 27W power supply.
- •Raspberry Pi 5 8GB can run quantized LLMs locally via Ollama with zero API costs
- •Llama 3.2-3B is the best balance of speed and quality at 12-15 tokens/sec
- •NVMe SSD and active cooling are essential; SD card storage is too slow and wears out
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



