Back to feed
Dev.to
Dev.to
8/4/2026
LLMs on Consumer Hardware — Part 2: Prefill and the Failure of the AI PC

LLMs on Consumer Hardware — Part 2: Prefill and the Failure of the AI PC

Short summary

This article benchmarks LLM inference across three consumer machines, revealing that prefill speed varies 18x while generation varies less than 2x — making prefill the real bottleneck for large-prompt workloads. Model-load time is storage-bound, with SATA SSDs taking 50s versus 8s on NVMe. The author critiques 'AI PC' marketing, showing that NPUs rated at 16 TOPS are useless for multi-billion-parameter model inference, leaving a 13-minute stall on large system prompts.

  • Prefill speed varies 18x across machines while generation varies <2x; prefill dominates large-prompt workloads
  • Model load time is storage-bound: SATA SSD takes 50s vs 8s on NVMe for an 18GB model
  • AI PC NPUs (16 TOPS) are inert for LLM inference; CPU prefill rate is the deciding factor

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more