Dev.to
7/18/2026

Underpowered Tests Lie
Short summary
An engineer auditing an AI ad-testing agent discovered a 20x sample size error: Lehr's equation prescribed 1,111 users per arm when the exact formula at a 1.9% base rate required 22,278. A second audit found raw p<0.05 across three-arm tests produced a 14% false-positive rate instead of the assumed 5%. The core lesson: statistical rules of thumb are calibrated for someone else's base rate, and underpowered tests never signal uncertainty — they confidently deliver wrong answers.
- •Lehr's shortcut (16/effect²) underestimates sample size 20x at low base rates like 1.9% CTR
- •Three-arm tests at raw p<0.05 have ~14% false-positive rate, not 5% — Holm-Bonferroni fixes this
- •Underpowered tests never report uncertainty; they confidently declare wrong winners that get acted on
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



