Dev.to
7/14/2026

We Audited an E-Commerce Support Bot, Fixed the Bugs, Then Re-Tested It. The Score Jumped 86 to 91.
Short summary
The authors audited an e-commerce support bot using BotCritic across four customer personas, scoring 86/100. After applying targeted system prompt fixes addressing false order lookups, missing return conditions, ambiguous refund triggers, and proactive eligibility warnings, the score jumped to 91/100. The biggest improvement was in robustness, showing that precise prompt engineering can resolve edge-case handling without changing underlying infrastructure.
- •E-commerce support bot audited with BotCritic scored 86/100 across 4 customer personas
- •Targeted system prompt fixes addressed false order lookups, missing return conditions, and ambiguous refund triggers
- •Re-test score rose to 91/100 with biggest gains in robustness, proving prompt-level fixes work
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



