Dev.to
7/17/2026

AI Audit Case Study: Detecting Systematic Data Exclusion in Model Evaluation Pipelines
Original: Stratagems #16: Mark Left a Hole in His AI Audit. Lena Counted Every Layer.
Short summary
Mark audits FairPay's AI model evaluation pipeline and discovers systematic exclusion of low-confidence samples via an auto-labeling config file — the same signature as Pulse AI's previous failures. He documents the evidence across four layers but buries the key findings in an appendix, leaving it to the client to discover the truth. The story illustrates how automated pipeline errors can silently inflate benchmark metrics and how auditors navigate the tension between full disclosure and client politics.
- •Auditor finds auto-labeling config systematically excluding low-score samples from evaluation datasets
- •Same exclusion-rule signature as Pulse AI's prior issues — automated pipeline error, not deliberate fraud
- •Auditor strategically buries key findings in appendix, testing whether client follows the evidence trail
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
