Dev.to
7/3/2026

Your AI Agent Is Leaking Data Right Now — And Every Tool Call Looks Safe
Short summary
SafetyDrift, an open-source tool based on recent research, detects multi-step attacks on AI agents that pass individual tool-call guardrails. It tracks data exposure, tool escalation, and reversibility across a session using Markov chain analysis, achieving 100% F1 on 200 synthetic traces with zero false positives. Ships as an MCP server for easy integration into Claude Code, Cursor, and other AI agent frameworks.
- •SafetyDrift detects sequence attacks on AI agents that individual guardrails miss
- •Uses Markov chain analysis to predict violation risk within 5 steps
- •Integrates as MCP server with 100% F1 accuracy on attack/benign classification
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



