Dev.to
6/29/2026

Your AI agent's leak risk depends more on the model than the prompt
Short summary
An empirical security study found that model choice is the dominant factor in system prompt leakage risk, with disclosure rates ranging 0–96% across five models—far outweighing prompt hardening efforts. The author released open-source scanning tool (agentproof-scan) and argues teams must measure their specific deployed model rather than relying on compliance audits or prompt engineering alone.
- •Model selection drives prompt leakage risk far more than prompt engineering; tested models showed 0–96% disclosure rates with identical probes
- •Practical testing against your specific deployed model is critical; compliance audits and tight prompts are insufficient guardrails
- •Open-source tool (agentproof-scan) now available; author emphasizes reproducible measurement over theoretical hardening and governance accountability
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



