AR
arXiv CS.AI
7/21/2026

PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
Short summary
PlanFlip introduces four planning-phase prompt injection attacks against multi-agent LLM systems that corrupt downstream sub-tasks via the Planner agent. Testing nine frontier LLMs across 3,479 episodes reveals that stronger models like GPT-5 are more vulnerable (ASR=0.68), homogeneous pipelines have a correlated-agent blind spot, and reasoning-augmented models like DeepSeek-R1 resist injections. The authors propose two defenses—GoalAnchorCheck and CrossAgentConsensus—achieving detection rates up to 1.00, concluding that heterogeneous model diversity is a security prerequisite.
- •Four novel planning-phase prompt injection attacks corrupt all downstream agent sub-tasks simultaneously
- •Stronger models are more vulnerable: GPT-5 has highest attack success rate (0.68), contradicting scale-equals-safety assumptions
- •Heterogeneous model diversity across agents is essential; same-backbone redundancy provides no protection
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


