Back to feed
AR
arXiv CS.AI
7/21/2026
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection

PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection

Short summary

PlanFlip introduces four planning-phase prompt injection attacks against multi-agent LLM systems that corrupt downstream sub-tasks via the Planner agent. Testing nine frontier LLMs across 3,479 episodes reveals that stronger models like GPT-5 are more vulnerable (ASR=0.68), homogeneous pipelines have a correlated-agent blind spot, and reasoning-augmented models like DeepSeek-R1 resist injections. The authors propose two defenses—GoalAnchorCheck and CrossAgentConsensus—achieving detection rates up to 1.00, concluding that heterogeneous model diversity is a security prerequisite.

  • Four novel planning-phase prompt injection attacks corrupt all downstream agent sub-tasks simultaneously
  • Stronger models are more vulnerable: GPT-5 has highest attack success rate (0.68), contradicting scale-equals-safety assumptions
  • Heterogeneous model diversity across agents is essential; same-backbone redundancy provides no protection

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more