Back to feed
arXiv cs.CL
arXiv cs.CL
7/16/2026
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

Short summary

Researchers systematically apply persona vectors—behavioral directions in activation space—to audit two open-weight LLMs across a 53-trait inventory. They find models default to helpful, task-oriented behavior with all agentic traits being natural, while traits like hyperbole, hallucination, and sycophancy are steerable-latent. Standard extraction fails on extreme traits like 'evil,' but vectors transferred from fine-tuned variants can recover them, with residual refusals appearing in chain-of-thought. Persona vectors are most useful as probes of behavioral organization rather than control mechanisms.

  • 53-trait persona vector audit reveals which LLM behaviors are natural, steerable, or intractable
  • Default model behavior is helpful and task-oriented; agentic traits are all naturally expressed
  • Steering produces largest gains on excluded traits like hyperbole, hallucination, and sycophancy

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more