arXiv cs.CL
7/16/2026

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
Short summary
Researchers systematically apply persona vectors—behavioral directions in activation space—to audit two open-weight LLMs across a 53-trait inventory. They find models default to helpful, task-oriented behavior with all agentic traits being natural, while traits like hyperbole, hallucination, and sycophancy are steerable-latent. Standard extraction fails on extreme traits like 'evil,' but vectors transferred from fine-tuned variants can recover them, with residual refusals appearing in chain-of-thought. Persona vectors are most useful as probes of behavioral organization rather than control mechanisms.
- •53-trait persona vector audit reveals which LLM behaviors are natural, steerable, or intractable
- •Default model behavior is helpful and task-oriented; agentic traits are all naturally expressed
- •Steering produces largest gains on excluded traits like hyperbole, hallucination, and sycophancy
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

