
Latent Communication Between Language Model Agents: Channels, Alignment, and the Limits of Text
Short summary
This study investigates whether LLM agents lose information when communicating via text versus latent channels, using Sparse Autoencoder (SAE) feature analysis. The SAE-sparse channel retains 99.4% probe accuracy at 28-fold compression versus 80.4% for text, and Procrustes alignment achieves 92% top-1 retrieval between Llama and Mistral. However, text round-trip analysis shows 88% of SAE features are destroyed and replaced. Critically, task-level evaluation finds the latent channel matches but never exceeds text on cross-lingual concept tasks, and text augmentation with latent features provides no benefit—leading to negative conclusions: lost features mostly encode surface form, not task-relevant semantics.
- •SAE-sparse latent channel retains 99.4% probe accuracy vs 80.4% for text, but text destroys 88% of SAE features
- •Cross-architecture alignment achieves 92% top-1 retrieval between Llama and Mistral via Procrustes alignment
- •Negative result: latent channel never exceeds text on task-level evaluation; lost features encode surface form not semantics
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


