Alignment Forum
6/22/2026

LLM-Driven Feature Discovery
Short summary
LLM-Driven Feature Discovery uses language models to analyze chat transcripts and uncover interesting model behaviors through unsupervised feature clustering, without needing access to model internals. Unlike Sparse Autoencoders, this approach produces human-interpretable clusters at the conversation level rather than per-token activations. Analysis of 100k Gemini transcripts yielded 20k features revealing behavior patterns, though the method remains preliminary and computationally expensive to scale.
- •Black-box feature discovery: LLMs identify behaviors from transcripts without model internals
- •Compared to SAEs: clearer interpretations, higher-level features, no activation access needed
- •Preliminary results on 100k Gemini transcripts show meaningful behavior clusters
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



