Back to feed
Alignment Forum
Alignment Forum
6/22/2026
LLM-Driven Feature Discovery

LLM-Driven Feature Discovery

Short summary

LLM-Driven Feature Discovery uses language models to analyze chat transcripts and uncover interesting model behaviors through unsupervised feature clustering, without needing access to model internals. Unlike Sparse Autoencoders, this approach produces human-interpretable clusters at the conversation level rather than per-token activations. Analysis of 100k Gemini transcripts yielded 20k features revealing behavior patterns, though the method remains preliminary and computationally expensive to scale.

  • Black-box feature discovery: LLMs identify behaviors from transcripts without model internals
  • Compared to SAEs: clearer interpretations, higher-level features, no activation access needed
  • Preliminary results on 100k Gemini transcripts show meaningful behavior clusters

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more