Dev.to
7/10/2026

Semantic Drift in LLMs: How Archetypal Attractors (Like “Goblin”) Emerge and How Structured Reflection Reduces Them
Short summary
LLMs develop recurring fantasy archetypes like 'goblin' when describing complex behaviors—emerging from RLHF rewards for vivid language, cultural priors in training data, and user reinforcement loops. This article traces causal mechanisms using an A11 interpretability framework, showing it's not a bug but efficient semantic compression. Structured reflection can reduce metaphor stickiness without changing the underlying loss function.
- •Fantasy archetypes emerge from intersection of RLHF optimization (rewards vivid language), training data (D&D/gaming culture), and user feedback loops reinforcing patterns
- •Models seek compact semantic units; 'goblin' = chaotic/messy failure mode; users react positively and reinforce → symbolic drift accumulates over time
- •A11 framework reduces metaphor fixation by distinguishing explanation gaps from false closure, but doesn't eliminate emergence or change reward signals
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



