arXiv cs.CL
7/13/2026

Creativity, honesty and designed forgetting emerge in small hyperbolic language models
Short summary
Researchers show that small language models (146M–3B parameters) on a hyperbolic substrate can detect companion-AI failure modes—sycophancy, dependence-fostering, and confabulated memories—better than frontier zero-shot judges (AUROC 0.804 vs 0.721). A 146M behavioural auditor achieves 90.7% binary-compliance accuracy where human raters fail to agree (Fleiss kappa 0.074). A memory OS implementing designed forgetting via exponential decay demonstrates emergent retrieval-gating behavior, offering a small-model path to trustworthy companion AI.
- •Small models (146M–3B) on hyperbolic substrate detect sycophancy and confabulation better than frontier judges
- •146M auditor achieves 90.7% compliance accuracy where human raters disagree (kappa 0.074)
- •Memory OS with exponential forgetting M(t)=S*exp(-λt) shows emergent selective retrieval gating
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

