Back to feed
arXiv cs.CL
arXiv cs.CL
7/13/2026
Creativity, honesty and designed forgetting emerge in small hyperbolic language models

Creativity, honesty and designed forgetting emerge in small hyperbolic language models

Short summary

Researchers show that small language models (146M–3B parameters) on a hyperbolic substrate can detect companion-AI failure modes—sycophancy, dependence-fostering, and confabulated memories—better than frontier zero-shot judges (AUROC 0.804 vs 0.721). A 146M behavioural auditor achieves 90.7% binary-compliance accuracy where human raters fail to agree (Fleiss kappa 0.074). A memory OS implementing designed forgetting via exponential decay demonstrates emergent retrieval-gating behavior, offering a small-model path to trustworthy companion AI.

  • Small models (146M–3B) on hyperbolic substrate detect sycophancy and confabulation better than frontier judges
  • 146M auditor achieves 90.7% compliance accuracy where human raters disagree (kappa 0.074)
  • Memory OS with exponential forgetting M(t)=S*exp(-λt) shows emergent selective retrieval gating

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more