Back to feed
arXiv cs.LG
arXiv cs.LG
7/13/2026
A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions

Short summary

This paper proposes a unified framework for understanding knowledge distillation in LLMs by decomposing output scores into interactions. The authors find that KD methods work by sparsifying interactions, and better methods achieve higher sparsity of complex interactions. They introduce a Complex Interaction Penalty (CIP) loss function that consistently improves performance across diverse KD methods on in-domain and out-of-distribution benchmarks.

  • Knowledge distillation in LLMs works by sparsifying interactions in student models
  • Performance variance across KD methods stems from handling of complex interactions
  • New CIP loss function improves diverse KD methods on in-domain and OOD benchmarks

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more