arXiv cs.CL
8/3/2026

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models
Short summary
This paper shows that knowledge distillation in small instruction-tuned LLMs has asymmetric bias effects: it improves context-following on unambiguous tasks but degrades refusal calibration on ambiguous ones. The authors trace the calibration loss to insufficient refusal-shaped training data and show that aggregate bias metrics conceal per-item harm. They propose PCCD, a three-step protocol that catches both asymmetric bias and trivial-refuser failures missed by standard evaluations.
- •Distillation improves unambiguous-task accuracy but causes 15% of correctly-abstained ambiguous items to receive stereotype answers
- •Silence-loss and filled-silence effects are uncorrelated, arising from distinct mechanisms
- •Proposed PCCD protocol detects calibration failures that aggregate metrics like CrowS-Pairs and BBQ miss
Generated with AI, which can make mistakes.
Is this a good recommendation for you?