arXiv cs.CL
7/9/2026

Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation
Short summary
This paper identifies behavior leverage imbalance in multi-teacher on-policy distillation for agentic LLMs, where vanilla distillation improves tool-call recall but causes over-calling on examples that should be answered directly. The authors propose Soft Clamp, a per-token divergence calibration method that compresses extreme token-level Jensen-Shannon divergence at mode-entry and structural positions while preserving gradients. On APIGen-MT, Soft Clamp reduces over-calling from 13.7% to 9.0% relative to vanilla GKD while matching decision accuracy, and also lowers tool-call loops in BFCL multi-turn diagnostics.
- •Multi-teacher on-policy distillation can cause invisible behavior shifts toward tool over-calling in agentic LLMs
- •Soft Clamp calibrates per-token divergence at mode-entry positions, reducing over-calling from 13.7% to 9.0% while maintaining accuracy
- •Results show multi-teacher OPD should monitor where teacher signals act locally, not just aggregate loss magnitude
Generated with AI, which can make mistakes.
Is this a good recommendation for you?