Back to feed
arXiv cs.CL
arXiv cs.CL
7/9/2026
Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation

Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation

Short summary

This paper identifies behavior leverage imbalance in multi-teacher on-policy distillation for agentic LLMs, where vanilla distillation improves tool-call recall but causes over-calling on examples that should be answered directly. The authors propose Soft Clamp, a per-token divergence calibration method that compresses extreme token-level Jensen-Shannon divergence at mode-entry and structural positions while preserving gradients. On APIGen-MT, Soft Clamp reduces over-calling from 13.7% to 9.0% relative to vanilla GKD while matching decision accuracy, and also lowers tool-call loops in BFCL multi-turn diagnostics.

  • Multi-teacher on-policy distillation can cause invisible behavior shifts toward tool over-calling in agentic LLMs
  • Soft Clamp calibrates per-token divergence at mode-entry positions, reducing over-calling from 13.7% to 9.0% while maintaining accuracy
  • Results show multi-teacher OPD should monitor where teacher signals act locally, not just aggregate loss magnitude

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more