Back to feed
arXiv cs.CL
arXiv cs.CL
7/15/2026
TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation

TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation

Short summary

TAKE introduces a text dataset distillation framework that compresses corpora to as little as 0.1% of original size while preserving downstream task accuracy. It uses influence functions convolved along the training trajectory to score each sample's knowledge contribution, then applies discrete Optimal Transport to select prototypes from synthetic candidates. Evaluated on text classification and NLI tasks at extreme compression (20 samples/class), the method shows data efficiency is achievable without sacrificing fidelity.

  • Reduces text corpora to 0.1% while preserving downstream accuracy
  • Uses trajectory-aware influence functions to score sample importance
  • Code released at github.com/votrinhan88/take

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more