arXiv cs.CL
7/15/2026

TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation
Short summary
TAKE introduces a text dataset distillation framework that compresses corpora to as little as 0.1% of original size while preserving downstream task accuracy. It uses influence functions convolved along the training trajectory to score each sample's knowledge contribution, then applies discrete Optimal Transport to select prototypes from synthetic candidates. Evaluated on text classification and NLI tasks at extreme compression (20 samples/class), the method shows data efficiency is achievable without sacrificing fidelity.
- •Reduces text corpora to 0.1% while preserving downstream accuracy
- •Uses trajectory-aware influence functions to score sample importance
- •Code released at github.com/votrinhan88/take
Generated with AI, which can make mistakes.
Is this a good recommendation for you?