arXiv cs.LG
8/4/2026

Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models
Short summary
This paper shows that aggregate ROC-AUC metrics hide real training-data leakage risks in language models. A blind bag-of-words baseline achieves AUC 0.97 on WikiMIA, yet sampling-based attacks verbatim-extract identifiers from 16.6% of Pile documents on Pythia-6.9B—a risk invisible to aggregate metrics. The authors release leakit, a black-box extraction-audit tool, and recommend per-document, domain-decomposed privacy reporting.
- •Blind baselines match sampling-based MIA on aggregate AUC, but per-document extraction reveals real leaks invisible to aggregate metrics
- •Identifier leakage grows with model capacity (5.6% to 16.6%) and is ~3x stronger in code than prose
- •Corpus deduplication shows no reduction in leakage; leakit tool released for black-box extraction audits
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
