Back to feed
arXiv cs.LG
arXiv cs.LG
8/4/2026
Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models

Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models

Short summary

This paper shows that aggregate ROC-AUC metrics hide real training-data leakage risks in language models. A blind bag-of-words baseline achieves AUC 0.97 on WikiMIA, yet sampling-based attacks verbatim-extract identifiers from 16.6% of Pile documents on Pythia-6.9B—a risk invisible to aggregate metrics. The authors release leakit, a black-box extraction-audit tool, and recommend per-document, domain-decomposed privacy reporting.

  • Blind baselines match sampling-based MIA on aggregate AUC, but per-document extraction reveals real leaks invisible to aggregate metrics
  • Identifier leakage grows with model capacity (5.6% to 16.6%) and is ~3x stronger in code than prose
  • Corpus deduplication shows no reduction in leakage; leakit tool released for black-box extraction audits

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more