The Verge
6/25/2026

How to train your data | The Vergecast
Short summary
Alex Reisner investigates how AI companies acquire and process training data from the open web, academia, and online platforms. The episode covers Common Crawl indexing, content filtering, and fair compensation debates. Core insight: training data acquisition remains the AI industry's least transparent and most ethically contested challenge.
- •AI companies source training data from Common Crawl, academic papers, and platforms like YouTube
- •Transparency around data sources is deliberately low; companies cite IP and legal risks
- •Creator compensation and fair data use remain unresolved industry questions
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

