Back to feed
The Verge
The Verge
6/25/2026
How to train your data | The Vergecast

How to train your data | The Vergecast

Short summary

Alex Reisner investigates how AI companies acquire and process training data from the open web, academia, and online platforms. The episode covers Common Crawl indexing, content filtering, and fair compensation debates. Core insight: training data acquisition remains the AI industry's least transparent and most ethically contested challenge.

  • AI companies source training data from Common Crawl, academic papers, and platforms like YouTube
  • Transparency around data sources is deliberately low; companies cite IP and legal risks
  • Creator compensation and fair data use remain unresolved industry questions

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more