arXiv cs.CL
7/10/2026

Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration
Short summary
Researchers introduce EspanStereo, a Spanish-language dataset for evaluating stereotypes in large language models across multiple Spanish-speaking regions, built through human-LLM collaborative annotation. The framework addresses the lack of non-English stereotype benchmarks and reveals significant regional variation in LLM bias. The methodology scales to other languages, offering a pathway toward comprehensive multilingual and cross-cultural bias evaluation.
- •New dataset (EspanStereo) enables stereotype evaluation for Spanish LLMs across multiple regions
- •Human-LLM collaborative framework reduces costs and identifies region-specific biases
- •Scalable methodology applicable to other languages and cultural contexts
Generated with AI, which can make mistakes.
Is this a good recommendation for you?