arXiv cs.CL
7/2/2026

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
Short summary
Researchers benchmarked frontier LLMs on Arabic cultural and sociolinguistic knowledge using 103 expert-validated prompts graded by native speakers. GPT-5.4 emerged as the most reliable automated judge, while all models showed implicit cultural reasoning as their primary failure mode. Models notably outperformed on Egyptian Arabic vs. Iraqi Arabic, raising questions about actual capability gaps versus human grader variation.
- •Cross-evaluation framework with 103 human-validated prompts across Egyptian and Iraqi Arabic dialects
- •GPT-5.4 most reliable judge; four of five LLM judges show systematic leniency
- •Implicit cultural reasoning identified as primary failure mode for all models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?