arXiv cs.CL
7/29/2026

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
Short summary
CogArena is a procedurally generated 13-paradigm benchmark testing whether LLM cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, a common axis explains ~50% of variance but within-grouping advantages are small and scoring-sensitive. Targeted scaffolds show a slight matched-grouping advantage, but no contrast survives multiplicity correction and the frozen confirmation criterion fails, suggesting LLM cognitive scores do not yet form stable five-dimensional profiles.
- •13-paradigm benchmark evaluating cognitive ability structure across 55 LLMs
- •Common axis explains ~50% variance but stable five-dimensional profiles not established
- •No scaffold-specific contrast survives multiplicity correction across model families
Generated with AI, which can make mistakes.
Is this a good recommendation for you?