arXiv cs.CL
8/3/2026

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
Short summary
This study evaluates whether LLMs can predict item difficulty levels on large-scale Reading and Writing tests. Zero-shot GPT-4.1 achieved a QWK of 0.578, but was outperformed by ConvBERT (QWK=0.625). All LLMs struggled with hard items, and GPT-5.4 tended to underestimate difficulty, suggesting caution when using LLMs for targeted-difficulty item generation.
- •GPT-4.1 zero-shot predicts item difficulty with QWK=0.578, below ConvBERT's 0.625
- •LLMs struggle to label hard items; GPT-5.4 underestimates difficulty
- •Semantic embeddings alone are insufficient for difficulty prediction
Generated with AI, which can make mistakes.
Is this a good recommendation for you?