Back to feed
arXiv cs.CL
arXiv cs.CL
8/3/2026
Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

Short summary

This study evaluates whether LLMs can predict item difficulty levels on large-scale Reading and Writing tests. Zero-shot GPT-4.1 achieved a QWK of 0.578, but was outperformed by ConvBERT (QWK=0.625). All LLMs struggled with hard items, and GPT-5.4 tended to underestimate difficulty, suggesting caution when using LLMs for targeted-difficulty item generation.

  • GPT-4.1 zero-shot predicts item difficulty with QWK=0.578, below ConvBERT's 0.625
  • LLMs struggle to label hard items; GPT-5.4 underestimates difficulty
  • Semantic embeddings alone are insufficient for difficulty prediction

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more