
From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings
Short summary
This paper proposes a framework for predicting psychometric item parameters (difficulty, discrimination, guessing) from text embeddings using regularized regression, with reliability and design ceilings to contextualize results. Applied to a math item bank (EEDI) and medical licensure benchmark (BEA 2024), item difficulty is highly predictable from text (R²=0.53, ~57% of reliability ceiling), while discrimination and guessing parameters are less predictable due to low target reliability rather than weak text signal. The authors demonstrate that single train-test splits inflate R² by 0.1–0.15, arguing for repeated cross-validation in calibration and benchmarking.
- •Framework predicts psychometric item parameters from text embeddings with reliability and design ceilings
- •Item difficulty is predictable from text (R²=0.53); discrimination and guessing less so due to low reliability ceilings
- •Single train-test splits inflate R² by 0.1–0.15; repeated cross-validation is essential for calibration work
Generated with AI, which can make mistakes.
Is this a good recommendation for you?