Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques
Hotaka Maeda, Yikai Lu
cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: 19 pages, 4 figures, 3 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation.
Terminology
Abstract
Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy =.75, F1 =.64, AUC =.80, Sensitivity =.64, Specificity =.81). Prediction for math (F1 =.73, AUC =.86) was considerably more accurate than ELA (F1 =.51, AUC =.72). Lowering the decision threshold from.5 to.25 raised average sensitivity for ELA and math to.88 and.91, while reducing specificity to.31 and.56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.
Sources
- Two-Stage Reasoning-Infused Learning: Improving Classification with LLM-Generated Reasoning
- Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
- Reasoning and Sampling-Augmented MCQ Difficulty Prediction via LLMs
- Estimating Item Difficulty with Large Language Models as Experts
- Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Small Language Models: Survey, Measurements, and Insights
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Analyzing Examinee Comments using DistilBERT and Machine Learning to Ensure Quality Control in Exam Content
- Attention Is All You Need
- GPT-4o System Card
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection