The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale
cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 31 pages, 8 figures, 16 tables. Companion paper: arXiv:2608.01000
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present CLAIM, a production pipeline for K-12 assessment-item generation coupling a two-stage generate-then-attack protocol (the model drafts as a "curriculum architect", then re-enters the same
Terminology
Abstract
We present CLAIM, a production pipeline for K-12 assessment-item generation coupling a two-stage generate-then-attack protocol (the model drafts as a "curriculum architect", then re-enters the same conversation as a hostile adversarial reviewer), bi-directional few-shot conditioning on accepted and rejected items, the latter carrying the evaluator's diagnosis, and a knowledge dictionary of 44,844 error-correction rules mined from that feedback and retrieved per standard and item type. Across 43,227 scored items over 755 Common Core ELA standards, three item types, and ten LLMs, the pipeline reaches a 97.8% expert-evaluator pass rate on a 9,074-item production run. We then ask what that rate certifies. Re-scoring a stratified sample with three judges from other vendors, blind to the deployed verdict, reproduces the format ordering under every judge and recovers a larger open-set deficit than the deployed evaluator does; but agreement on the accept/reject binary is weak at production prevalence (kappa about 0.13), and the judges agree with each other no better. The level is therefore judge-relative, and with no student-response data our quality evidence is evaluator-judged throughout. The corpus also exposes a robust asymmetry. Multiple-choice and multiple-select generation saturate at 98% or above for both frontier models under a dozen static rules, whereas fill-in-the-blank generation is capability-tiered (82.8-96.7% across five models under a matched rule set, standards, and judge) and plateaus under prompt-only optimization, with error mass shifting between answer-key over-inclusion and omission as rules accumulate. We analyze this as open-set boundary determination, a task autoregressive decoders are structurally ill-equipped to solve, and show the asymmetry recurring when the evaluator itself is distilled: fail-recall rises from 8% to 63% while F1 saturates at 0.25.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Distractor generation for multiple-choice questions with predictive prompting and large language models
- Answer Matching Outperforms Multiple Choice for Language Model Evaluation
- Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets
- Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study
- Sequential Enumeration in Large Language Models
- Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
- How Many Instructions Can LLMs Follow at Once?
- SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
- Language Models (Mostly) Know What They Know
- Self-critiquing models for assisting human evaluators
- Large Language Models for Education: A Survey and Outlook
- EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
- TextGrad: Automatic "Differentiation" via Text
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection