Small, Private Language Models as Teammates for Educational Assessment Design
cs.AI, cs.CL, cs.HC
Submitted: 2026-05-14
Updated: 2026-05-14
DOI: 10.1007/978-3-032-29755-6_15
Code: https://github.com/kastle-lab/bloom-taxonomy-lm-analysis
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical
Terminology
Abstract
Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bloom's taxonomy). However, they often rely on subjective or limited evaluation methods; focus primarily on proprietary models; or rarely systematically examine generation, evaluation, or deployment constraints in real educational settings. Meanwhile, Small Language Models (SLMs) have emerged as local alternatives that better address privacy and resource limitations; yet their effectiveness for assessment tasks remains underexplored. To address this gap, we systematically compare LLMs and SLMs for assessment question design; evaluate generation quality across Bloom's taxonomy levels using reproducible, pedagogically grounded metrics; and further assess model-based judging against expert-informed evaluation by analyzing reliability and agreement patterns. Results show that SLMs achieve competitive performance across key pedagogically motivated quality dimensions while enabling local, privacy-sensitive deployment. However, model-based evaluations also exhibit systematic inconsistencies and bias relative to expert ratings. These findings provide evidence to posit language models as bounded assistants in assessment workflows; underscore the necessity of Human-in-the-Loop; and advance the automated educational question generation field by examining quality, reliability, and deployment-aware trade-offs.
Sources
- Exploring the Capabilities of Prompted Large Language Models in Educational and Assessment Applications
- Small Models, Big Support: A Local LLM Framework for Educator-Centric Content Creation and Assessment with RAG and CAG
- It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
- Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
- GPT-4 as a Homework Tutor can Improve Student Engagement and Learning Outcomes
- From Cloud to Edge: Rethinking Generative AI for Low-Resource Design Challenges
- BERTScore: Evaluating Text Generation with BERT
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection