SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children
cs.CL
Submitted: 2026-08-04
Updated: 2026-08-04
Code: https://github.com/LEAP-LABKUS/SpecialEduBench
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Language is the target of most early intervention for autistic children.
Terminology
Abstract
Language is the target of most early intervention for autistic children. Because the goal and the method change from child to child, the work falls to a teacher who takes one child at a time and judges each scene as it unfolds. Artificial intelligence is now being brought to that work, yet the benchmarks that reach special education ask what a model knows rather than what it does in front of a child. Building one is not straightforward, since whether a response is good teaching depends on what the child has just done, so no answer key applies. The evidence that settles it is visual as much as verbal, since the length of a wait, a shift of gaze, and the child's uptake leave no trace in a transcript. We introduce SpecialEduBench, which measures pedagogical competence along knowledge, skill, and attitude, with 4,537 knowledge items and with 200 skill items and 68 attitude items built on recorded intervention, the attitude items crossing pressure with monitoring into 192 response cells. Seven special-education experts wrote, scored, and reviewed the items, and we revised the judge model's instruction against the reference scores they set. Across eight frontier vision-language models no axis is saturated, since the strongest still fails about a tenth of the honesty cells. The models converge where the knowledge is factual and separate where the task is situated, and the failures gather where pressure is applied. We intend the benchmark as an audit to run before deployment and as a starting point for models built for this domain.
Sources
- ASD-Chat: An Innovative Dialogue Intervention System for Children with Autism based on LLM and VB-MAPP
- Alignment faking in large language models
- TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
- Evaluating Gemini in an arena for learning
- OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models
- Benchmarking the Pedagogical Knowledge of Large Language Models
- TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
- MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering