A Benchmark for LLM's Understanding of Middle School and High School Science Topics
cs.AI, cs.CY, cs.HC
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/InviteInstitute/MS-NGSS-Aligned-Benchmark
Terminology
Sources
- LLMs in Education: Novel Perspectives, Challenges, and Opportunities
- LFM2 Technical Report
- The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models
- Large Databases Need Small, Open-Weight Language Models
- Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
- Measuring Massive Multitask Language Understanding
- Benchmarking the Pedagogical Knowledge of Large Language Models
- 2 OLMo 2 Furious
- TaMPERing with Large Language Models: A Field Guide for using Generative AI in Public Administration Research
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners
- A Survey on Data Synthesis and Augmentation for Large Language Models
- PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
- ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
- Large Language Models for Education: A Survey
- PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models
- Efficient Large Language Models: A Survey
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection