LLM-as-a-judge validity is strongly task-dependent across physics assessment formats
physics.ed-ph, cs.CL
Submitted: 2026-03-16
Updated: 2026-09-26
Terminology
Sources
- Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study
- Using Large Language Models to Assign Partial Credit to Students' Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-facing Feedback
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
- LLM Evaluators Recognize and Favor Their Own Generations
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Report on the Scoping Workshop on AI in Science Education Research 2025
- AI-assisted Automated Short Answer Grading of Handwritten University Level Mathematics Exams