A Unified Framework to Elicit Structured Feedback for Interpretable Multi-Trait Essay Scoring

arXiv:2608.28407 · cs.CL · Submitted 2026-08-28 · Read on arXiv

cs.CL

Submitted: 2026-08-28

Updated: 2026-08-28

Comments: 14 pages, accepted to EMNLP 2026 Findings. Code: https://github.com/Atiyahsama/HiFTS

Code: https://github.com/Atiyahsama/HiFTS

License: http://creativecommons.org/licenses/by/4.0/

The gist: Multi-trait Automated Essay Scoring (AES) requires rubric-grounded reasoning across interdependent traits, rather than isolated score prediction.

Terminology

Abstract

Multi-trait Automated Essay Scoring (AES) requires rubric-grounded reasoning across interdependent traits, rather than isolated score prediction. Existing feedback-enhanced methods often decouple feedback from scoring or assess traits independently, weakening score--feedback consistency and rubric alignment. We propose HiFTS, a unified autoregressive framework that generates hierarchical CoT feedback before predicting trait-level and holistic scores. HiFTS distills rubric-grounded hierarchical CoT feedback from a teacher LLM and trains student models to jointly generate feedback and scores. HiFTS further applies Group Relative Policy Optimization with a composite reward balancing score agreement, calibration, feedback quality, and structural validity. At inference, a lightweight global prior provides holistic guidance to reduce drift during long-form reasoning. We also introduce CFMS-34, a Chinese multi-trait AES dataset with 951 essays annotated with holistic scores and 34 rubric-based traits. Experiments on CFMS-34 and ASAP++ show that HiFTS achieves strong holistic and trait-level scoring while producing coherent, rubric-aligned feedback.

Related papers