Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction

arXiv:2609.15559 · cs.CL, cs.AI · Submitted 2026-09-14 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-09-14

Updated: 2026-09-14

Comments: EMNLP 2026 Findings

Code: https://github.com/hayeonggg/SURE

License: http://creativecommons.org/licenses/by/4.0/

The gist: Grammatical error correction (GEC) evaluation has traditionally relied on reference or edit overlap, which can penalize valid rewrites that differ from gold corrections.

Terminology

Abstract

Grammatical error correction (GEC) evaluation has traditionally relied on reference or edit overlap, which can penalize valid rewrites that differ from gold corrections. Reference-free metrics reduce this dependence, but evaluating whether a fluent output is a valid correction of the source remains challenging. We propose SURE, a source-conditioned reward evaluator trained on within-source preferences spanning minimal-edit and rewrite-oriented corrections. SURE jointly learns an overall reward with criteria-level supervision for grammaticality, faithfulness, and fluency, together with span-level grounding for source-side error resolution. Experiments on SEEDA show that SURE performs competitively against strong baselines, with particular gains on rewrite-style corrections and more disentangled criteria-level diagnostics. Our code is available at https://github.com/hayeonggg/SURE.

Sources

Related papers