yukajii / MT digest

Reward models for GEC and gender bias in MT metrics

arXiv announcements of 20 September 2026 · 2 papers

A reference-based GEC scorer can be wrong for the reason that matters most: if the reference set misses a valid edit, M² and ERRANT will still mark it down. RM-EVAL, trained on SEEDA preference data, is a reference-free meta-evaluator that tracks human judgments at both full- and partial-sequence level, and the same reward model can be used as a learning signal for GEC generation. On the MT side, a WMT 2026 analysis of the occupation-balanced GAMBIT+ subset finds that evaluation metrics can inherit gender preferences even when the source leaves gender unspecified. The paper compares seven metrics on masculine versus feminine realizations, which makes the bias visible as an evaluation problem, not just a decoding one.

Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

Use reference-free reward models for GEC evaluation when gold edits miss valid paraphrased fixes, and consider reward-guided decoding, which improves generation without unfreezing the base model.

Abstract

Reference-based metrics for Grammatical Error Correction (GEC) such as M^2 and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We introduce RM-EVAL, a reward model trained on human preference data from SEEDA, as a reference-free meta-evaluator that predicts human-like quality judgments at both full-sequence and partial-sequence levels. Beyond evaluation, we show that the same reward model can be used as a learning signal to improve GEC generation via Reward-Guided Text Generation (RGTG), which keeps a base GEC model frozen and performs online, reward-driven decoding. Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.

Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations

Expect MT evaluation metrics to favor masculine forms for unspecified-gender occupations, with 1,308 paired examples per language showing stereotype-driven score gaps that vary by evaluator and target language.

Abstract

Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.