yukajii / MT digest

Speech translation reward learning and synthetic domain translation

arXiv announcements of 22 September 2026 · 2 papers

A speech-translation study on Qwen2.5-Omni-3B across four languages uses group relative policy optimization to score both transcripts and translations, tackling the train–test mismatch between reference transcripts in SFT and the model’s own transcripts at inference. Alongside it, TransBERT shows that a French life-sciences language model can be pretrained entirely on synthetically translated text, with TransCorpus as the translation pipeline behind the data.

Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation

For LLM speech translation, GRPO on model-generated transcripts and translations narrows the training-inference mismatch and gives CoT 1.77 BLEU over direct ST on CoVoST 2, so you should reward both steps, not just the output text.

Abstract

In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.

TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling

For domain NLP in French life sciences, TransBERT relies entirely on synthetic translation and still reaches strong downstream results, so synthetic parallel data can bootstrap low-resource domain resources when real text is scarce.

Abstract

The scarcity of non-English language data in specialized domains significantly limits the development of effective Natural Language Processing (NLP) tools. We present TransBERT, a novel framework for pre-training language models using exclusively synthetically translated text, and introduce TransCorpus, a scalable translation toolkit. Focusing on the life sciences domain in French, our approach demonstrates that state-of-the-art performance on various downstream tasks can be achieved solely by leveraging synthetically translated data. We release the TransCorpus toolkit, the TransCorpus-bio-fr corpus (36.4GB of French life sciences text), TransBERT-bio-fr, its associated pre-trained language model and reproducible code for both pre-training and fine-tuning. Our results highlight the viability of synthetic translation in a high-resource translation direction for building high-quality NLP resources in low-resource language/domain pairs.