yukajii / MT digest
LLM translation fine-tuning and cross-lingual legal QA
arXiv announcements of 23 September 2026 · 3 papers
If your translation system is going into production, the failure mode to worry about is not just lower quality but losing the controls users rely on: formality, grammatical gender, length, and other MT-specific instructions can slip after fine-tuning even when generic forgetting metrics look fine. A study on Llama 3.2 1B Instruct and Llama 3.1 8B tests forgetting-mitigation methods for MT fine-tuning, while a bilingual Vietnamese–English suite on Vietnamese labour law probes retrieval, translation, and verifier-guided correction across 231 question–answer pairs. The third paper is a survey of brain-to-language decoding, covering invasive and non-invasive signals, text generation, streaming personalised speech, and facial animation.
Elastic Weight Consolidation best protects general capabilities in MT fine-tuning, but only mixing control-task examples preserves formality and grammatical-gender control, and that benefit does not transfer to unseen prompts.
Abstract
Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific instruction following (MT-IF): instructions that modify a translation, such as formality, grammatical gender, and length control. We compare methods anchored to auxiliary data, to model outputs, and to the base model parameters, first in a screening study with Llama 3.2 1B Instruct, then on Llama 3.1 8B Instruct fine-tuned on bidirectional Arabic-English or Spanish-English data. Elastic Weight Consolidation preserves general capabilities best in both stages; on the 8B Spanish model the average score on general benchmarks drops 1.7 points versus 11.0 for standard fine-tuning, yet its scores for formality and grammatical gender control remain close to standard fine-tuning. Only data mixing with control-task examples preserves these controls, but its gains do not transfer to unseen prompts for the same task.
Brain-to-language decoding is a neurotech survey, but MT teams can borrow its emphasis on shared representations, online calibration, and user control when evaluating streaming or personalised speech interfaces.
Abstract
Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a means of studying how the brain represents language. Advances in neural recording and representation learning have expanded the field from constrained recognition and acoustic reconstruction to text generation, streaming personalised speech and facial animation. This survey synthesises these developments across invasive and non-invasive measurements, drawing on a search without a lower year limit and source-led updates through September 2026. We connect Articulated, Inner and Perceived tasks to the neural populations they engage, the representations available to decoders and the outputs those representations can support. We examine model development, public resources and the evolution of evaluation, and compare published performance and communication costs within their reported protocols. The synthesis identifies complementary routes to progress: phonetic, acoustic and semantic targets preserve different aspects of a message; shared representations support reuse across recording conditions and tasks; and online communication increasingly depends on calibration, feedback and user control alongside decoding accuracy. Shared benchmarks enable algorithmic comparisons, while longitudinal studies reveal the demands of sustained use. We discuss these developments and their remaining limitations, then outline a prospective five-level trajectory from commands and language to meaning, scenarios and bidirectional cognitive exchange
Dense retrieval beats learned-sparse retrieval for English-to-Vietnamese labour-law QA at R@5 0.358 versus 0.032, and verifier-guided correction mainly improves citation preservation by 0.022--0.034.
Abstract
Cross-lingual legal question answering must retrieve statutes across languages while preventing unsupported legal claims. We introduce a bilingual evaluation suite of 231 Vietnamese--English question--answer pairs from Vietnamese labour law. Of these, 75 are additionally annotated for five challenging legal reasoning phenomena. We evaluate a verifier-guided pipeline that decomposes answers into claims, checks citation reachability and entailment, and corrects citation failures and contradictions. We also introduce six automatic diagnostics for faithfulness to retrieved evidence, covering citations, modality, exceptions, procedures, conclusions, and evidential support. Experiments show that learned-sparse retrieval performs poorly for English-to-Vietnamese retrieval (R@5~=~0.032), whereas dense retrieval reaches 0.358 and slightly outperforms hybrid retrieval. Translation placement has no statistically detectable effect on these automatic diagnostics in our controlled comparison and supporting sensitivity analyses. Verifier-guided correction improves citation preservation by 0.022--0.034 at the system level but produces no reliable gains in the remaining dimensions. Human evaluation further shows that the automatic diagnostics do not fully align with human judgements of answer quality.