yukajii / MT digest

Communicative success in machine-translated conversation and Telugu spoken QA

arXiv announcements of 17 September 2026 · 3 papers

Shipping translation systems means more than getting sentence-level fidelity right: if your MT-powered interpreter stumbles on pragmatics, register, or socially awkward phrasing, the conversation still fails. The main paper here tackles that gap directly, while the rest of the issue is mostly adjacent NLP: a scientific image-quality RAG system and VākQA, a Telugu spoken QA benchmark with 2,001 factoid pairs and 2.53 hours of audio.

Evaluating Communicative Success in Machine-Translated Conversation

Use this checklist-and-judge framework for interpreter-style MT when you need to score semantic, pragmatic, and cultural-social success, because sentence-level fidelity metrics miss failures in multi-turn conversation.

Abstract

Interpreter agents built on machine translation (MT) increasingly mediate live conversation between people who do not share a language, yet we still evaluate them with metrics built for isolated sentences, which measure fidelity rather than whether communication succeeds. We introduce a reusable three-layer checklist-and-judge framework that evaluates interpreter-mediated conversation across semantic, pragmatic, and cultural-social dimensions, covering the naturalness, intent, and social appropriateness that fidelity metrics leave unmeasured. It runs in both single-turn and interactive multi-turn settings, where simulated users reply to translated messages as the conversation unfolds and each turn is scored alongside the conversation as a whole. We extensively validate it through controlled perturbations, cross-judge comparisons, and human annotations. Our main single-turn benchmark evaluates 10 interpreter setups across Arabic, Bengali, Indonesian, and Korean from 5,624 OpenSubtitles-derived scenarios spanning 12 translation directions, and our multi-turn study covers all 6 language pairs in scripted and live modes. Results show a consistent decline from semantic to pragmatic and cultural-social success, while conventional MT metrics overlook failures among stronger interpreters, and prompt ablations show that scenario context, structured instructions, and cultural context improve communicative success, although gains vary across setups. Our work thus provides an evaluation framework and benchmark for interpreter agents in conversation, and highlights the importance of communicative success alongside existing translation metrics.

Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

Scientific image QA has no translation component, but MT practitioners should note the RAG pattern: retrieve fine-grained exemplars and fuse them before judging, which may help reference-based evaluation setups.

Abstract

This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

Treat Telugu spoken QA as a caution for speech translation: open-weight judges over-penalise valid paraphrases, and cascaded ASR-MT errors compound, so human validation matters for non-English speech outputs.

Abstract

Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.