| 2026-10-01 | Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models |
| 2026-10-01 | Counting and Min-Cost Encoding for Tokenization in Large Language Models |
| 2026-09-30 | Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation |
| 2026-09-30 | Index-Translate: A Multilingual Translation Model Family -- Text, Speech, Controlled Dubbing, and Long-Document Translation |
| 2026-09-30 | The Invisible Language Tax: Token Premiums of French and Regional Languages in 2026 LLM Tokenizers, and a French-Optimized Prototype |
| 2026-09-30 | When Scientific Contradictions Are Lost in Translation |
| 2026-09-30 | GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning |
| 2026-09-29 | Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision |
| 2026-09-29 | Generating Edit-Inducing Questions for AI Research Manuscripts |
| 2026-09-29 | MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization |
| 2026-09-29 | Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts |
| 2026-09-28 | Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL$\rightarrow$FOL |
| 2026-09-28 | BaatCheet: A Multilingual Corpus for Dialogue Translation in Indian Languages |
| 2026-09-28 | NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech |
| 2026-09-28 | MixDetect: Word-Level Localization and Quantification of AI Editing |
| 2026-09-28 | Program-Verified Self-Evolution for Vision-Language Models |
| 2026-09-27 | Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting |
| 2026-09-27 | G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation |
| 2026-09-24 | COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages |
| 2026-09-24 | Tag-Aware Structured Text Translation: Towards a Systematic Understanding |
| 2026-09-24 | TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification |
| 2026-09-24 | Scoring Both Directions: LLMs realize the MRS they cannot reliably parse |
| 2026-09-23 | Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following |
| 2026-09-23 | Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond |
| 2026-09-23 | Cross-Lingual Legal QA for Vietnamese Labour Law: Retrieval, Translation, and Verifier-Guided Correction |
| 2026-09-22 | Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation |
| 2026-09-22 | TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling |
| 2026-09-21 | Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation |
| 2026-09-21 | Machine-Interpretable Information: Compiling Documents into Searchable and Readable Protocol States |
| 2026-09-21 | When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs |
| 2026-09-20 | Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction |
| 2026-09-20 | Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations |
| 2026-09-17 | Evaluating Communicative Success in Machine-Translated Conversation |
| 2026-09-17 | Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation |
| 2026-09-17 | VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering |
| 2026-09-16 | LocQE: Principled Domain Adaptation for Localisation Quality Estimation by Leveraging Post-Edits |
| 2026-09-16 | TACTICS: Taxonomy-Aware Intelligent Corpus Sampling for Machine Translation |
| 2026-09-16 | Machine Translation between English and Syriac (East Syriac Dialect) using Statistical Machine Learning |
| 2026-09-16 | TeochewBench: A Human-Reviewed Benchmark for Teochew Hanzi Translation |
| 2026-09-16 | A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages |
| 2026-09-15 | Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM |
| 2026-09-15 | Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs |
| 2026-09-15 | Target-Language Generation in Multilingual Models: Activation Steering and Optimal Control |
| 2026-09-15 | Challenges of Auditing: Variability in Outputs of Large Language Models for Health |
| 2026-09-15 | Enhancing Accessibility of Medical Texts through Large Language Model-Driven Plain Language Adaptation |
| 2026-09-14 | Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction |
| 2026-09-14 | ReMova: Fine-tuning LLMs for English to Belarusian translation |
| 2026-09-14 | StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation |
| 2026-09-14 | Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA |
| 2026-09-14 | Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation |
| 2026-09-13 | Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking |
| 2026-09-13 | Investigating Temporal Motion Features for Pose-to-Text Indian Sign Language Translation |
| 2026-09-11 | In the Blind: Building Pseudo-References for MT Evaluation |
| 2026-09-11 | RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines |
| 2026-09-11 | Generative Interpretability via Scalable Neuro-Symbolic Models |
| 2026-09-10 | TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs |
| 2026-09-10 | Domain-Specific Hallucination Detection in Large Language Models |
| 2026-09-10 | SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ |
| 2026-09-10 | Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting |
| 2026-09-10 | LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation |
| 2026-09-09 | SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers |
| 2026-09-09 | $S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants |
| 2026-09-09 | BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation |
| 2026-09-09 | Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction |
| 2026-09-09 | Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs |
| 2026-09-07 | Replicating a Disjoint-Set Union Experiment over Various Notions of Micro Units to assess Translation Effort |
| 2026-09-07 | Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition |
| 2026-09-07 | Retrieval-Augmented Multi-Prompt Ensemble for Minor-Grain Breeding Information Extraction |
| 2026-09-07 | LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error and Robustness Nuances |
| 2026-09-07 | Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining |
| 2026-09-04 | EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages |
| 2026-09-04 | Discourse Dependency: A Continuous Criterion for Translation Difficulty |
| 2026-09-04 | Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3 |
| 2026-09-04 | LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics |
| 2026-09-04 | MoirfEolas and CríochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology |
| 2026-09-03 | Last Translation Benchmark |
| 2026-09-03 | Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks |
| 2026-09-03 | Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation |
| 2026-09-03 | HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews |
| 2026-09-03 | Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language |
| 2026-09-02 | SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval |
| 2026-09-02 | MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts |
| 2026-09-02 | A Layered Taxonomy for Chinese Learner Grammatical Error Annotation |
| 2026-09-02 | No country for old linguists: LLM-brain alignment underdetermines neural computation |
| 2026-09-02 | A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models |
| 2026-09-01 | Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking |
| 2026-09-01 | The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space |
| 2026-09-01 | On the Design Fundamentals of Pixel Text Representation Learning |
| 2026-09-01 | TWIX: a Two-Stage Approach for End-To-End Named Entity Recognition and Relation Extraction |
| 2026-09-01 | Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding |
| 2026-08-31 | OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding |
| 2026-08-31 | Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation |
| 2026-08-31 | MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts |
| 2026-08-31 | From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education |
| 2026-08-31 | Latent Mechanisms of Language Control in Multilingual Language Models |
| 2026-08-28 | Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study |
| 2026-08-28 | NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry |
| 2026-08-28 | CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms |
| 2026-08-28 | SimpCue: Cue-Based Prompting for Multilingual Text Simplification |
| 2026-08-28 | Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images |
| 2026-08-27 | STAR : Sentence Translation Alignment Rate for Document-to-Document Machine Translation |
| 2026-08-27 | Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper |
| 2026-08-27 | TTPO: Test-Time Policy Optimization |
| 2026-08-27 | Reasoning about In-Context Samples for Machine-Translation |
| 2026-08-27 | Beyond Reflection: Affirmation as a Promising Behavioral Marker Associated with Quality in Text-Based Counseling |
| 2026-08-26 | Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models |
| 2026-08-26 | Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training |
| 2026-08-26 | VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text |
| 2026-08-26 | Cross-lingual Representation Learning via Centroid Intervention Fusion |
| 2026-08-26 | One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography |
| 2026-08-25 | A Primer on Computational Semantics for Artificial Intelligence Systems |
| 2026-08-25 | Linear Probing Provides Robust and Efficient Detection of Machine-Generated Text |
| 2026-08-25 | Shortcut Before Circuit: Document Statistics Time In-Context Conflict Resolution |
| 2026-08-25 | Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts |
| 2026-08-25 | DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration |
| 2026-08-24 | Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study |
| 2026-08-24 | Language Chain in Alignment: Cross-lingual Ranking Preference Optimization |
| 2026-08-24 | Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text |
| 2026-08-24 | The Multilingual FrameNet Corpus |
| 2026-08-24 | How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines |
| 2026-08-21 | Source-Free MT Evaluation Is Not MT Evaluation |
| 2026-08-21 | Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images |
| 2026-08-21 | COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models |
| 2026-08-21 | Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation |
| 2026-08-21 | Scaling Unsupervised Word Alignment to Documents via Structural Constraints |
| 2026-08-20 | HealMed: Multilingual Evaluation of Large Language Models in Medicine |
| 2026-08-20 | CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance |
| 2026-08-20 | G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation |
| 2026-08-20 | FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models |
| 2026-08-20 | ContractScrub: A benchmark for final review of legal contracts |
| 2026-08-19 | Assessing Quality of Experience in Natural Language Generation of German Text |
| 2026-08-19 | OmniAlign: A Unified Multilingual Aligner for Word and Sentence Alignment |
| 2026-08-19 | Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale |
| 2026-08-19 | TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation |
| 2026-08-19 | When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation |
| 2026-08-18 | Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges |
| 2026-08-18 | BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models |
| 2026-08-18 | Code as Representation: A Compilable Parsing Paradigm for Academic Documents |
| 2026-08-18 | An Investigation of Translationese in the Generations of Multilingual Large Language Models |
| 2026-08-18 | TokEval: A Tokenizer Evaluation Suite |
| 2026-08-17 | IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages |
| 2026-08-17 | Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers? |
| 2026-08-17 | Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization |
| 2026-08-17 | Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation |
| 2026-08-17 | Uncertainty-Aware Decision Making in Multimodal Large Language Models |
| 2026-08-14 | Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models |
| 2026-08-14 | From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving |
| 2026-08-14 | Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision |
| 2026-08-14 | Seeing Red, Thinking Bad: Color Bias in Vision Language Models |
| 2026-08-14 | A Survey of Large Models in Sports |
| 2026-08-13 | BM25-Augmented Many-Shot Translation for Low-Resource North-Eastern Indian Languages |
| 2026-08-13 | Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering |
| 2026-08-13 | How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures |
| 2026-08-13 | When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics |
| 2026-08-13 | ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval |
| 2026-08-12 | Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects |
| 2026-08-12 | When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use |
| 2026-08-12 | TELLME: Test-Enhanced Learning for Language Model Enrichment |
| 2026-08-12 | When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation |
| 2026-08-12 | Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages |
| 2026-08-11 | Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation |
| 2026-08-11 | On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation |
| 2026-08-11 | InSight-doc: Agentic Visual Perception for Long-Document Understanding |
| 2026-08-11 | Most biomedical publications show signs of LLM-assisted writing |
| 2026-08-11 | VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? |
| 2026-08-10 | Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness |
| 2026-08-10 | Failure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation System |
| 2026-08-10 | Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing |
| 2026-08-10 | ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models |
| 2026-08-10 | Subjective Multi-Bias Detection with Large Language Models |
| 2026-08-07 | Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation |
| 2026-08-07 | SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators |
| 2026-08-07 | LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering |
| 2026-08-07 | Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control |
| 2026-08-07 | Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages |
| 2026-08-06 | Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI |
| 2026-08-06 | Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation |
| 2026-08-06 | Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents |
| 2026-08-06 | Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness |
| 2026-08-06 | M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding |
| 2026-08-05 | AI Literacy for Legal Translation: Developing Digital Resilience |
| 2026-08-05 | Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings |
| 2026-08-05 | Analysis of Numerical Localisation in LLM Translations |
| 2026-08-05 | Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference |
| 2026-08-05 | Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos |
| 2026-08-04 | Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation |
| 2026-08-04 | Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation |
| 2026-08-04 | Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study |
| 2026-08-04 | PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation |
| 2026-08-04 | TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation |
| 2026-08-03 | Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study |
| 2026-08-03 | CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning |
| 2026-08-03 | Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer |
| 2026-08-03 | Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction |
| 2026-08-03 | Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation |
| 2026-07-31 | Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation |
| 2026-07-31 | Studying quantization trade-offs for efficient inference deployment in machine translation |
| 2026-07-31 | Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art |
| 2026-07-31 | Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models |
| 2026-07-31 | Cross-Lingual Transfer for Machine Translation in Turkic Languages |
| 2026-07-30 | AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification |
| 2026-07-30 | CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising |
| 2026-07-30 | Challenges in annotations by humans and LLMs: A case study of evaluative language |
| 2026-07-30 | DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation |
| 2026-07-30 | Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B |
| 2026-07-29 | Contrastive ESA: Human Evaluation of Multiple Translations at Once |
| 2026-07-29 | MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis |
| 2026-07-29 | DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models |
| 2026-07-29 | DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search |
| 2026-07-29 | Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation |
| 2026-07-28 | Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications |
| 2026-07-28 | AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review |
| 2026-07-28 | Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation |
| 2026-07-28 | MemSFT: Mitigating Alignment Tax with an External Parametric Memory |
| 2026-07-28 | Evaluation of forced alignment of code-mixed speech: the case of Hindi-English |
| 2026-07-27 | Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets |
| 2026-07-27 | Occluded Oculus: Operationalizing Stylistic Obscurement |
| 2026-07-27 | From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages |
| 2026-07-27 | BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes |
| 2026-07-27 | DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification |
| 2026-07-24 | grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP |
| 2026-07-24 | A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books |
| 2026-07-24 | PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs |
| 2026-07-24 | Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging |
| 2026-07-24 | Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge |
| 2026-07-23 | Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms |
| 2026-07-23 | REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning |
| 2026-07-23 | news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling |
| 2026-07-23 | An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations |
| 2026-07-23 | Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin |
| 2026-07-22 | On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens |
| 2026-07-22 | TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management |
| 2026-07-22 | Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering |
| 2026-07-22 | Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering |
| 2026-07-22 | Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking |
| 2026-06-05 | Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses |
| 2026-06-05 | MMAE: A Massive Multitask Audio Editing Benchmark |
| 2026-06-05 | Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning |
| 2026-06-05 | Auditing Training Data in Domain-adapted LLMs: LoRA-MINT |
| 2026-06-05 | M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions |
| 2026-06-04 | Automatic Labelling of Speech Translation Errors |
| 2026-06-04 | Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation |
| 2026-06-04 | Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach |
| 2026-06-04 | "Chi nas dal soch el sent de legn" -- Auditing Text Corpora for Lombard |
| 2026-06-04 | Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition |
| 2026-06-03 | ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation |
| 2026-06-03 | A French Corpus Annotated for Multiword Expressions with Adverbial Function |
| 2026-06-03 | NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models |
| 2026-06-03 | M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks |
| 2026-06-03 | Automatic Generation of Titles for Research Papers Using Language Models |
| 2026-06-02 | G^2C-MT: Graph-Guided Context Selection for Document-Level Machine Translation |
| 2026-06-02 | KletterMix: Climbing Toward High-Quality German Pretraining Data |
| 2026-06-02 | Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models |
| 2026-06-02 | Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation? |
| 2026-06-02 | Benchmarking Speech-to-Speech Translation Models |
| 2026-06-01 | CARTE: A Benchmark for Mapping Language Model Knowledge Across France |
| 2026-06-01 | HLL: Can Agents Cross Humanity's Last Line of Verification? |
| 2026-06-01 | Translating Classical Poetry into Modern Prose |
| 2026-06-01 | Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French |
| 2026-06-01 | Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling |
| 2026-05-29 | Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows |
| 2026-05-29 | Model-Based Quality Assessment for Massively Multilingual Parallel Data |
| 2026-05-29 | Unlocking Fine-Grained Translation Quality Estimation in LRMs through Synergistically Evolving Implicit and Explicit Reasoning |
| 2026-05-29 | TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices |
| 2026-05-29 | Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion |
| 2026-05-28 | Comparative Evaluation of Machine Translation Systems on Images with Text |
| 2026-05-28 | Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection |
| 2026-05-28 | Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering |
| 2026-05-28 | Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation |
| 2026-05-28 | AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation |
| 2026-05-27 | Why We Need Speech to Evaluate Speech Translation |
| 2026-05-27 | IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following |
| 2026-05-27 | Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction |
| 2026-05-27 | Breaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin Scripts |
| 2026-05-27 | Challenges in Explaining Pretrained Clinical Text Classifiers |
| 2026-05-26 | BhashaSetu: A Data-Centric Approach to Low-Resource Machine Translation |
| 2026-05-26 | GraphReview: Scientific Paper Evaluation via LLM-Based Graph Message Passing |
| 2026-05-26 | Slide Deck Q&A Quality Assurance App: A Multi-Stage Pipeline for Pedagogical Question Generation |
| 2026-05-26 | Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases |
| 2026-05-26 | Rethinking the Multilingual Reasoning Gap with Layer Swap |
| 2026-05-25 | Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC |
| 2026-05-25 | Testing the Deliteralization Hypothesis in Human and Machine Translation |
| 2026-05-25 | Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals |
| 2026-05-25 | ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence |
| 2026-05-25 | What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA |
| 2026-05-22 | Hidden Human-Like Nature of Machine-Generated Texts: Theory and Detection Enhancement |
| 2026-05-22 | NLG Evaluation: Past, Present, Future |
| 2026-05-22 | How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework |
| 2026-05-22 | ETCHR: Editing To Clarify and Harness Reasoning |
| 2026-05-22 | ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models |
| 2026-05-20 | Enhancing Scientific Discourse: Machine Translation for the Scientific Domain |
| 2026-05-20 | Smarter edits? Post-editing with error highlights and translation suggestions |
| 2026-05-20 | Metaphors in Literary Post-Editing: Opening Pandora's Box? |
| 2026-05-20 | MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks |
| 2026-05-20 | ACL-Verbatim: hallucination-free question answering for research |
| 2026-05-19 | AI Technologies in Language Access: Attitudes Towards AI and the Human Value of Language Access Managers |
| 2026-05-19 | NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding |
| 2026-05-19 | A Data-Driven Approach to Idiomaticity Based on Experts' Criteria in Theoretical Linguistics |
| 2026-05-19 | Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP |
| 2026-05-19 | What Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading Experience |
| 2026-05-18 | Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models |
| 2026-05-18 | A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration |
| 2026-05-18 | From Documents to Segments: A Contextual Reformulation for Topic Assignment |
| 2026-05-18 | FOL2NS: Generating Natural Sentences from First-Order Logic |
| 2026-05-18 | How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking |
| 2026-05-16 | Agentic AI Translate: An Agentic Translator Prototype for Translation as Communication Design |
| 2026-05-16 | How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study |
| 2026-05-16 | Closing the Gap at CRAC 2026: Two-Stage Adaptation for LLM-Based Multilingual Coreference Resolution |
| 2026-05-16 | UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation |
| 2026-05-16 | PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media |
| 2026-05-15 | CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs |
| 2026-05-15 | Reference-Free Reinforcement Learning Fine-Tuning for MT: A Seq2Seq Perspective |
| 2026-05-15 | ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation |
| 2026-05-15 | DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection |
| 2026-05-15 | Reasoners or Translators? Contamination-aware Evaluation and Neuro-Symbolic Robustness in Tax Law |
| 2026-05-12 | Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs |
| 2026-05-12 | Towards Visually-Guided Movie Subtitle Translation for Indic Languages |
| 2026-05-12 | DiffScore: Text Evaluation Beyond Autoregressive Likelihood |
| 2026-05-12 | DocAtlas: Multilingual Document Understanding Across 80+ Languages |
| 2026-05-12 | Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability |
| 2026-05-11 | Evolving Knowledge Distillation for Lightweight Neural Machine Translation |
| 2026-05-11 | BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation |
| 2026-05-11 | ICT-NLP at SemEval-2026 Task 3: Less Is More -- Multilingual Encoder with Joint Training and Adaptive Ensemble for Dimensional Aspect Sentiment Regression |
| 2026-05-11 | Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs |
| 2026-05-11 | Coherency through formalisations of Structured Natural Language, A case study on FRETish |
| 2026-05-10 | Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification |
| 2026-05-10 | Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution |
| 2026-05-10 | Towards Compact Sign Language Translation: Frame Rate and Model Size Trade-offs |
| 2026-05-10 | Edit-Based Refinement for Parallel Masked Diffusion Language Models |
| 2026-05-10 | TAD: Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM |
| 2026-05-09 | Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan |
| 2026-05-09 | Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation |
| 2026-05-09 | 100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts |
| 2026-05-09 | LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language |
| 2026-05-09 | DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding |
| 2026-05-07 | YEZE at SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization via Heterogeneous Ensembling |
| 2026-05-07 | Reflections and New Directions for Human-Centered Large Language Models |
| 2026-05-07 | Algospeak, Hiding in the Open: The Trade-off Between Legible Meaning and Detection Avoidance |
| 2026-05-07 | Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text |
| 2026-05-07 | MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media |
| 2026-05-06 | Beyond BLEU: A Semantic Evaluation Method for Code Translation |
| 2026-05-06 | Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models |
| 2026-05-06 | Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting |
| 2026-05-06 | DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation |
| 2026-05-06 | MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge |
| 2026-05-05 | Natural Language Processing: A Comprehensive Practical Guide from Tokenisation to RLHF |
| 2026-05-05 | A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language |
| 2026-05-05 | From prompting to evidence-based translation: A RAG+prompt system for Japanese-Chinese translation and its pedagogical potential |
| 2026-05-05 | Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages |
| 2026-05-05 | A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition |
| 2026-05-04 | A multilingual hallucination benchmark: MultiWikiQHalluA |
| 2026-05-04 | PC-MNet: Dual-Level Congruity Modeling for Multimodal Sarcasm Detection via Polarity-Modulated Attention |
| 2026-05-04 | SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures |
| 2026-05-04 | A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures |
| 2026-05-04 | HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs |
| 2026-05-03 | A Multimodal Dataset for Visually Grounded Ambiguity in Machine Translation |
| 2026-05-03 | RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences |
| 2026-05-03 | EditPropBench: Measuring Factual Edit Propagation in Scientific Manuscripts |
| 2026-05-03 | Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-Gated LLM |
| 2026-05-03 | Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality |
| 2026-05-02 | Artificial intelligence language technologies in multilingual healthcare: Grand challenges ahead |
| 2026-05-02 | Benchmarking LightGBM and BiLSTM for Sentiment Analysis on Indonesian E-Commerce Reviews |
| 2026-05-02 | Sentiment Analysis of Mobile Legends App Reviews Using Machine Learning and LSTM-Based Deep Learning Models |
| 2026-05-02 | OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice |
| 2026-05-02 | Fine-Tuning Pre-Trained Code Models for AI-Generated Code Detection |
| 2026-05-01 | Language-free Experience at Expo 2025 Osaka |
| 2026-05-01 | Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus |
| 2026-05-01 | Borrowed Geometry: Computational Reuse of Frozen Text-Pretrained Transformer Weights Across Modalities |
| 2026-05-01 | Escaping Mode Collapse in LLM Generation via Geometric Regulation |
| 2026-05-01 | Democratizing the medieval English legal tradition |
| 2026-04-29 | Text Style Transfer with Machine Translation for Graphic Designs |
| 2026-04-29 | TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models |
| 2026-04-29 | Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI |
| 2026-04-29 | Classification of Public Opinion on the Free Nutritional Meal Program on YouTube Media Using the LSTM Method |
| 2026-04-29 | Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping |
| 2026-04-28 | Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation |
| 2026-04-28 | MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors |
| 2026-04-28 | Language corpora for the Dutch medical domain |
| 2026-04-28 | One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech |
| 2026-04-28 | LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization |
| 2026-04-26 | Pref-CTRL: Preference Driven LLM Alignment using Representation Editing |
| 2026-04-26 | Neural Grammatical Error Correction for Romanian |
| 2026-04-26 | Applications of the Transformer Architecture in AI-Assisted English Reading Comprehension |
| 2026-04-26 | Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms |
| 2026-04-26 | Translate or Simplify First: An Analysis of Cross-lingual Text Simplification in English and French |
| 2026-04-25 | $\mathcal{S}^2$IT: Stepwise Syntax Integration Tuning for Large Language Models in Aspect Sentiment Quad Prediction |
| 2026-04-25 | Evaluating Large Language Models on Computer Science University Exams in Data Structures |
| 2026-04-25 | Au-M-ol: A Unified Model for Medical Audio and Language Understanding |
| 2026-04-25 | Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective |
| 2026-04-25 | A Parametric Memory Head for Continual Generative Retrieval |
| 2026-04-24 | TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction |
| 2026-04-24 | Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models |
| 2026-04-24 | Evaluating Temporal Consistency in Multi-Turn Language Models |
| 2026-04-24 | TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis |
| 2026-04-24 | RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment |
| 2026-04-23 | Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages |
| 2026-04-23 | Incentivizing Neuro-symbolic Language-based Reasoning in VLMs via Reinforcement Learning |
| 2026-04-23 | mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code |
| 2026-04-23 | When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs |
| 2026-04-23 | Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models |
| 2026-04-21 | ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation |
| 2026-04-21 | SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models |
| 2026-04-21 | IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text |
| 2026-04-21 | Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews |
| 2026-04-21 | Bangla Key2Text: Text Generation from Keywords for a Low Resource Language |
| 2026-04-20 | LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation |
| 2026-04-20 | ltzGLUE: Luxembourgish General Language Understanding Evaluation |
| 2026-04-20 | Multilingual Training and Evaluation Resources for Vision-Language Models |
| 2026-04-20 | Model in Distress: Sentiment Analysis on French Synthetic Social Media |
| 2026-04-20 | Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck? |
| 2026-04-19 | Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains |
| 2026-04-19 | Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining |
| 2026-04-19 | Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions |
| 2026-04-19 | Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation |
| 2026-04-19 | MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation |
| 2026-04-18 | MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation |
| 2026-04-18 | MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning |
| 2026-04-18 | Beyond Black-Box Labels: Interpretable Criteria for Diagnosing SubjectiveNLP Tasks |
| 2026-04-18 | Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation |
| 2026-04-18 | How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them |
| 2026-04-17 | The impact of postediting on AI generative translation in Yemeni context: Translating literary prose by ChatGPT |
| 2026-04-17 | MUSCAT: MUltilingual, SCientific ConversATion Benchmark |
| 2026-04-17 | Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures |
| 2026-04-17 | Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models |
| 2026-04-17 | JFinTEB: Japanese Financial Text Embedding Benchmark |
| 2026-04-16 | Fabricator or dynamic translator? |
| 2026-04-16 | XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics |
| 2026-04-16 | Domain Fine-Tuning FinBERT on Finnish Histopathological Reports: Train-Time Signals and Downstream Correlations |
| 2026-04-16 | MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events |
| 2026-04-16 | IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation |
| 2026-04-15 | MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging |
| 2026-04-15 | EuropeMedQA Study Protocol: A Multilingual, Multimodal Medical Examination Dataset for Language Model Evaluation |
| 2026-04-15 | Applied Explainability for Large Language Models: A Comparative Study |
| 2026-04-15 | Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance |
| 2026-04-15 | MARCA: A Checklist-Based Benchmark for Multilingual Web Search |
| 2026-04-14 | Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation |
| 2026-04-14 | MetFuse: Figurative Fusion between Metonymy and Metaphor |
| 2026-04-14 | ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection |
| 2026-04-14 | Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection |
| 2026-04-14 | Calibrated Confidence Estimation for Tabular Question Answering |
| 2026-04-13 | INDOTABVQA: A Benchmark for Cross-Lingual Table Understanding in Bahasa Indonesia Documents |
| 2026-04-13 | Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning |
| 2026-04-13 | Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges |
| 2026-04-13 | Please Make it Sound like Human: Encoder-Decoder vs. Decoder-Only Transformers for AI-to-Human Text Style Transfer |
| 2026-04-13 | NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment |
| 2026-04-12 | Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance |
| 2026-04-12 | Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs |
| 2026-04-12 | When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities |
| 2026-04-12 | Lost in Diffusion: Uncovering Hallucination Patterns and Failure Modes in Diffusion Large Language Models |
| 2026-04-12 | Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS |
| 2026-04-11 | Comparative Analysis of Large Language Models in Healthcare |
| 2026-04-11 | Think in Sentences: Explicit Sentence Boundaries Enhance Language Model's Capabilities |
| 2026-04-11 | ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification |
| 2026-04-11 | Hijacking Text Heritage: Hiding the Human Signature through Homoglyphic Substitution |
| 2026-04-11 | Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models |
| 2026-04-10 | Should We be Pedantic About Reasoning Errors in Machine Translation? |
| 2026-04-10 | MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator |
| 2026-04-10 | Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated Probabilities |
| 2026-04-10 | Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning |
| 2026-04-10 | Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition |
| 2026-04-09 | The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training |
| 2026-04-09 | AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation |
| 2026-04-09 | TEMPER: Testing Emotional Perturbation in Quantitative Reasoning |
| 2026-04-09 | Quantum Vision Theory Applied to Audio Classification for Deepfake Speech Detection |
| 2026-04-09 | LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs |
| 2026-04-08 | ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding |
| 2026-04-08 | Geometric Properties of the Voronoi Tessellation in Latent Semantic Manifolds of Large Language Models |
| 2026-04-08 | Agent-Driven Corpus Linguistics: A Framework for Autonomous Linguistic Discovery |
| 2026-04-08 | Video-guided Machine Translation with Global Video Context |
| 2026-04-08 | Discourse Coherence and Response-Guided Context Rewriting for Multi-Party Dialogue Generation |
| 2026-04-07 | The Model Agreed, But Didn't Learn: Diagnosing Surface Compliance in Large Language Models |
| 2026-04-07 | Context-Aware Dialectal Arabic Machine Translation with Interactive Region and Register Selection |
| 2026-04-07 | Confidence Should Be Calibrated More Than One Turn Deep |
| 2026-04-07 | DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions |
| 2026-04-07 | GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding |
| 2026-04-06 | MERIT: Multilingual Expert-Reward Informed Tuning for Chinese-Centric Low-Resource Machine Translation |
| 2026-04-06 | EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content |
| 2026-04-06 | HUKUKBERT: Domain-Specific Language Model for Turkish Law |
| 2026-04-06 | Watch Before You Answer: Learning from Visually Grounded Post-Training |
| 2026-04-06 | Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding |
| 2026-04-05 | Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges |
| 2026-04-05 | Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language Models |
| 2026-04-05 | Lexical Indicators of Mind Perception in Human-AI Companionship |
| 2026-04-05 | RUQuant: Towards Refining Uniform Quantization for Large Language Models |
| 2026-04-05 | AdaptFuse: Training-Free Sequential Preference Learning via Externalized Bayesian Inference |
| 2026-03-26 | Toward domain-specific machine translation and quality estimation systems |
| 2026-03-26 | Cross-Preference Learning for Sentence-Level and Context-Aware Machine Translation |
| 2026-03-26 | Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages |
| 2026-03-26 | Beyond Detection: Rethinking Education in the Age of AI-writing |
| 2026-03-26 | Bilingual Text-to-Motion Generation: A New Benchmark and Baselines |
| 2026-03-25 | Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset |
| 2026-03-25 | Improving Lean4 Autoformalization via Cycle Consistency Fine-tuning |
| 2026-03-25 | Thinking with Tables: Enhancing Multi-Modal Tabular Understanding via Neuro-Symbolic Reasoning |
| 2026-03-25 | Robust Multilingual Text-to-Pictogram Mapping for Scalable Reading Rehabilitation |
| 2026-03-25 | Self-Distillation for Multi-Token Prediction |
| 2026-03-24 | Multilingual KokoroChat: A Multi-LLM Ensemble Translation Method for Creating a Multilingual Counseling Dialogue Dataset |
| 2026-03-24 | Swiss-Bench SBP-002: A Frontier Model Comparison on Swiss Legal and Regulatory Tasks |
| 2026-03-24 | LLMORPH: Automated Metamorphic Testing of Large Language Models |
| 2026-03-24 | From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service |
| 2026-03-24 | Is AI Catching Up to Human Expression? Exploring Emotion, Personality, Authorship, and Linguistic Style in English and Arabic with Six Large Language Models |
| 2026-03-23 | Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation |
| 2026-03-23 | SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding |
| 2026-03-23 | Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages |
| 2026-03-23 | Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning |
| 2026-03-23 | Parameter-Efficient Fine-Tuning for Medical Text Summarization: A Comparative Study of Lora, Prompt Tuning, and Full Fine-Tuning |
| 2026-03-22 | ReasonScaffold: A Scaffolded Reasoning-based Annotation Protocol for Human-AI Co-Annotation |
| 2026-03-22 | CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs |
| 2026-03-22 | Reading Between the Lines: How Electronic Nonverbal Cues shape Emotion Decoding |
| 2026-03-22 | Enhancing reasoning accuracy in large language models during inference time |
| 2026-03-22 | Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF |
| 2026-03-21 | Can ChatGPT Really Understand Modern Chinese Poetry? |
| 2026-03-21 | PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs |
| 2026-03-21 | Mitigating Shortcut Reasoning in Language Models: A Gradient-Aware Training Approach |
| 2026-03-21 | NoveltyAgent: Autonomous Novelty Reporting Agent with Point-wise Novelty Analysis and Self-Validation |
| 2026-03-21 | MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages |
| 2026-03-20 | Span-Level Machine Translation Meta-Evaluation |
| 2026-03-20 | EvidenceRL: Reinforcing Evidence Consistency for Trustworthy Language Models |
| 2026-03-20 | Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study |
| 2026-03-20 | Diffutron: A Masked Diffusion Language Model for Turkish Language |
| 2026-03-20 | Current LLMs still cannot 'talk much' about grammar modules: Evidence from syntax |
| 2026-03-19 | Automatic detection of Gen-AI texts: A comparative framework of neural models |
| 2026-03-19 | Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders |
| 2026-03-19 | VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models |
| 2026-03-19 | TARo: Token-level Adaptive Routing for LLM Test-time Alignment |
| 2026-03-19 | Detecting Basic Values in A Noisy Russian Social Media Text Data: A Multi-Stage Classification Framework |
| 2026-03-18 | From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation |
| 2026-03-18 | From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation |
| 2026-03-18 | ConGA: Guidelines for Contextual Gender Annotation. A Framework for Annotating Gender in Machine Translation |
| 2026-03-18 | Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions |
| 2026-03-18 | Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures |
| 2026-03-17 | Omnilingual MT: Machine Translation for 1,600 Languages |
| 2026-03-17 | Ensemble Self-Training for Unsupervised Machine Translation |
| 2026-03-17 | Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic |
| 2026-03-17 | Multilingual Reference Need Assessment System for Wikipedia |
| 2026-03-17 | Evaluating Ill-Defined Tasks in Large Language Models |
| 2026-03-16 | Towards Privacy-Preserving Machine Translation at the Inference Stage: A New Task and Benchmark |
| 2026-03-16 | Machine Translation in the Wild: User Reaction to Xiaohongshu's Built-In Translation Feature |
| 2026-03-16 | Developing an English-Efik Corpus and Machine Translation System for Digitization Inclusion |
| 2026-03-16 | Interpretable Predictability-Based AI Text Detection: A Replication Study |
| 2026-03-16 | SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia |
| 2026-03-15 | Parameter-Efficient Quality Estimation via Frozen Recursive Models |
| 2026-03-15 | Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective |
| 2026-03-15 | ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation |
| 2026-03-15 | Fine-tuning MLLMs Without Forgetting Is Easier Than You Think |
| 2026-03-15 | AI Can Learn Scientific Taste |
| 2026-03-14 | NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments |
| 2026-03-14 | Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models |
| 2026-03-14 | Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering |
| 2026-03-14 | MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos |
| 2026-03-14 | CMHL: Contrastive Multi-Head Learning for Emotionally Consistent Text Classification |
| 2026-03-13 | Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation |
| 2026-03-13 | Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation |
| 2026-03-13 | From Text to Forecasts: Bridging Modality Gap with Temporal Evolution Semantic Space |
| 2026-03-13 | Continual Learning in Large Language Models: Methods, Challenges, and Opportunities |
| 2026-03-13 | 98$\times$ Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router |
| 2026-03-12 | Semi-Synthetic Parallel Data for Translation Quality Estimation: A Case Study of Dataset Building for an Under-Resourced Language Pair |
| 2026-03-12 | Streaming Translation and Transcription Through Speech-to-Text Causal Alignment |
| 2026-03-12 | Trust Oriented Explainable AI for Fake News Detection |
| 2026-03-12 | BLooP: Zero-Shot Abstractive Summarization using Large Language Models with Bigram Lookahead Promotion |
| 2026-03-12 | EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models |
| 2026-03-11 | Large Language Models as Annotators for Machine Translation Quality Estimation |
| 2026-03-11 | Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation |
| 2026-03-11 | mAceReason-Math: A Dataset of High-Quality Multilingual Math Problems Ready For RLVR |
| 2026-03-11 | MUNIChus: Multilingual News Image Captioning Benchmark |
| 2026-03-11 | End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering |
| 2026-03-10 | EPIC-EuroParl-UdS: Information-Theoretic Perspectives on Translation and Interpreting |
| 2026-03-10 | Calibration-Reasoning Framework for Descriptive Speech Quality Assessment |
| 2026-03-10 | CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research? |
| 2026-03-10 | LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation |
| 2026-03-10 | RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation |
| 2026-03-09 | \$OneMillion-Bench: How Far are Language Agents from Human Experts? |
| 2026-03-09 | NCL-UoR at SemEval-2026 Task 5: Embedding-Based Methods, Fine-Tuning, and LLMs for Word Sense Plausibility Rating |
| 2026-03-09 | MultiGraSCCo: A Multilingual Anonymization Benchmark with Annotations of Personal Identifiers |
| 2026-03-09 | Gender Bias in MT for a Genderless Language: New Benchmarks for Basque |
| 2026-03-09 | Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference |
| 2026-03-08 | Image Generation Models: A Technical History |
| 2026-03-08 | Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning |
| 2026-03-08 | An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data |
| 2026-03-08 | Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data |
| 2026-03-08 | Benchmarking Large Language Models for Quebec Insurance: From Closed-Book to Retrieval-Augmented Generation |
| 2026-03-07 | Domain-Specific Quality Estimation for Machine Translation in Low-Resource Scenarios |
| 2026-03-07 | How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection |
| 2026-03-07 | To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise |
| 2026-03-07 | RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts |
| 2026-03-07 | Emotion Transcription in Conversation: A Benchmark for Capturing Subtle and Complex Emotional States through Natural Language |
| 2026-03-06 | MedInjection-FR: Exploring the Role of Native, Synthetic, and Translated Data in Biomedical Instruction Tuning |
| 2026-03-06 | Transparent AI for Mathematics: Transformer-Based Large Language Models for Mathematical Entity Relationship Extraction with XAI |
| 2026-03-06 | PVminerLLM: Structured Extraction of Patient Voice from Patient-Generated Text using Large Language Models |
| 2026-03-06 | VerChol -- Grammar-First Tokenization for Agglutinative Languages |
| 2026-03-06 | From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring |
| 2026-03-05 | NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution |
| 2026-03-05 | Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution |
| 2026-03-05 | PersianPunc: A Large-Scale Dataset and BERT-Based Approach for Persian Punctuation Restoration |
| 2026-03-05 | AILS-NTUA at SemEval-2026 Task 3: Efficient Dimensional Aspect-Based Sentiment Analysis |
| 2026-03-05 | NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension |
| 2026-03-04 | Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation |
| 2026-03-04 | Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA |
| 2026-03-04 | FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation |
| 2026-03-04 | Bielik-Q2-Sharp: A Comparative Study of Extreme 2-bit Quantization Methods for a Polish 11B Language Model |
| 2026-03-04 | A Neural Topic Method Using a Large-Language-Model-in-the-Loop for Business Research |
| 2026-03-03 | APRES: An Agentic Paper Revision and Evaluation System |
| 2026-03-03 | OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets |
| 2026-03-03 | A Browser-based Open Source Assistant for Multimodal Content Verification |
| 2026-03-03 | TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning |
| 2026-03-03 | ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs |
| 2026-03-02 | QIME: Constructing Interpretable Medical Text Embeddings via Ontology-Grounded Questions |
| 2026-03-02 | nchellwig at SemEval-2026 Task 3: Self-Consistent Structured Generation (SCSG) for Dimensional Aspect-Based Sentiment Analysis using Large Language Models |
| 2026-03-02 | From Variance to Invariance: Qualitative Content Analysis for Narrative Graph Annotation |
| 2026-03-02 | Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain |
| 2026-03-02 | EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training |
| 2026-03-01 | Hybrid Neural-LLM Pipeline for Morphological Glossing in Endangered Language Documentation: A Case Study of Jungar Tuvan |
| 2026-03-01 | VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling |
| 2026-03-01 | CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning |
| 2026-03-01 | DEP: A Decentralized Large Language Model Evaluation Protocol |
| 2026-03-01 | GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant |
| 2026-02-28 | RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesis |
| 2026-02-28 | A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs |
| 2026-02-28 | BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages |
| 2026-02-28 | QQ: A Toolkit for Language Identifiers and Metadata |
| 2026-02-28 | LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models |
| 2026-02-27 | Terminology Rarity Predicts Catastrophic Failure in LLM Translation of Low-Resource Ancient Languages: Evidence from Ancient Greek |
| 2026-02-27 | When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation |
| 2026-02-27 | CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing |
| 2026-02-27 | Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry |
| 2026-02-27 | GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models |
| 2026-02-26 | MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations |
| 2026-02-26 | AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors |
| 2026-02-26 | Towards Faithful Industrial RAG: A Reinforced Co-adaptation Framework for Advertising QA |
| 2026-02-26 | Cognitive Models and AI Algorithms Provide Templates for Designing Language Agents |
| 2026-02-26 | IDP Accelerator: Agentic Document Intelligence from Extraction to Compliance Validation |
| 2026-02-25 | Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion |
| 2026-02-25 | Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets |
| 2026-02-25 | Importance of Prompt Optimisation for Error Detection in Medical Notes Using Language Models |
| 2026-02-25 | Mitigating Structural Noise in Low-Resource S2TT: An Optimized Cascaded Nepali-English Pipeline with Punctuation Restoration |
| 2026-02-25 | Enhancing Multilingual Embeddings via Multi-Way Parallel Text Alignment |
| 2026-02-24 | MrBERT: Modern Multilingual Encoders via Vocabulary, Domain, and Dimensional Adaptation |
| 2026-02-24 | On Data Engineering for Scaling LLM Terminal Capabilities |
| 2026-02-24 | Enhancing Hate Speech Detection on Social Media: A Comparative Analysis of Machine Learning Models and Text Transformation Approaches |
| 2026-02-24 | PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A |
| 2026-02-24 | Beyond the Star Rating: A Scalable Framework for Aspect-Based Sentiment Analysis Using LLMs and Text Classification |
| 2026-02-23 | Natural Language Processing Models for Robust Document Categorization |
| 2026-02-23 | DEEP: Docker-based Execution and Evaluation Platform |
| 2026-02-23 | BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop |
| 2026-02-23 | Classroom Final Exam: An Instructor-Tested Reasoning Benchmark |
| 2026-02-23 | To Reason or Not to: Selective Chain-of-Thought in Medical Question Answering |
| 2026-02-22 | IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning |
| 2026-02-22 | How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders |
| 2026-02-22 | Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content |
| 2026-02-22 | TurkicNLP: An NLP Toolkit for Turkic Languages |
| 2026-02-22 | Uncovering Context Reliance in Unstructured Knowledge Editing |
| 2026-02-21 | BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models |
| 2026-02-21 | Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models |
| 2026-02-21 | From Trial by Fire To Sleep Like a Baby: A Lexicon of Anxiety Associations for 20k English Multiword Expressions |
| 2026-02-21 | ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models |
| 2026-02-21 | MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Elastic LLMs |
| 2026-02-20 | PolyFrame at MWE-2026 AdMIRe 2: When Words Are Not Enough: Multimodal Idiom Disambiguation |
| 2026-02-20 | Thinking by Subtraction: Confidence-Driven Contrastive Decoding for LLM Reasoning |
| 2026-02-20 | DP-RFT: Learning to Generate Synthetic Text via Differentially Private Reinforcement Fine-Tuning |
| 2026-02-20 | Simplifying Outcomes of Language Model Component Analyses with ELIA |
| 2026-02-20 | Click it or Leave it: Detecting and Spoiling Clickbait with Informativeness Measures and Large Language Models |
| 2026-02-19 | Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics |
| 2026-02-19 | Representation Collapse in Machine Translation Through the Lens of Angular Dispersion |
| 2026-02-19 | Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study |
| 2026-02-19 | CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts |
| 2026-02-19 | Auditing Reciprocal Sentiment Alignment: Inversion Risk, Dialect Representation and Intent Misalignment in Transformers |
| 2026-02-18 | The Validity of Coreference-based Evaluations of Natural Language Understanding |
| 2026-02-18 | BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization |
| 2026-02-18 | Training Models on Dialects of Translationese Shows How Lexical Diversity and Source-Target Syntactic Similarity Shape Learning |
| 2026-02-18 | MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models |
| 2026-02-18 | IndicEval: A Bilingual Indian Educational Evaluation Framework for Large Language Models |
| 2026-02-17 | Towards Expectation Detection in Language: A Case Study on Treatment Expectations in Reddit |
| 2026-02-17 | LuxMT Technical Report |
| 2026-02-17 | ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models |
| 2026-02-17 | *-PLUIE: Personalisable metric with Llm Used for Improved Evaluation |
| 2026-02-17 | DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting |
| 2026-02-16 | Unlocking Reasoning Capability on Machine Translation in Large Language Models |
| 2026-02-16 | How to Train Your Long-Context Visual Document Model |
| 2026-02-16 | Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation |
| 2026-02-16 | Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets |
| 2026-02-16 | Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography |
| 2026-01-12 | Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation |
| 2026-01-12 | Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset |
| 2026-01-12 | Order in the Evaluation Court: A Critical Analysis of NLG Evaluation Trends |
| 2026-01-12 | PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs |
| 2026-01-12 | From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP |
| 2026-01-11 | TurkBench: A Benchmark for Evaluating Turkish Large Language Models |
| 2026-01-11 | Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models |
| 2026-01-11 | UETQuintet at BioCreative IX -- MedHopQA: Enhancing Biomedical QA with Selective Multi-hop Reasoning and Contextual Retrieval |
| 2026-01-11 | MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues |
| 2026-01-11 | Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models |
| 2026-01-10 | AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages |
| 2026-01-10 | EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generation |
| 2026-01-10 | InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs |
| 2026-01-10 | Evaluating Accounting Reasoning Capabilities of Large Language Models |
| 2026-01-10 | BizFinBench.v2: A Unified Dual-Mode Bilingual Benchmark for Expert-Level Financial Capability Alignment |
| 2026-01-09 | A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality |
| 2026-01-09 | What do the metrics mean? A critical analysis of the use of Automated Evaluation Metrics in Interpreting |
| 2026-01-09 | CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning |
| 2026-01-09 | Pantagruel: Unified Self-Supervised Encoders for French Text and Speech |
| 2026-01-09 | How well can off-the-shelf LLMs elucidate molecular structures from mass spectra using chain-of-thought reasoning? |
| 2026-01-08 | Glitter: Visualizing Lexical Surprisal for Readability in Administrative Texts |
| 2026-01-08 | GenProve: Learning to Generate Text with Fine-Grained Provenance |
| 2026-01-08 | DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization |
| 2026-01-08 | Advancing Language Models for Code-related Tasks |
| 2026-01-08 | LANGSAE EDITING: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal |
| 2026-01-07 | EASLT: Emotion-Aware Sign Language Translation |
| 2026-01-07 | AI Generated Text Detection |
| 2026-01-07 | STELLA: Self-Reflective Terminology-Aware Framework for Building an Aerospace Information Retrieval Benchmark |
| 2026-01-07 | Analyzing and Improving Cross-lingual Knowledge Transfer for Machine Translation |
| 2026-01-07 | When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question-Answering |
| 2026-01-06 | Pearmut: Human Evaluation of Translation Made Trivial |
| 2026-01-06 | Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing |
| 2026-01-06 | Beyond the Black Box: Theory and Mechanism of Large Language Models |
| 2026-01-06 | Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning |
| 2026-01-06 | Automatic Prompt Engineering with No Task Cues and No Tuning |
| 2026-01-05 | pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs |
| 2026-01-05 | Scalable Construction of a Lung Cancer Knowledge Base: Profiling Semantic Reasoning in LLMs |
| 2026-01-05 | DeCode: Decoupling Content and Delivery for Medical QA |
| 2026-01-05 | Estimating Text Temperature |
| 2026-01-05 | DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs |
| 2026-01-02 | Exploring the Performance of Large Language Models on Subjective Span Identification Tasks |
| 2026-01-02 | Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends |
| 2026-01-02 | Physio-DPO: Aligning Large Language Models with the Protein Energy Landscape to Eliminate Structural Hallucinations |
| 2026-01-02 | Memory Bank Compression for Continual Adaptation of Large Language Models |
| 2026-01-02 | Sigmoid Head for Quality Estimation under Language Ambiguity |
| 2026-01-01 | Talk Less, Verify More: Improving LLM Assistants with Semantic Checks and Execution Feedback |
| 2026-01-01 | Comparative Efficiency Analysis of Lightweight Transformer Models: A Multi-Domain Empirical Benchmark for Enterprise NLP Deployment |
| 2026-01-01 | Noise-Aware Named Entity Recognition for Historical VET Documents |
| 2026-01-01 | The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining |
| 2026-01-01 | JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation |
| 2025-12-31 | Big AI is accelerating the metacrisis: What can we do? |
| 2025-12-31 | HaluNet: Multi-Granular Uncertainty Modeling for Efficient Hallucination Detection in LLM Question Answering |
| 2025-12-31 | MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models |
| 2025-12-31 | CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement |
| 2025-12-31 | MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes |
| 2025-12-30 | HY-MT1.5 Technical Report |
| 2025-12-30 | Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech |
| 2025-12-30 | AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives |
| 2025-12-30 | Factorized Learning for Temporally Grounded Video-Language Models |
| 2025-12-30 | QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs |
| 2025-12-29 | Not too long do read: Evaluating LLM-generated extreme scientific summaries |
| 2025-12-29 | MiMo-Audio: Audio Language Models are Few-Shot Learners |
| 2025-12-29 | AI4Reading: Chinese Audiobook Interpretation System Based on Multi-Agent Collaboration |
| 2025-12-29 | Training AI Co-Scientists Using Rubric Rewards |
| 2025-12-29 | Chinese Morph Resolution in E-commerce Live Streaming Scenarios |
| 2025-12-28 | Fake News Classification in Urdu: A Domain Adaptation Approach for a Low-Resource Language |
| 2025-12-28 | CNSight: Evaluation of Clinical Note Segmentation Tools |
| 2025-12-28 | Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis |
| 2025-12-28 | A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms |
| 2025-12-28 | TabiBERT: A Large-Scale ModernBERT Foundation Model and Unified Benchmarking Framework for Turkish |
| 2025-12-27 | Chain-of-thought Reviewing and Correction for Time Series Question Answering |
| 2025-12-27 | Hallucination Detection and Evaluation of Large Language Model |
| 2025-12-27 | Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages |
| 2025-12-27 | M2G-Eval: Enhancing and Evaluating Multi-granularity Multilingual Code Generation |
| 2025-12-27 | ManchuTTS: Towards High-Quality Manchu Speech Synthesis via Flow Matching and Hierarchical Text Representation |
| 2025-12-26 | Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis |
| 2025-12-26 | AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts |
| 2025-12-26 | Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management |
| 2025-12-26 | TimeBill: Time-Budgeted Inference for Large Language Models |
| 2025-12-26 | SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence |
| 2025-12-25 | Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation |
| 2025-12-25 | Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers |
| 2025-12-25 | Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures |
| 2025-12-25 | Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM |
| 2025-12-25 | A Unified Definition of Hallucination, Or: It's the World Model, Stupid |
| 2025-12-24 | Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy |
| 2025-12-24 | Neural Probe-Based Hallucination Detection for Large Language Models |
| 2025-12-24 | MultiMind at SemEval-2025 Task 7: Crosslingual Fact-Checked Claim Retrieval via Multi-Source Alignment |
| 2025-12-24 | ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models |
| 2025-12-24 | Automatic Replication of LLM Mistakes in Medical Conversations |
| 2025-12-23 | Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings |
| 2025-12-23 | TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior |
| 2025-12-23 | Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning |
| 2025-12-23 | Can LLMs Solve My Grandma's Riddle? Evaluating Multilingual Large Language Models on Reasoning Traditional Bangla Tricky Riddles |
| 2025-12-23 | MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs |
| 2025-12-22 | From Speech to Subtitles: Evaluating ASR Models in Subtitling Italian Television Programs |
| 2025-12-22 | SiamGPT: Quality-First Fine-Tuning for Stable Thai Text Generation |
| 2025-12-22 | CodeSimpleQA: Scaling Factuality in Code Large Language Models |
| 2025-12-22 | Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding |
| 2025-12-22 | Algerian Dialect |
| 2025-12-21 | From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation |
| 2025-12-21 | Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations |
| 2025-12-21 | On Finding Inconsistencies in Documents |
| 2025-12-21 | LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction |
| 2025-12-21 | Application of deep learning approaches for medieval historical documents transcription |
| 2025-12-20 | LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators |
| 2025-12-20 | DACE For Railway Acronym Disambiguation |
| 2025-12-20 | SRS-Stories: Vocabulary-constrained multilingual story generation for language learning |
| 2025-12-20 | Generalization Gaps in Political Fake News Detection: An Empirical Study on the LIAR Dataset |
| 2025-12-20 | InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning |
| 2025-12-19 | Subjective Question Generation and Answer Evaluation using NLP |
| 2025-12-19 | When the Gold Standard isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content |
| 2025-12-19 | Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models |
| 2025-12-19 | AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora |
| 2025-12-19 | Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers |
| 2025-12-18 | Hacking Neural Evaluation Metrics with Single Hub Text |
| 2025-12-18 | Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics |
| 2025-12-18 | Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs |
| 2025-12-18 | Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs |
| 2025-12-18 | UM_FHS at the CLEF 2025 SimpleText Track: Comparing No-Context and Fine-Tune Approaches for GPT-4.1 Models in Sentence and Document-Level Text Simplification |
| 2025-12-17 | From NLG Evaluation to Modern Student Assessment in the Era of ChatGPT: The Great Misalignment Problem and Pedagogical Multi-Factor Assessment (P-MFA) |
| 2025-12-17 | An Empirical Study on Chinese Character Decomposition in Multiword Expression-Aware Neural Machine Translation |
| 2025-12-17 | Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models |
| 2025-12-17 | Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams |
| 2025-12-17 | Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024 |
| 2025-12-16 | A Comparative Analysis of Retrieval-Augmented Generation Techniques for Bengali Standard-to-Dialect Machine Translation Using LLMs |
| 2025-12-16 | JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction |
| 2025-12-16 | TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs |
| 2025-12-16 | Linguists should learn to love speech-based deep learning models |
| 2025-12-16 | Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies |
| 2025-12-15 | Building from Scratch: A Multi-Agent Framework with Human-in-the-Loop for Multilingual Legal Terminology Mapping |
| 2025-12-15 | FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models |
| 2025-12-15 | Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing |
| 2025-12-15 | An Open and Reproducible Deep Research Agent for Long-Form Question Answering |
| 2025-12-15 | PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation |
| 2025-12-14 | HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks |
| 2025-12-14 | What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation |
| 2025-12-14 | Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space |
| 2025-12-14 | Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining |
| 2025-12-14 | NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents |
| 2025-12-13 | Large language models have learned to use language |
| 2025-12-13 | Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics |
| 2025-12-13 | Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking |
| 2025-12-13 | Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors |
| 2025-12-13 | F5-TTS-RO: Extending F5-TTS to Romanian TTS via Lightweight Input Adaptation |
| 2025-12-12 | Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis |
| 2025-12-12 | Task-Specific Sparse Feature Masks for Molecular Toxicity Prediction with Chemical Language Models |
| 2025-12-12 | Multi-Intent Spoken Language Understanding: Methods, Trends, and Challenges |
| 2025-12-12 | Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling |
| 2025-12-12 | DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry |
| 2025-12-11 | MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data |
| 2025-12-11 | Multilingual VLM Training: Adapting an English-Trained VLM to French |
| 2025-12-11 | Enhancing Next-Generation Language Models with Knowledge Graphs: Extending Claude, Mistral IA, and GPT-4 via KG-BERT |
| 2025-12-11 | Applying NLP to iMessages: Understanding Topic Avoidance, Responsiveness, and Sentiment |
| 2025-12-11 | XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs |
| 2025-12-10 | LLMs in Interpreting Legal Documents |
| 2025-12-10 | Efficient Continual Learning in Neural Machine Translation: A Low-Rank Adaptation Approach |
| 2025-12-10 | Neurosymbolic Information Extraction from Transactional Documents |
| 2025-12-10 | Language models as tools for investigating the distinction between possible and impossible natural languages |
| 2025-12-10 | MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment |
| 2025-12-09 | What Triggers my Model? Contrastive Explanations Inform Gender Choices by Translation Models |
| 2025-12-09 | HealthcareNLP: where are we and what is next? |
| 2025-12-09 | Automatic Essay Scoring and Feedback Generation in Basque Language Learning |
| 2025-12-09 | Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages |
| 2025-12-09 | MindShift: Analyzing Language Models' Reactions to Psychological Prompts |
| 2025-12-08 | TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation |
| 2025-12-08 | HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs |
| 2025-12-08 | SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents |
| 2025-12-08 | Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models |
| 2025-12-08 | Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation |
| 2025-12-07 | MATEX: A Multi-Agent Framework for Explaining Ethereum Transactions |
| 2025-12-07 | Large Language Models and Forensic Linguistics: Navigating Opportunities and Threats in the Age of Generative AI |
| 2025-12-07 | CMV-Fuse: Cross Modal-View Fusion of AMR, Syntax, and Knowledge Representations for Aspect Based Sentiment Analysis |
| 2025-12-07 | Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models |
| 2025-12-07 | Progress Ratio Embeddings: An Impatience Signal for Robust Length Control in Neural Text Generation |
| 2025-12-06 | Adapting AlignScore Mertic for Factual Consistency Evaluation of Text in Russian: A Student Abstract |
| 2025-12-06 | Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models |
| 2025-12-06 | Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety |
| 2025-12-06 | Less Is More for Multi-Step Logical Reasoning of LLM Generalisation Under Rule Removal, Paraphrasing, and Compression |
| 2025-12-06 | Classifying German Language Proficiency Levels Using Large Language Models |
| 2025-12-05 | LMSpell: Neural Spell Checking for Low-Resource Languages |
| 2025-12-05 | Grounded Multilingual Medical Reasoning for Question Answering with Large Language Models |
| 2025-12-05 | To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis |
| 2025-12-05 | Empathy by Design: Aligning Large Language Models for Healthcare Dialogue |
| 2025-12-05 | Learning from Self Critique and Refinement for Faithful LLM Summarization |
| 2025-12-04 | Structured Document Translation via Format Reinforcement Learning |
| 2025-12-04 | MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation |
| 2025-12-04 | AdiBhashaa: A Community-Curated Benchmark for Machine Translation into Indian Tribal Languages |
| 2025-12-04 | Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective |
| 2025-12-04 | EvoEdit: Lifelong Free-Text Knowledge Editing through Latent Perturbation Augmentation and Knowledge-driven Parameter Fusion |
| 2025-12-02 | Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic |
| 2025-12-02 | A Concise Review of Hallucinations in LLMs and their Mitigation |
| 2025-12-02 | BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion |
| 2025-12-02 | PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models |
| 2025-12-02 | Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies |
| 2025-12-01 | MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages |
| 2025-12-01 | MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark |
| 2025-12-01 | BHRAM-IL: A Benchmark for Hallucination Recognition and Assessment in Multiple Indian Languages |
| 2025-12-01 | Conveying Imagistic Thinking in Traditional Chinese Medicine Translation: A Prompt Engineering and LLM-Based Evaluation Framework |
| 2025-12-01 | Agreement-Constrained Probabilistic Minimum Bayes Risk Decoding |
| 2025-11-30 | Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study |
| 2025-11-30 | WaterSearch: A Quality-Aware Search-based Watermarking Framework for Large Language Models |
| 2025-11-30 | Advancing Academic Chatbots: Evaluation of Non Traditional Outputs |
| 2025-11-30 | Accelerating Bangla NLP Tasks with Automatic Mixed Precision: Resource-Efficient Training Preserving Model Efficacy |
| 2025-11-30 | How do we measure privacy in text? A survey of text anonymization metrics |
| 2025-11-29 | A Taxonomy of Errors in English as she is spoke: Toward an AI-Based Method of Error Analysis for EFL Writing Instruction |
| 2025-11-29 | Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets |
| 2025-11-29 | Developing a Comprehensive Framework for Sentiment Analysis in Turkish |
| 2025-11-29 | Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs |
| 2025-11-29 | CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning |
| 2025-11-28 | Conveying Imagistic Thinking in TCM Translation: A Prompt Engineering and LLM-Based Evaluation Framework |
| 2025-11-28 | OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion |
| 2025-11-28 | BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning |
| 2025-11-28 | FEANEL: A Benchmark for Fine-Grained Error Analysis in K-12 English Writing |
| 2025-11-28 | Tourism Question Answer System in Indian Language using Domain-Adapted Foundation Models |
| 2025-11-27 | Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter? |
| 2025-11-27 | ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering |
| 2025-11-27 | Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration |
| 2025-11-27 | Sentiment Analysis Of Shopee Product Reviews Using Distilbert |
| 2025-11-27 | From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation |
| 2025-11-26 | Bangla Sign Language Translation: Dataset Creation Challenges, Benchmarking and Prospects |
| 2025-11-26 | FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers |
| 2025-11-26 | Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text |
| 2025-11-26 | Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices |
| 2025-11-26 | RoParQ: Paraphrase-Aware Alignment of Large Language Models Towards Robustness to Paraphrased Questions |
| 2025-11-25 | CounterVQA: Evaluating and Improving Counterfactual Reasoning in Vision-Language Models for Video Understanding |
| 2025-11-25 | Breaking Bad: Norms for Valence, Arousal, and Dominance for over 10k English Multiword Expressions |
| 2025-11-25 | Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach |
| 2025-11-25 | Length-MAX Tokenizer for Language Models |
| 2025-11-25 | Geometry of Decision Making in Language Models |
| 2025-11-24 | A symbolic Perl algorithm for the unification of Nahuatl word spellings |
| 2025-11-24 | Skeletons Matter: Dynamic Data Augmentation for Text-to-Query |
| 2025-11-24 | Generating Reading Comprehension Exercises with Large Language Models for Educational Applications |
| 2025-11-24 | A Reproducible Framework for Neural Topic Modeling in Focus Group Analysis |
| 2025-11-24 | Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search |
| 2025-11-23 | SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data |
| 2025-11-23 | "AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa |
| 2025-11-23 | MindEval: Benchmarking Language Models on Multi-turn Mental Health Support |
| 2025-11-23 | From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence |
| 2025-11-23 | Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models |
| 2025-11-21 | LangMark: A Multilingual Dataset for Automatic Post-Editing |
| 2025-11-21 | Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables |
| 2025-11-21 | Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT |
| 2025-11-21 | MUCH: A Multilingual Claim Hallucination Benchmark |
| 2025-11-21 | PUCP-Metrix: A Comprehensive Open-Source Repository of Linguistic Metrics for Spanish |
| 2025-11-20 | AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser |
| 2025-11-20 | OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe |
| 2025-11-20 | Classification of worldwide news articles by perceived quality, 2018-2024 |
| 2025-11-20 | Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation |
| 2025-11-20 | Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions |
| 2025-11-19 | Evaluating Multimodal Large Language Models on Vertically Written Japanese Text |
| 2025-11-19 | Mathematical Analysis of Hallucination Dynamics in Large Language Models: Uncertainty Quantification, Advanced Decoding, and Principled Mitigation |
| 2025-11-19 | MAPROC at AHaSIS Shared Task: Few-Shot and Sentence Transformer for Sentiment Analysis of Arabic Hotel Reviews |
| 2025-11-19 | Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language |
| 2025-11-19 | HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples |
| 2025-11-18 | Subword Tokenization Strategies for Kurdish Word Embeddings |
| 2025-11-18 | MuCPT: Music-related Natural Language Model Continued Pretraining |
| 2025-11-18 | The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models |
| 2025-11-18 | SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction |
| 2025-11-18 | Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak |
| 2025-11-17 | Can QE-informed (Re)Translation lead to Error Correction? |
| 2025-11-17 | Non-Linear Scoring Model for Translation Quality Evaluation |
| 2025-11-17 | Translation Entropy: A Statistical Framework for Evaluating Translation Systems |
| 2025-11-17 | Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study |
| 2025-11-17 | From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models |
| 2025-11-16 | QA-Noun: Representing Nominal Semantics via Natural Language Question-Answer Pairs |
| 2025-11-16 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models |
| 2025-11-16 | Evaluating Autoformalization Robustness via Semantically Similar Paraphrasing |
| 2025-11-16 | Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing |
| 2025-11-16 | Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing |
| 2025-11-15 | Exploring Parameter-Efficient Fine-Tuning and Backtranslation for the WMT 25 General Translation Task |
| 2025-11-15 | Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations |
| 2025-11-15 | A Reasoning Paradigm for Named Entity Recognition |
| 2025-11-15 | MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues |
| 2025-11-15 | AugAbEx : Way Forward for Extractive Case Summarization |
| 2025-11-14 | DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains |
| 2025-11-14 | Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion |
| 2025-11-14 | Context-Emotion Aware Therapeutic Dialogue Generation: A Multi-component Reinforcement Learning Approach to Language Models for Mental Health Support |
| 2025-11-14 | M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text |
| 2025-11-14 | Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB |
| 2025-11-13 | NumPert: Numerical Perturbations to Probe Language Models for Veracity Prediction |
| 2025-11-13 | TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain |
| 2025-11-13 | Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning |
| 2025-11-13 | Faithful Summarization of Consumer Health Queries: A Cross-Lingual Framework with LLMs |
| 2025-11-13 | BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages |
| 2025-11-12 | MTQ-Eval: Multilingual Text Quality Evaluation for Language Models |
| 2025-11-12 | mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models |
| 2025-11-12 | A Neurosymbolic Approach to Natural Language Formalization and Verification |
| 2025-11-12 | How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation |
| 2025-11-12 | MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique |
| 2025-11-06 | RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods |
| 2025-11-06 | Dynamic Jointly Batch Selection for Data Efficient Machine Translation Fine-Tuning |
| 2025-11-06 | T-FIX: Text-Based Explanations with Features Interpretable to eXperts |
| 2025-11-06 | VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks |
| 2025-11-06 | MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation |
| 2025-11-05 | ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation |
| 2025-11-05 | How to Evaluate Speech Translation with Source-Aware Neural MT Metrics |
| 2025-11-05 | Segmentation Beyond Defaults: Asymmetrical Byte Pair Encoding for Optimal Machine Translation Performance |
| 2025-11-05 | BanglaSTEM: A Parallel Corpus for Technical Domain Bangla-English Translation |
| 2025-11-05 | Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG |
| 2025-11-04 | The Analysis of Lexical Errors in Machine Translation from English into Romanian |
| 2025-11-04 | Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model |
| 2025-11-04 | PragExTra: A Multilingual Corpus of Pragmatic Explicitation in Translation |
| 2025-11-04 | Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results |
| 2025-11-04 | Targeted Error Correction in Knowledge Distillation: Small Language Models Surpass GPT |
| 2025-11-03 | Imperfect Language, Artificial Intelligence, and the Human Mind: An Interdisciplinary Approach to Linguistic Errors in Native Spanish Speakers |
| 2025-11-03 | DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection |
| 2025-11-03 | $\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles |
| 2025-11-03 | BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification |
| 2025-11-03 | EngChain: A Symbolic Benchmark for Verifiable Multi-Step Reasoning in Engineering |
| 2025-11-02 | HPLT 3.0: Very Large-Scale Multilingual Resources for LLM and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models |
| 2025-11-02 | Advancing Machine-Generated Text Detection from an Easy to Hard Supervision Perspective |
| 2025-11-02 | Building a Silver-Standard Dataset from NICE Guidelines for Clinical LLMs |
| 2025-11-02 | The Biased Oracle: Assessing LLMs' Understandability and Empathy in Medical Diagnoses |
| 2025-11-02 | ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval |
| 2025-11-01 | Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus |
| 2025-11-01 | PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks |
| 2025-11-01 | MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts |
| 2025-11-01 | LingGym: How Far Are LLMs from Thinking Like Field Linguists? |
| 2025-11-01 | Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction |
| 2025-10-31 | TransAlign: Machine Translation Encoders are Strong Word Aligners, Too |
| 2025-10-31 | From the Rock Floor to the Cloud: A Systematic Survey of State-of-the-Art NLP in Battery Life Cycle |
| 2025-10-31 | Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap |
| 2025-10-31 | MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models |
| 2025-10-31 | POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation |
| 2025-10-30 | SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level |
| 2025-10-30 | Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark |
| 2025-10-30 | Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations |
| 2025-10-30 | Bayesian Network Fusion of Large Language Models for Sentiment Analysis |
| 2025-10-30 | Quantitative Intertextuality from the Digital Humanities Perspective: A Survey |
| 2025-10-29 | A Critical Study of Automatic Evaluation in Sign Language Translation |
| 2025-10-29 | A Survey on Efficient Large Language Model Training: From Data-centric Perspectives |
| 2025-10-29 | RLMEval: Evaluating Research-Level Neural Theorem Proving |
| 2025-10-29 | Fine-Tuned Language Models for Domain-Specific Summarization and Tagging |
| 2025-10-29 | Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation |
| 2025-10-28 | MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation |
| 2025-10-28 | Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation |
| 2025-10-28 | MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task |
| 2025-10-28 | Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices |
| 2025-10-28 | Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants |
| 2025-10-27 | A U-Net and Transformer Pipeline for Multilingual Image Translation |
| 2025-10-27 | Quality-Aware Translation Tagging in Multilingual RAG system |
| 2025-10-27 | M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark |
| 2025-10-27 | M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset |
| 2025-10-27 | AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages |
| 2025-10-26 | VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions |
| 2025-10-26 | Integrating Linguistics and AI: Morphological Analysis and Corpus development of Endangered Toto Language of West Bengal |
| 2025-10-26 | Iterative Layer Pruning for Efficient Translation Inference |
| 2025-10-26 | Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models |
| 2025-10-26 | A Comprehensive Dataset for Human vs. AI Generated Text Detection |
| 2025-10-25 | Multilingual Target-Stance Extraction |
| 2025-10-25 | Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection |
| 2025-10-25 | Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs |
| 2025-10-25 | Irony Detection in Urdu Text: A Comparative Study Using Machine Learning Models and Large Language Models |
| 2025-10-25 | WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models |
| 2025-10-24 | Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics |
| 2025-10-24 | Typoglycemia under the Hood: Investigating Language Models' Understanding of Scrambled Words |
| 2025-10-24 | Document Understanding, Measurement, and Manipulation Using Category Theory |
| 2025-10-24 | Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering |
| 2025-10-24 | Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech Models |
| 2025-10-23 | Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost |
| 2025-10-23 | Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models |
| 2025-10-23 | VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation |
| 2025-10-23 | Framework for Machine Evaluation of Reasoning Completeness in Large Language Models For Classification Tasks |
| 2025-10-23 | \textsc{CantoNLU}: A benchmark for Cantonese natural language understanding |
| 2025-10-21 | BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks |
| 2025-10-21 | IMB: An Italian Medical Benchmark for Question Answering |
| 2025-10-21 | DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code |
| 2025-10-21 | SemiAdapt and SemiLoRA: Efficient Domain Adaptation for Transformer-based Low-Resource Language Translation with a Case Study on Irish |
| 2025-10-21 | MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training |
| 2025-10-20 | Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations |
| 2025-10-20 | Evaluating Large Language Models on Urdu Idiom Translation |
| 2025-10-20 | Chain-of-Thought Reasoning Improves Context-Aware Translation with Large Language Models |
| 2025-10-20 | Transformer-Based Low-Resource Language Translation: A Study on Standard Bengali to Sylheti |
| 2025-10-20 | Lingua Custodi's participation at the WMT 2025 Terminology shared task |
| 2025-10-19 | LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding |
| 2025-10-19 | ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models |
| 2025-10-19 | Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation |
| 2025-10-19 | DiscoTrack: A Multilingual LLM Benchmark for Discourse Tracking |
| 2025-10-19 | Interpretability Framework for LLMs in Undergraduate Calculus |
| 2025-10-18 | Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review |
| 2025-10-18 | Fine-tuning of Large Language Models for Constituency Parsing Using a Sequence to Sequence Approach |
| 2025-10-18 | AI-Generated Text Detection in Low-Resource Languages: A Case Study on Urdu |
| 2025-10-18 | ReviewGuard: Enhancing Deficient Peer Review Detection via LLM-Driven Data Augmentation |
| 2025-10-18 | Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models |
| 2025-10-17 | On Non-interactive Evaluation of Animal Communication Translators |
| 2025-10-17 | BiMax: Bidirectional MaxSim Score for Document-Level Alignment |
| 2025-10-17 | Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs |
| 2025-10-17 | HypoSpace: Evaluating LLM Creativity as Set-Valued Hypothesis Generators under Underdetermination |
| 2025-10-17 | Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing |
| 2025-10-16 | Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks |
| 2025-10-16 | MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking |
| 2025-10-16 | From Binary to Bilingual: How the National Weather Service is Using Artificial Intelligence to Develop a Comprehensive Translation Program |
| 2025-10-16 | Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models |
| 2025-10-16 | Your Next Token Prediction: A Multilingual Benchmark for Personalized Response Generation |
| 2025-10-15 | Beyond Single-Reward: Multi-Pair, Multi-Perspective Preference Optimization for Machine Translation |
| 2025-10-15 | MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models |
| 2025-10-15 | LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA |
| 2025-10-15 | Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps |
| 2025-10-15 | Document Intelligence in the Era of Large Language Models: A Survey |
| 2025-10-14 | DPO-Tuned Large Language Models for Segmentation in Simultaneous Speech Translation |
| 2025-10-14 | Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions |
| 2025-10-14 | When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection |
| 2025-10-14 | ACADATA: Parallel Dataset of Academic Data for Machine Translation |
| 2025-10-14 | A Survey on Parallel Reasoning |
| 2025-10-13 | Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering |
| 2025-10-13 | End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF: A Reproducibility Study |
| 2025-10-13 | LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens |
| 2025-10-13 | Automating Structural Engineering Workflows with Large Language Model Agents |
| 2025-10-13 | Investigating Large Language Models' Linguistic Abilities for Text Preprocessing |
| 2025-10-12 | Happiness is Sharing a Vocabulary: A Study of Transliteration Methods |
| 2025-10-12 | Bhasha-Rupantarika: Algorithm-Hardware Co-design approach for Multilingual Neural Machine Translation |
| 2025-10-12 | STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models |
| 2025-10-12 | Sarcasm Detection Using Deep Convolutional Neural Networks: A Modular Deep Learning Framework |
| 2025-10-12 | Assessing Large Language Models for Structured Medical Order Extraction |
| 2025-10-11 | Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations |
| 2025-10-11 | MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction |
| 2025-10-11 | You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs |
| 2025-10-11 | End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs |
| 2025-10-11 | HUME: Measuring the Human-Model Performance Gap in Text Embedding Task |
| 2025-10-10 | Quality Estimation Reranking for Document-Level Translation |
| 2025-10-10 | DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation |
| 2025-10-10 | ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability |
| 2025-10-10 | ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering |
| 2025-10-10 | A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages |
| 2025-10-09 | ChatGPT as a Translation Engine: A Case Study on Japanese-English |
| 2025-10-09 | Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains |
| 2025-10-09 | AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment |
| 2025-10-09 | A$^2$Search: Ambiguity-Aware Question Answering with Reinforcement Learning |
| 2025-10-09 | Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation |
| 2025-10-07 | The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP |
| 2025-10-07 | Test-Time Scaling of Reasoning Models for Machine Translation |
| 2025-10-07 | Type and Complexity Signals in Multilingual Question Representations |
| 2025-10-07 | FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering |
| 2025-10-07 | MathRobust-LV: Evaluation of Large Language Models' Robustness to Linguistic Variations in Mathematical Reasoning |
| 2025-10-06 | COLE: a Comprehensive Benchmark for French Language Understanding Evaluation |
| 2025-10-06 | Aligning Language Models with Clinical Expertise: DPO for Heart Failure Nursing Documentation in Critical Care |
| 2025-10-06 | Learning to Interpret Weight Differences in Language Models |
| 2025-10-06 | A Set of Quebec-French Corpus of Regional Expressions and Terms |
| 2025-10-06 | When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA |
| 2025-10-05 | Time Is Effort: Estimating Human Post-Editing Time for Grammar Error Correction Tool Evaluation |
| 2025-10-05 | Large Language Models Hallucination: A Comprehensive Survey |
| 2025-10-05 | Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought |
| 2025-10-05 | Evaluation of Clinical Trials Reporting Quality using Large Language Models |
| 2025-10-05 | LongTail-Swap: benchmarking language models' abilities on rare words |
| 2025-10-03 | Self-Improvement in Multimodal Large Language Models: A Survey |
| 2025-10-03 | EditLens: Quantifying the Extent of AI Editing in Text |
| 2025-10-03 | Cache-to-Cache: Direct Semantic Communication Between Large Language Models |
| 2025-10-03 | Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer |
| 2025-10-03 | Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation |
| 2025-10-02 | Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention |
| 2025-10-02 | F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data |
| 2025-10-02 | Detecting LLM-Generated Spam Reviews by Integrating Language Model Embeddings and Graph Neural Network |
| 2025-10-02 | RAG-BioQA Retrieval-Augmented Generation for Long-Form Biomedical Question Answering |
| 2025-10-02 | Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction |
| 2025-10-01 | MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation |
| 2025-10-01 | Research on the Integration of Embodied Intelligence and Reinforcement Learning in Textual Domains |
| 2025-10-01 | mR3: Multilingual Rubric-Agnostic Reward Reasoning Models |
| 2025-10-01 | A-VERT: Agnostic Verification with Embedding Ranking Targets |
| 2025-10-01 | Copy-Paste to Mitigate Large Language Model Hallucinations |
| 2025-09-30 | PerQ: Efficient Evaluation of Multilingual Text Personalization Quality |
| 2025-09-30 | CEAID: Benchmark of Multilingual Machine-Generated Text Detection Methods for Central European Languages |
| 2025-09-30 | Generating Difficult-to-Translate Texts |
| 2025-09-30 | TASER: Translation Assessment via Systematic Evaluation and Reasoning |
| 2025-09-30 | Explaining novel senses using definition generation with open language models |
| 2025-09-29 | MMRQA: Signal-Enhanced Multimodal Large Language Models for MRI Quality Assessment |
| 2025-09-29 | Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey |
| 2025-09-29 | Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation |
| 2025-09-29 | Alternatives To Next Token Prediction In Text Generation -- A Survey |
| 2025-09-29 | Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs |
| 2025-09-28 | The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact |
| 2025-09-28 | Aligning LLMs for Multilingual Consistency in Enterprise Applications |
| 2025-09-28 | Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment |
| 2025-09-28 | EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos |
| 2025-09-28 | Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm |
| 2025-09-27 | How to Make Large Language Models Generate 100% Valid Molecules? |
| 2025-09-27 | Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation |
| 2025-09-27 | MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction |
| 2025-09-27 | Comparison of Scoring Rationales Between Large Language Models and Human Raters |
| 2025-09-27 | From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis |
| 2025-09-26 | JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA |
| 2025-09-26 | KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues |
| 2025-09-26 | Mixture of Detectors: A Compact View of Machine-Generated Text Detection |
| 2025-09-26 | The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems |
| 2025-09-26 | Multilingual Vision-Language Models, A Survey |
| 2025-09-25 | PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints |
| 2025-09-25 | "Be My Cheese?": Assessing Cultural Nuance in Multilingual LLM Translations |
| 2025-09-25 | Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning |
| 2025-09-25 | Towards Transparent AI: A Survey on Explainable Language Models |
| 2025-09-25 | BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback |
| 2025-09-24 | Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks |
| 2025-09-24 | CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems |
| 2025-09-24 | EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation |
| 2025-09-24 | SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages |
| 2025-09-24 | Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation |
| 2025-09-23 | Evaluating Language Translation Models by Playing Telephone |
| 2025-09-23 | Investigating Test-Time Scaling with Reranking for Machine Translation |
| 2025-09-23 | Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector |
| 2025-09-23 | TsqLoRA: Towards Sensitivity and Quality Low-Rank Adaptation for Efficient Fine-Tuning |
| 2025-09-23 | DTW-Align: Bridging the Modality Gap in End-to-End Speech Translation with Dynamic Time Warping Alignment |
| 2025-09-22 | Crosslingual Optimized Metric for Translation Assessment of Indian Languages |
| 2025-09-22 | Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text |
| 2025-09-22 | Specification-Aware Machine Translation and Evaluation for Purpose Alignment |
| 2025-09-22 | Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages |
| 2025-09-22 | MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents |
| 2025-09-21 | Extending Automatic Machine Translation Evaluation to Book-Length Documents |
| 2025-09-21 | AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation |
| 2025-09-21 | CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages |
| 2025-09-21 | Probabilistic Token Alignment for Large Language Model Fusion |
| 2025-09-21 | TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies? |
| 2025-09-20 | Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle |
| 2025-09-20 | ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions |
| 2025-09-20 | Angular Dispersion Accelerates $k$-Nearest Neighbors Machine Translation |
| 2025-09-20 | Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data |
| 2025-09-20 | Time to Revist Exact Match |
| 2025-09-19 | A method for improving multilingual quality and diversity of instruction fine-tuning datasets |
| 2025-09-19 | Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation |
| 2025-09-19 | UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations |
| 2025-09-19 | Whisper-UT: A Unified Translation Framework for Speech and Text |
| 2025-09-19 | Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems |
| 2025-09-18 | Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages |
| 2025-09-18 | Real, Fake, or Manipulated? Detecting Machine-Influenced Text |
| 2025-09-18 | CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models |
| 2025-09-18 | Frustratingly Easy Data Augmentation for Low-Resource ASR |
| 2025-09-18 | Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction |
| 2025-09-17 | Long-context Reference-based MT Quality Estimation |
| 2025-09-17 | Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality |
| 2025-09-17 | Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification |
| 2025-09-17 | Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST |
| 2025-09-17 | You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models |
| 2025-09-16 | Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents |
| 2025-09-16 | Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviews |
| 2025-09-16 | HistoryBankQA: Multilingual Temporal Question Answering on Historical Events |
| 2025-09-16 | Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO |
| 2025-09-16 | Multi-Model Synthetic Training for Mission-Critical Small Language Models |
| 2025-09-15 | A comparison of pipelines for the translation of a low resource language based on transformers |
| 2025-09-15 | MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering |
| 2025-09-15 | Preservation of Language Understanding Capabilities in Speech-aware Large Language Models |
| 2025-09-15 | Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges |
| 2025-09-15 | DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models |
| 2025-09-14 | Improving LLMs' Learning for Coreference Resolution |
| 2025-09-14 | Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity |
| 2025-09-14 | !MSA at AraHealthQA 2025 Shared Task: Enhancing LLM Performance for Arabic Clinical Question Answering through Prompt Engineering and Ensemble Learning |
| 2025-09-14 | RanAT4BIE: Random Adversarial Training for Biomedical Information Extraction |
| 2025-09-14 | FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs |
| 2025-09-13 | An Interpretable Benchmark for Clickbait Detection and Tactic Attribution |
| 2025-09-13 | CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis |
| 2025-09-13 | ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER |
| 2025-09-13 | Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents |
| 2025-09-13 | Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding |
| 2025-09-12 | Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models |
| 2025-09-12 | Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient Records |
| 2025-09-12 | Arabic Large Language Models for Medical Text Generation |
| 2025-09-12 | Large Language Models Meet Legal Artificial Intelligence: A Survey |
| 2025-09-12 | SearchInstruct: Enhancing Domain Adaptation via Retrieval-Based Instruction Dataset Creation |
| 2025-09-11 | Mitigating Language Barriers in Education: Developing Multilingual Digital Learning Materials with Machine Translation |
| 2025-09-11 | GmSLM : Generative Marmoset Spoken Language Modeling |
| 2025-09-11 | Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization |
| 2025-09-11 | Agentic LLMs for Question Answering over Tabular Data |
| 2025-09-11 | MR-UIE: Multi-Perspective Reasoning with Reinforcement Learning for Universal Information Extraction |
| 2025-09-10 | CM-Align: Consistency-based Multilingual Alignment for Large Language Models |
| 2025-09-10 | Toward Subtrait-Level Model Explainability in Automated Writing Evaluation |
| 2025-09-10 | MoVoC: Morphology-Aware Subword Construction for Geez Script Languages |
| 2025-09-10 | Automatic Detection of Inauthentic Templated Responses in English Language Assessments |
| 2025-09-10 | Natural Language Translation of Formal Proofs through Informalization of Proof Steps and Recursive Summarization along Proof Structure |
| 2025-09-09 | From Scarcity to Efficiency: Investigating the Effects of Data Augmentation on African Machine Translation |
| 2025-09-09 | SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge |
| 2025-09-09 | SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery |
| 2025-09-09 | Small Open Models Achieve Near Parity with Large Models in Low Resource Literary Translation at a Fraction of the Cost |
| 2025-09-09 | Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning |
| 2025-09-08 | UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction |
| 2025-09-08 | EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models |
| 2025-09-08 | Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models |
| 2025-09-08 | MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations |
| 2025-09-08 | Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector |
| 2025-09-07 | KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino |
| 2025-09-07 | Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation |
| 2025-09-07 | Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods |
| 2025-09-07 | MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries |
| 2025-09-07 | MSLEF: Multi-Segment LLM Ensemble Finetuning in Recruitment |
| 2025-09-06 | LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization |
| 2025-09-06 | A Survey of the State-of-the-Art in Conversational Question Answering Systems |
| 2025-09-06 | QCSE: A Pretrained Quantum Context-Sensitive Word Embedding for Natural Language Processing |
| 2025-09-06 | Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation |
| 2025-09-06 | Using Contrastive Learning to Improve Two-Way Reasoning in Large Language Models: The Obfuscation Task as a Case Study |
| 2025-09-05 | Hunyuan-MT Technical Report |
| 2025-09-05 | PRIM: Towards Practical In-Image Multilingual Machine Translation |
| 2025-09-05 | No Translation Needed: Forecasting Quality from Fertility and Metadata |
| 2025-09-05 | The Token Tax: Systematic Bias in Multilingual Tokenization |
| 2025-09-05 | A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning |
| 2025-09-04 | Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation |
| 2025-09-04 | Align-then-Slide: A complete evaluation framework for Ultra-Long Document-Level Machine Translation |
| 2025-09-04 | Exploring NLP Benchmarks in an Extremely Low-Resource Setting |
| 2025-09-04 | MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages |
| 2025-09-04 | MTQA:Matrix of Thought for Enhanced Reasoning in Complex Question Answering |
| 2025-09-03 | English Pronunciation Evaluation without Complex Joint Training: LoRA Fine-tuned Speech Multimodal LLM |
| 2025-09-03 | Expanding the WMT24++ Benchmark with Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader |
| 2025-09-03 | LatPhon: Lightweight Multilingual G2P for Romance Languages and English |
| 2025-09-03 | Training LLMs to be Better Text Embedders through Bidirectional Reconstruction |
| 2025-09-03 | A Long Short-Term Memory (LSTM) Model for Business Sentiment Analysis Based on Recurrent Neural Network |
| 2025-09-02 | FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain |
| 2025-09-02 | Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation |
| 2025-09-02 | LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue |
| 2025-09-02 | Implicit Reasoning in Large Language Models: A Comprehensive Survey |
| 2025-09-02 | A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation |
| 2025-09-01 | Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective |
| 2025-09-01 | chDzDT: Word-level morphology-aware language model for Algerian social media text |
| 2025-09-01 | CSRM-LLM: Embracing Multilingual LLMs for Cold-Start Relevance Matching in Emerging E-commerce Markets |
| 2025-09-01 | MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model |
| 2025-09-01 | ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links |
| 2025-08-31 | TMT: A Simple Way to Translate Topic Models Using Dictionaries |
| 2025-08-31 | CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA |
| 2025-08-31 | Performance Analysis of Supervised Machine Learning Algorithms for Text Classification |
| 2025-08-31 | MedCOD: Enhancing English-to-Spanish Medical Translation of Large Language Models Using Enriched Chain-of-Dictionary Framework |
| 2025-08-31 | EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes |
| 2025-08-30 | A Multi-Strategy Approach for AI-Generated Text Detection |
| 2025-08-30 | Can Multi-turn Self-refined Single Agent LMs with Retrieval Solve Hard Coding Problems? |
| 2025-08-30 | The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang |
| 2025-08-30 | ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics |
| 2025-08-30 | TECP: Token-Entropy Conformal Prediction for LLMs |
| 2025-08-29 | BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning |
| 2025-08-29 | Quantum-Enhanced Natural Language Generation: A Multi-Model Framework with Hybrid Quantum-Classical Architectures |
| 2025-08-29 | Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models |
| 2025-08-29 | Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR |
| 2025-08-29 | Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval |
| 2025-08-28 | Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark |
| 2025-08-28 | The Uneven Impact of Post-Training Quantization in Machine Translation |
| 2025-08-28 | MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers |
| 2025-08-28 | CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance |
| 2025-08-28 | Generative Annotation for ASR Named Entity Correction |
| 2025-08-27 | Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis |
| 2025-08-27 | Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement |
| 2025-08-27 | Survey of Specialized Large Language Model |
| 2025-08-27 | Social Bias in Multilingual Language Models: A Survey |
| 2025-08-27 | Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models |
| 2025-08-26 | Improving Low-Resource Translation with Dictionary-Guided Fine-Tuning and RL: A Spanish-to-Wayuunaiki Study |
| 2025-08-26 | A New NMT Model for Translating Clinical Texts from English to Spanish |
| 2025-08-26 | Automatic Question & Answer Generation Using Generative Large Language Model (LLM) |
| 2025-08-26 | LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordination |
| 2025-08-26 | Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap |
| 2025-08-25 | COMET-poly: Machine Translation Metric Grounded in Other Candidates |
| 2025-08-25 | German4All - A Dataset and Model for Readability-Controlled Paraphrasing in German |
| 2025-08-25 | Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models |
| 2025-08-25 | Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs |
| 2025-08-25 | Information availability in different languages and various technological constraints related to multilinguism on the Internet |
| 2025-08-24 | Evaluating the Impact of Verbal Multiword Expressions on Machine Translation |
| 2025-08-24 | Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models |
| 2025-08-24 | From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users |
| 2025-08-24 | MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models |
| 2025-08-24 | DS@GT at CheckThat! 2025: A Simple Retrieval-First, LLM-Backed Framework for Claim Normalization |
| 2025-08-23 | Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens |
| 2025-08-23 | QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments |
| 2025-08-23 | ObjexMT: Objective Extraction and Metacognitive Calibration for LLM-as-a-Judge under Multi-Turn Jailbreaks |
| 2025-08-23 | Planning for Success: Exploring LLM Long-term Planning Capabilities in Table Understanding |
| 2025-08-23 | DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation |
| 2025-08-22 | OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages |
| 2025-08-22 | M3TQA: Massively Multilingual Multitask Table Question Answering |
| 2025-08-22 | MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering |
| 2025-08-22 | MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian |
| 2025-08-22 | CEQuest: Benchmarking Large Language Models for Construction Estimation |
| 2025-08-21 | Principle Methods of Rendering Non-equivalent Words from Uzbek and Dari to Russian and English |
| 2025-08-21 | Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing |
| 2025-08-21 | Language-Guided Tuning: Enhancing Numeric Optimization with Textual Feedback |
| 2025-08-21 | A Survey on Large Language Model Benchmarks |
| 2025-08-21 | Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs |
| 2025-08-20 | In2x at WMT25 Translation Task |
| 2025-08-20 | Improving LLMs for Machine Translation Using Synthetic Preference Data |
| 2025-08-20 | The Digital Sous Chef -- A Comparative Study on Fine-Tuning Language Models for Recipe Generation |
| 2025-08-20 | Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference |
| 2025-08-20 | Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs |
| 2025-08-19 | MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models |
| 2025-08-19 | Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling |
| 2025-08-19 | MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation |
| 2025-08-19 | AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings |
| 2025-08-19 | Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text |
| 2025-08-18 | DocHPLT: A Massively Multilingual Document-Level Translation Dataset |
| 2025-08-18 | From SALAMANDRA to SALAMANDRATA: BSC Submission for WMT25 General Machine Translation Shared Task |
| 2025-08-18 | Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT |
| 2025-08-18 | RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns |
| 2025-08-18 | Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi |
| 2025-08-17 | Legal$Δ$: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thought Guided Information Gain |
| 2025-08-17 | A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation |
| 2025-08-17 | LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages |
| 2025-08-17 | Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges |
| 2025-08-17 | What do Speech Foundation Models Learn? Analysis and Applications |
| 2025-08-16 | CAMF: Collaborative Adversarial Multi-agent Framework for Machine Generated Text Detection |
| 2025-08-16 | LLMs Struggle with NLI for Perfect Aspect: A Cross-Linguistic Study in Chinese and Japanese |
| 2025-08-16 | Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation |
| 2025-08-16 | EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models |
| 2025-08-16 | Optimizing Token Choice for Code Watermarking: A RL Approach |
| 2025-08-15 | ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection |
| 2025-08-15 | Using Natural Language for Human-Robot Collaboration in the Real World |
| 2025-08-15 | When Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism Detection |
| 2025-08-15 | When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs |
| 2025-08-15 | LLM-Guided Planning and Summary-Based Scientific Text Simplification: DS@GT at CLEF 2025 SimpleText |
| 2025-08-14 | Neural Machine Translation for Coptic-French: Strategies for Low-Resource Ancient Languages |
| 2025-08-14 | From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms |
| 2025-08-14 | Empowering Multimodal LLMs with External Tools: A Comprehensive Survey |
| 2025-08-14 | Can Multi-modal (reasoning) LLMs detect document manipulation? |
| 2025-08-14 | Continuous Bangla Sign Language Translation: Mitigating the Expense of Gloss Annotation with the Assistance of Graph |
| 2025-08-13 | UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech |
| 2025-08-13 | Estimating Machine Translation Difficulty |
| 2025-08-13 | A Comprehensive Evaluation framework of Alignment Techniques for LLMs |
| 2025-08-13 | AINL-Eval 2025 Shared Task: Detection of AI-Generated Scientific Abstracts in Russian |
| 2025-08-13 | Evaluating the Role of Large Language Models in Legal Practice in India |
| 2025-08-12 | A Survey on Training-free Alignment of Large Language Models |
| 2025-08-12 | TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation |
| 2025-08-12 | A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models |
| 2025-08-12 | Optimizing Retrieval-Augmented Generation (RAG) for Colloquial Cantonese: A LoRA-Based Systematic Review |
| 2025-08-12 | TiMoE: Time-Aware Mixture of Language Experts |
| 2025-08-11 | What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction |
| 2025-08-11 | Large Language Models for Subjective Language Understanding: A Survey |
| 2025-08-11 | CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data Augmentation |
| 2025-08-11 | Toward Machine Interpreting: Lessons from Human Interpreting Studies |
| 2025-08-11 | Re:Verse -- Can Your VLM Read a Manga? |
| 2025-08-10 | ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models |
| 2025-08-10 | CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation |
| 2025-08-10 | ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering |
| 2025-08-10 | Event-Aware Sentiment Factors from LLM-Augmented Financial Tweets: A Transparent Framework for Interpretable Quant Trading |
| 2025-08-10 | Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Semantic Alignment |
| 2025-08-09 | AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance |
| 2025-08-09 | BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context |
| 2025-08-09 | Model-Agnostic Sentiment Distribution Stability Analysis for Robust LLM-Generated Texts Detection |
| 2025-08-09 | Text to Speech System for Meitei Mayek Script |
| 2025-08-09 | Annotating Errors in English Learners' Written Language Production: Advancing Automated Written Feedback Systems |
| 2025-08-08 | Testing the Limits of Machine Translation from One Book |
| 2025-08-08 | Evaluating Style-Personalized Text Generation: Challenges and Directions |
| 2025-08-08 | HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning |
| 2025-08-08 | Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents |
| 2025-08-08 | Matrix-Driven Instant Review: Confident Detection and Reconstruction of LLM Plagiarism on PC |
| 2025-08-07 | TASE: Token Awareness and Structured Evaluation for Multilingual Language Models |
| 2025-08-07 | MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs |
| 2025-08-07 | ATLANTIS at SemEval-2025 Task 3: Detecting Hallucinated Text Spans in Question Answering |
| 2025-08-07 | Navigating Through Paper Flood: Advancing LLM-based Paper Evaluation through Domain-Aware Retrieval and Latent Reasoning |
| 2025-08-07 | Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025 |
| 2025-08-06 | Step More: Going Beyond Single Backpropagation in Meta Learning Based Model Editing |
| 2025-08-06 | Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap |
| 2025-08-06 | Multilingual Source Tracing of Speech Deepfakes: A First Benchmark |
| 2025-08-06 | Lightweight Transformers for Zero-Shot and Fine-Tuned Text-to-SQL Generation Using Spider |
| 2025-08-06 | P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis |
| 2025-08-05 | fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval |
| 2025-08-05 | Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models |
| 2025-08-05 | VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation |
| 2025-08-05 | Are We on the Right Way for Assessing Document Retrieval-Augmented Generation? |
| 2025-08-05 | CardiffNLP at CLEARS-2025: Prompting Large Language Models for Plain Language and Easy-to-Read Text Rewriting |
| 2025-08-04 | Test Set Quality in Multilingual LLM Evaluation |
| 2025-08-04 | SHAMI-MT: A Syrian Arabic Dialect to Modern Standard Arabic Bidirectional Machine Translation System |
| 2025-08-04 | The SMeL Test: A simple benchmark for media literacy in language models |
| 2025-08-04 | Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models |
| 2025-08-04 | A French Version of the OLDI Seed Corpus |
| 2025-08-03 | HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark |
| 2025-08-03 | Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback |
| 2025-08-03 | Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe |
| 2025-08-03 | Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language |
| 2025-08-03 | A comprehensive taxonomy of hallucinations in Large Language Models |
| 2025-08-02 | ArzEn-MultiGenre: An aligned parallel dataset of Egyptian Arabic song lyrics, novels, and subtitles, with English translations |
| 2025-08-02 | Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan |
| 2025-08-02 | CSIRO-LT at SemEval-2025 Task 11: Adapting LLMs for Emotion Recognition for Multiple Languages |
| 2025-08-02 | MaRGen: Multi-Agent LLM Approach for Self-Directed Market Research and Analysis |
| 2025-08-02 | Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates |
| 2025-08-01 | MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language |
| 2025-08-01 | Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications |
| 2025-08-01 | Demo: TOSense -- What Did You Just Agree to? |
| 2025-08-01 | Do They Understand Them? An Updated Evaluation on Nonbinary Pronoun Handling in Large Language Models |
| 2025-08-01 | PaPaformer: Language Model from Pre-trained Paraller Paths |
| 2025-07-31 | Comparison of Large Language Models for Deployment Requirements |
| 2025-07-31 | Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis |
| 2025-07-31 | Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning |
| 2025-07-31 | Beyond the Cloud: Assessing the Benefits and Drawbacks of Local LLM Deployment for Translators |
| 2025-07-31 | MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks |
| 2025-07-30 | Opportunities and Challenges of LLMs in Education: An NLP Perspective |
| 2025-07-30 | LLM-Crowdsourced: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models |
| 2025-07-30 | Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation |
| 2025-07-30 | PATENTWRITER: A Benchmarking Study for Patent Drafting with LLMs |
| 2025-07-30 | C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations |
| 2025-07-29 | RL from Teacher-Model Refinement: Gradual Imitation Learning for Machine Translation |
| 2025-07-29 | AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models |
| 2025-07-29 | VN-MTEB: Vietnamese Massive Text Embedding Benchmark |
| 2025-07-29 | Overview of ADoBo at IberLEF 2025: Automatic Detection of Anglicisms in Spanish |
| 2025-07-29 | Post-Training Large Language Models via Reinforcement Learning from Self-Feedback |
| 2025-07-28 | Multilingual Self-Taught Faithfulness Evaluators |
| 2025-07-28 | FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models |
| 2025-07-28 | On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey |
| 2025-07-28 | When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification |
| 2025-07-28 | Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models |
| 2025-07-27 | Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks? |
| 2025-07-27 | AI-Driven Generation of Old English: A Framework for Low-Resource Languages |
| 2025-07-27 | Multi-Agent Interactive Question Generation Framework for Long Document Understanding |
| 2025-07-27 | Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation |
| 2025-07-27 | Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations |
| 2025-07-26 | Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam |
| 2025-07-26 | VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering |
| 2025-07-26 | Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text |
| 2025-07-26 | UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities |
| 2025-07-26 | PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training |
| 2025-07-25 | LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation |
| 2025-07-25 | MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks |
| 2025-07-25 | LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences |
| 2025-07-25 | Towards Domain Specification of Embedding Models in Medicine |
| 2025-07-25 | SpeechIQ: Speech Intelligence Quotient Across Cognitive Levels in Voice Understanding Large Language Models |
| 2025-07-24 | AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs |
| 2025-07-24 | GIIFT: Graph-guided Inductive Image-free Multimodal Machine Translation |
| 2025-07-24 | Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges |
| 2025-07-24 | Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection |
| 2025-07-24 | Checklists Are Better Than Reward Models For Aligning Language Models |
| 2025-07-23 | Natural Language Processing for Tigrinya: Current State and Future Directions |
| 2025-07-23 | Dual-branch Prompting for Multimodal Machine Translation |
| 2025-07-23 | Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challenge |
| 2025-07-23 | Evaluating the Performance of AI Text Detectors, Few-Shot and Chain-of-Thought Prompting Using DeepSeek Generated Text |
| 2025-07-23 | A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task |
| 2025-07-22 | Introducing Quality Estimation to Machine Translation Post-editing Workflow: An Empirical Study on Its Usefulness |
| 2025-07-22 | SiLQ: Simple Large Language Model Quantization-Aware Training |
| 2025-07-22 | GG-BBQ: German Gender Bias Benchmark for Question Answering |
| 2025-07-22 | Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning |
| 2025-07-22 | Beyond Isolated Dots: Benchmarking Structured Table Construction as Deep Knowledge Extraction |
| 2025-07-21 | Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification |
| 2025-07-21 | A Novel Self-Evolution Framework for Large Language Models |
| 2025-07-21 | BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning |
| 2025-07-21 | Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback |
| 2025-07-21 | ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling |
| 2025-07-20 | From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment |
| 2025-07-20 | A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script |
| 2025-07-20 | What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction |
| 2025-07-20 | Tiny language models |
| 2025-07-20 | MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction Tuning for Emotion-Cause Pair Extraction |
| 2025-07-19 | Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification |
| 2025-07-19 | Docopilot: Improving Multimodal Models for Document-Level Understanding |
| 2025-07-19 | Optimizing Legal Document Retrieval in Vietnamese with Semi-Hard Negative Mining |
| 2025-07-19 | Mangosteen: An Open Thai Corpus for Language Model Pretraining |
| 2025-07-19 | Rethinking Suicidal Ideation Detection: A Trustworthy Annotation Framework and Cross-Lingual Model Evaluation |
| 2025-07-18 | NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining |
| 2025-07-18 | Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks |
| 2025-07-18 | Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters |
| 2025-07-18 | Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study |
| 2025-07-18 | Lessons from the TREC Plain Language Adaptation of Biomedical Abstracts (PLABA) track |
| 2025-07-17 | MRT at IberLEF-2025 PRESTA Task: Maximizing Recovery from Tables with Multiple Steps |
| 2025-07-17 | A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models |
| 2025-07-17 | TransEvalnia: Reasoning-based Evaluation and Ranking of Translations |
| 2025-07-17 | MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models |
| 2025-07-17 | Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities |
| 2025-07-16 | The first open machine translation system for the Chechen language |
| 2025-07-16 | Mitigating Stylistic Biases of Machine Translation Systems via Monolingual Corpora Only |
| 2025-07-16 | Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation |
| 2025-07-16 | Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models |
| 2025-07-16 | Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese |
| 2025-07-15 | MapIQ: Benchmarking Multimodal Large Language Models for Map Question Answering |
| 2025-07-15 | FMC: Formalization of Natural Language Mathematical Competition Problems |
| 2025-07-15 | Cross-lingual Few-shot Learning for Persian Sentiment Analysis with Incremental Adaptation |
| 2025-07-15 | Seq vs Seq: An Open Suite of Paired Encoders and Decoders |
| 2025-07-15 | AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles |
| 2025-07-14 | Multiple Choice Learning of Low Rank Adapters for Language Modeling |
| 2025-07-14 | Abusive text transformation using LLMs |
| 2025-07-14 | Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers |
| 2025-07-14 | Absher: A Benchmark for Evaluating Large Language Models Understanding of Saudi Dialects |
| 2025-07-14 | Using AI to replicate human experimental results: a motion study |
| 2025-07-13 | The CoNLL-2013 Shared Task on Grammatical Error Correction |
| 2025-07-13 | An Exploration of Knowledge Editing for Arabic |
| 2025-07-13 | How Important is `Perfect' English for Machine Translation Prompts? |
| 2025-07-13 | Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models |
| 2025-07-13 | NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance |
| 2025-07-12 | Psychology-Driven Enhancement of Humour Translation |
| 2025-07-12 | DS@GT at Touché: Large Language Models for Retrieval-Augmented Debate |
| 2025-07-12 | Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models |
| 2025-07-12 | DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models |
| 2025-07-12 | PU-Lie: Lightweight Deception Detection in Imbalanced Diplomatic Dialogues via Positive-Unlabeled Learning |
| 2025-07-11 | Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency |
| 2025-07-11 | Evaluating LLMs in Medicine: A Call for Rigor, Transparency |
| 2025-07-11 | Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization |
| 2025-07-11 | ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition |
| 2025-07-11 | Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test? |
| 2025-07-10 | Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation |
| 2025-07-10 | Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review |
| 2025-07-10 | Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization |
| 2025-07-10 | Overview of the TREC 2023 deep learning track |
| 2025-07-10 | RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning |
| 2025-07-09 | ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning |
| 2025-07-09 | LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation |
| 2025-07-09 | FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation |
| 2025-07-09 | Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings |
| 2025-07-09 | Checklist Engineering Empowers Multilingual LLM Judges |
| 2025-07-08 | UQLM: A Python Package for Uncertainty Quantification in Large Language Models |
| 2025-07-08 | Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs |
| 2025-07-08 | Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems |
| 2025-07-08 | Evaluating Morphological Alignment of Tokenizers in 70 Languages |
| 2025-07-08 | DS@GT at CheckThat! 2025: Detecting Subjectivity via Transfer-Learning and Corrective Data Augmentation |
| 2025-07-07 | O_FT@EvalLLM2025 : étude comparative de choix de données et de stratégies d'apprentissage pour l'adaptation de modèles de langue à un domaine |
| 2025-07-07 | A Survey of Pun Generation: Datasets, Evaluations and Methodologies |
| 2025-07-07 | Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning |
| 2025-07-07 | Transcribing Spanish Texts from the Past: Experiments with Transkribus, Tesseract and Granite |
| 2025-07-07 | On the Semantics of Large Language Models |
| 2025-07-06 | The role of large language models in UI/UX design: A systematic literature review |
| 2025-07-06 | THM@SimpleText 2025 -- Task 1.1: Revisiting Text Simplification based on Complex Terms for Non-Experts |
| 2025-07-06 | Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts |
| 2025-07-06 | No Language Data Left Behind: A Comparative Study of CJK Language Datasets in the Hugging Face Ecosystem |
| 2025-07-06 | GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models |
| 2025-07-05 | An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand |
| 2025-07-05 | Losing our Tail -- Again: On (Un)Natural Selection And Multilingual Large Language Models |
| 2025-07-05 | LLMThinkBench: Towards Basic Math Reasoning and Overthinking in Large Language Models |
| 2025-07-05 | Large Language Models for Zero-Shot Multicultural Name Recognition |
| 2025-07-05 | Token Level Hallucination Detection via Variance in Language Models |
| 2025-07-04 | Learning to Translate Ambiguous Terminology by Preference Optimization on Post-Edits |
| 2025-07-04 | GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation |
| 2025-07-04 | WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia |
| 2025-07-04 | SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation |
| 2025-07-04 | H2HTalk: Evaluating Large Language Models as Emotional Companion |
| 2025-07-03 | Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models |
| 2025-07-03 | Cautious Next Token Prediction |
| 2025-07-03 | Answer Matching Outperforms Multiple Choice for Language Model Evaluation |
| 2025-07-03 | MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent |
| 2025-07-03 | MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks |
| 2025-07-02 | Adapting Language Models to Indonesian Local Languages: An Empirical Study of Language Transferability on Zero-Shot Settings |
| 2025-07-02 | Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla |
| 2025-07-02 | Confidence and Stability of Global and Pairwise Scores in NLP Evaluation |
| 2025-07-02 | MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining |
| 2025-07-02 | LLMs for Legal Subsumption in German Employment Contracts |
| 2025-07-01 | Transferable Modeling Strategies for Low-Resource LLM Tasks: A Prompt and Alignment-Based Approach |
| 2025-07-01 | Verifiable Natural Language to Linear Temporal Logic Translation: A Benchmark Dataset and Evaluation Suite |
| 2025-07-01 | TransLaw: Benchmarking Large Language Models in Multi-Agent Simulation of the Collaborative Translation |
| 2025-07-01 | Pitfalls of Evaluating Language Models with Open Benchmarks |
| 2025-07-01 | ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models |
| 2025-06-30 | AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data |
| 2025-06-30 | Machine Understanding of Scientific Language |
| 2025-06-30 | Natural language processing for African languages |
| 2025-06-30 | IMPACT: Inflectional Morphology Probes Across Complex Typologies |
| 2025-06-30 | EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning |
| 2025-06-29 | Two Spelling Normalization Approaches Based on Large Language Models |
| 2025-06-29 | Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family |
| 2025-06-29 | RiverText: A Python Library for Training and Evaluating Incremental Word Embeddings from Text Data Streams |
| 2025-06-29 | LLM-Assisted Question-Answering on Technical Documents Using Structured Data-Aware Retrieval Augmented Generation |
| 2025-06-29 | ATGen: A Framework for Active Text Generation |
| 2025-06-28 | MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs |
| 2025-06-28 | The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure |
| 2025-06-28 | Text Production and Comprehension by Human and Artificial Intelligence: Interdisciplinary Workshop Report |
| 2025-06-28 | MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering |
| 2025-06-28 | Jan-nano Technical Report |
| 2025-06-27 | Decoding Machine Translationese in English-Chinese News: LLMs vs. NMTs |
| 2025-06-27 | Can Peter Pan Survive MT? A Stylometric Study of LLMs, NMTs, and HTs in Children's Literature Translation |
| 2025-06-27 | Temperature Matters: Enhancing Watermark Robustness Against Paraphrasing Attacks |
| 2025-06-27 | LinguaSynth: Heterogeneous Linguistic Signals for News Classification |
| 2025-06-27 | The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements |
| 2025-06-26 | Optimising Language Models for Downstream Tasks: A Post-Training Perspective |
| 2025-06-26 | Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks |
| 2025-06-26 | Text2Cypher Across Languages: Evaluating Foundational Models Beyond English |
| 2025-06-26 | Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents |
| 2025-06-26 | Towards Transparent AI: A Survey on Explainable Large Language Models |
| 2025-06-25 | Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation |
| 2025-06-25 | Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content |
| 2025-06-25 | ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset |
| 2025-06-25 | Language Modeling by Language Models |
| 2025-06-25 | How to Retrieve Examples in In-context Learning to Improve Conversational Emotion Recognition using Large Language Models? |
| 2025-06-24 | Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress |
| 2025-06-24 | CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation |
| 2025-06-24 | Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge |
| 2025-06-24 | TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems |
| 2025-06-24 | ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing |
| 2025-06-23 | TranslationCorrect: A Unified Framework for Machine Translation Post-Editing with Predictive Error Assistance |
| 2025-06-23 | Comparative Evaluation of ChatGPT and DeepSeek Across Key NLP Tasks: Strengths, Weaknesses, and Domain-Specific Performance |
| 2025-06-23 | Enhancing Entity Aware Machine Translation with Multi-task Learning |
| 2025-06-23 | MLLP-VRAIN UPV system for the IWSLT 2025 Simultaneous Speech Translation Translation task |
| 2025-06-23 | Enhancing Document Retrieval in COVID-19 Research: Leveraging Large Language Models for Hidden Relation Extraction |
| 2025-06-22 | CareLab at #SMM4H-HeaRD 2025: Insomnia Detection and Food Safety Event Extraction with Domain-Aware Transformers |
| 2025-06-22 | QueueEDIT: Structural Self-Correction for Sequential Model Editing in LLMs |
| 2025-06-22 | $φ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models |
| 2025-06-22 | Statistical Multicriteria Evaluation of LLM-Generated Text |
| 2025-06-22 | LLMs for Customized Marketing Content Generation and Evaluation at Scale |
| 2025-06-21 | Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages |
| 2025-06-21 | TPTT: Transforming Pretrained Transformer into Titans |
| 2025-06-21 | Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning |
| 2025-06-21 | The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future |
| 2025-06-21 | Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights |
| 2025-06-20 | TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs |
| 2025-06-20 | Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs |
| 2025-06-20 | Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning |
| 2025-06-20 | Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025 |
| 2025-06-20 | Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages |
| 2025-06-19 | A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension |
| 2025-06-19 | From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation |
| 2025-06-19 | From LLM-anation to LLM-orchestrator: Coordinating Small Models for Data Labeling |
| 2025-06-19 | NepaliGPT: A Generative Language Model for the Nepali Language |
| 2025-06-19 | End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data |
| 2025-06-18 | Gender-Neutral Machine Translation Strategies in Practice |
| 2025-06-18 | Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation |
| 2025-06-18 | Oldies but Goldies: The Potential of Character N-grams for Romanian Texts |
| 2025-06-18 | WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts |
| 2025-06-18 | video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models |
| 2025-06-17 | MAS-LitEval : Multi-Agent System for Literary Translation Quality Assessment |
| 2025-06-17 | LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data |
| 2025-06-17 | Essential-Web v1.0: 24T tokens of organized web data |
| 2025-06-17 | Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality |
| 2025-06-17 | Memory Tokens: Large Language Models Can Generate Reversible Sentence Embeddings |
| 2025-06-16 | CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right |
| 2025-06-16 | An Interdisciplinary Approach to Human-Centered Machine Translation |
| 2025-06-16 | Edeflip: Supervised Word Translation between English and Yoruba |
| 2025-06-16 | Missing the human touch? A computational stylometry analysis of GPT-4 translations of online Chinese literature |
| 2025-06-16 | CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model |
| 2025-06-15 | Transforming Chatbot Text: A Sequence-to-Sequence Approach |
| 2025-06-15 | Assessing the Role of Data Quality in Training Bilingual Language Models |
| 2025-06-15 | JEBS: A Fine-grained Biomedical Lexical Simplification Task |
| 2025-06-15 | HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance |
| 2025-06-15 | Assessing the Performance Gap Between Lexical and Semantic Models for Information Retrieval With Formulaic Legal Language |
| 2025-06-13 | A Gamified Evaluation and Recruitment Platform for Low Resource Language Machine Translation Systems |
| 2025-06-13 | Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation |
| 2025-06-13 | KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models |
| 2025-06-13 | ImmunoFOMO: Are Language Models missing what oncologists see? |
| 2025-06-13 | AbsenceBench: Language Models Can't Tell What's Missing |
| 2025-06-12 | Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? |
| 2025-06-12 | Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning |
| 2025-06-12 | Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters |
| 2025-06-12 | GLAP: General contrastive audio-text pretraining across domains and languages |
| 2025-06-12 | Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? |
| 2025-06-11 | Towards Efficient and Effective Alignment of Large Language Models |
| 2025-06-11 | Gender Bias in English-to-Greek Machine Translation |
| 2025-06-11 | Taming SQL Complexity: LLM-Based Equivalence Evaluation for Text-to-SQL |
| 2025-06-11 | Unsupervised Elicitation of Language Models |
| 2025-06-11 | Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation |
| 2025-06-10 | TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration |
| 2025-06-10 | mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks |
| 2025-06-10 | Comparing human and LLM proofreading in L2 writing: Impact on lexical and syntactic features |
| 2025-06-10 | Evaluation empirique de la sécurisation et de l'alignement de ChatGPT et Gemini: analyse comparative des vulnérabilités par expérimentations de jailbreaks |
| 2025-06-10 | Advancing STT for Low-Resource Real-World Speech |
| 2025-06-09 | Beyond the Sentence: A Survey on Context-Aware Machine Translation with Large Language Models |
| 2025-06-09 | LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding |
| 2025-06-09 | Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions |
| 2025-06-09 | What Do Indonesians Really Need from Language Technology? A Nationwide Survey |
| 2025-06-09 | Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline |
| 2025-06-08 | Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning |
| 2025-06-08 | Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning |
| 2025-06-08 | Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs |
| 2025-06-08 | Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text |
| 2025-06-08 | ConfQA: Answer Only If You Are Confident |
| 2025-06-07 | MedCite: Can Language Models Generate Verifiable Text for Medicine? |
| 2025-06-07 | BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs |
| 2025-06-07 | Quantile Regression with Large Language Models for Price Prediction |
| 2025-06-07 | Mixture of Small and Large Models for Chinese Spelling Check |
| 2025-06-07 | SafeLawBench: Towards Safe Alignment of Large Language Models |
| 2025-06-06 | Building Models of Neurological Language |
| 2025-06-06 | Corrector Sampling in Language Models |
| 2025-06-06 | Generating Grounded Responses to Counter Misinformation via Learning Efficient Fine-Grained Critiques |
| 2025-06-06 | BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions |
| 2025-06-06 | Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approach |
| 2025-06-05 | SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? |
| 2025-06-05 | RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation |
| 2025-06-05 | ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT |
| 2025-06-05 | LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models |
| 2025-06-05 | Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps Translation |
| 2025-06-04 | Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts |
| 2025-06-04 | TokAlign: Efficient Vocabulary Adaptation via Token Alignment |
| 2025-06-04 | Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis |
| 2025-06-04 | MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP |
| 2025-06-04 | LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation |
| 2025-06-03 | M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset |
| 2025-06-03 | FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models |
| 2025-06-03 | HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring |
| 2025-06-03 | TL;DR: Too Long, Do Re-weighting for Effcient LLM Reasoning Compression |
| 2025-06-03 | It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems |
| 2025-06-02 | Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages |
| 2025-06-02 | Dictionaries to the Rescue: Cross-Lingual Vocabulary Transfer for Low-Resource Languages Using Bilingual Dictionaries |
| 2025-06-02 | MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation |
| 2025-06-02 | Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor |
| 2025-06-02 | Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books |
| 2025-06-01 | COMPKE: Complex Question Answering under Knowledge Editing |
| 2025-06-01 | Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge |
| 2025-06-01 | Trick or Neat: Adversarial Ambiguity and Language Model Evaluation |
| 2025-06-01 | Culturally-Grounded Chain-of-Thought (CG-CoT):Enhancing LLM Performance on Culturally-Specific Tasks in Low-Resource Languages |
| 2025-06-01 | LAQuer: Localized Attribution Queries in Content-grounded Generation |
| 2025-05-31 | Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations |
| 2025-05-31 | Length Aware Speech Translation for Video Dubbing |
| 2025-05-31 | Exploring In-context Example Generation for Machine Translation |
| 2025-05-31 | EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models |
| 2025-05-31 | Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data |
| 2025-05-30 | Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation |
| 2025-05-30 | LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text |
| 2025-05-30 | Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors |
| 2025-05-30 | Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios |
| 2025-05-30 | CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation |
| 2025-05-29 | Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement |
| 2025-05-29 | Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport |
| 2025-05-29 | SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods |
| 2025-05-29 | VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos |
| 2025-05-29 | Enhancing Large Language Models'Machine Translation via Dynamic Focus Anchoring |