When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models
cs.CL
Submitted: 2026-06-01
Updated: 2026-06-01
Journal ref: Findings of the Association for Computational Linguistics: ACL 2026
DOI: 10.18653/v1/2026.findings-acl.1661
Code: https://github.com/sarmistha-D/Hybrid_MOE
License: http://creativecommons.org/licenses/by/4.0/
The gist: In the contemporary epoch of multilingual education, learning idioms provides a fascinating gateway towards creativity, cultural values, historical context, and diverse perspectives inherent to
Terminology
Abstract
In the contemporary epoch of multilingual education, learning idioms provides a fascinating gateway towards creativity, cultural values, historical context, and diverse perspectives inherent to various linguistic traditions. This paper showcases the navigation of retaining figurative and cultural semantics in low-resource Southeast Asian languages such as Hindi, Bengali, and Thai, where culturally rich idioms pose significant obstacles for computational modeling and cross-linguistic transfer due to their deep metaphorical complexity. To tackle such complexity, we present Varnika, a reconstructed multimodal idiom corpus comprising 3,533 multilingual idioms, enriched with seven idiomatic tones aligned with both textual and visual representations. Additionally, to infer informative idiomatic understanding, we introduce a Hybrid Mixture-of-Experts (HybridMoE) framework that embeds multiple idiomatic expert opinions while mitigating expert sparsity by integrating outputs from both selected and unselected experts through controlled hybridization, further augmented with Idiomatic Property Signals via masked multimodal embeddings. To analyze the performance across multiple dimensions, we propose the IDIO-TONE and Idiomatic Validation Score, a three-stage evaluation pipeline measuring (i) literal translation fidelity, (ii) visual-semantic alignment, and (iii) idiomatic meaning retention. Empirical evaluations highlight that HybridMoE achieves 5--6% performance gains across advanced vision language models, demonstrating improved representation of figurative language and culturally embedded meaning in multilingual multimodal settings
Sources
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Hard Nut to Crack: Idiom Detection with Conversational Large Language Models
- Neural Metaphor Detection in Context
- Efficient Large Scale Language Modeling with Mixtures of Experts
- FLUTE: Figurative Language Understanding through Textual Explanations
- MERMAID: Metaphor Generation with Symbolism and Discriminative Decoding
- When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
- Cultural Learning-Based Culture Adaptation of Language Models
- Improved Baselines with Visual Instruction Tuning
- SmolVLM: Redefining small and efficient multimodal models
- TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
- KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
- Typhoon: Thai Large Language Models
- Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- PaliGemma 2: A Family of Versatile VLMs for Transfer
- IRFL: Image Recognition of Figurative Language
- Prototype-based HyperAdapter for Sample-Efficient Multi-task Tuning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering