When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates
cs.CL
Submitted: 2026-06-11
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026
Code: https://github.com/mbzuai-nlp/SemCog
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Aya 23: Open Weight Releases to Further Multilingual Progress
- Large-Scale Machine Translation between Arabic and Hebrew: Available Corpora and Initial Results
- An Arabic-Hebrew parallel corpus of TED talks
- DeepSeek-V3 Technical Report
- Symphonym: Universal Phonetic Embeddings for Cross-Script Toponym Matching
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- The Llama 3 Herd of Models
- Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech
- GPT-4 Technical Report
- Qwen2.5 Technical Report
- IMPACT: Inflectional Morphology Probes Across Complex Typologies
- Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models
- Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities
- OpenAI GPT-5 System Card
- Multilingual LLMs Struggle to Link Orthography and Semantics in Bilingual Word Processing
- Gemma 2: Improving Open Language Models at a Practical Size
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering