CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback
cs.CL, cs.AI, cs.LG
Submitted: 2024-11-13
Updated: 2026-09-12
Comments: full draft v3
Code: https://github.com/draftsubmt/CHAI-LLM
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) show strong performance across many tasks but remain weak at understanding code-mixed (CM) language.
Terminology
Abstract
Large language models (LLMs) show strong performance across many tasks but remain weak at understanding code-mixed (CM) language. Despite this limitation, improving LLMs for CM tasks has received little attention. To address this gap, we propose CHAI, a general-purpose framework for enhancing LLM performance on CM tasks, focusing on CM translation. CHAI leverages four key ideas. First, we investigate the use of LLMs as annotators to address the scarcity of high-quality CM datasets. Second, we leverage these LLM annotations to generate large-scale preference data and apply reinforcement learning from AI feedback (RLAIF) to improve CM translation. Third, we incorporate LLM-generated domain knowledge as a constitution, enabling iterative response refinement. Fourth, we perform extensive evaluations on real-world datasets and settings. Results show that CHAI-powered LLMs outperform state-of-the-art open-source models by 68.45% on average in human-adjudicated win rate on CM translation tasks. This work is a step toward more inclusive open-source code-mixed LLMs.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Unsupervised Cross-lingual Representation Learning at Scale
- The Llama 3 Herd of Models
- Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities
- Low-resource Languages: A Review of Past Work and Future Challenges
- GPT-4 Technical Report
- Multilingual Controlled Generation And Gold-Standard-Agnostic Evaluation of Code-Mixed Sentences
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Training language models to follow instructions with human feedback
- SemEval-2020 Task 9: Overview of Sentiment Analysis of Code-Mixed Tweets
- Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
- Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
- Cross-lingual Language Model Pretraining
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- HRLAIF: Improvements in Helpfulness and Harmlessness in Open-domain Reinforcement Learning From AI Feedback
- SentMix-3L: A Bangla-English-Hindi Code-Mixed Dataset for Sentiment Analysis
- COMET: A Neural Framework for MT Evaluation
- Proximal Policy Optimization Algorithms
- Multilingual Large Language Models Are Not (Yet) Code-Switchers
- HinGE: A Dataset for Generation and Evaluation of Code-Mixed Hinglish Text
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering