Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
cs.LG, cs.AI, cs.CL
Submitted: 2026-03-10
Updated: 2026-08-26
Comments: ICML 2026
Code: https://github.com/nam630/mutual_
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labeled data or external verifiers.
Terminology
Abstract
While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labeled data or external verifiers. Existing data has already been exploited, and new data is expensive to collect. Moreover, true intelligence goes far beyond verifiable tasks. Therefore, we need self-improvement frameworks that are less dependent on external signals and more broadly applicable to both verifiable and non-verifiable domains. We propose **Mutual Information Preference Optimization (MIPO)**, a contrastive data augmentation method that constructs preference pairs by generating a positive response conditioning on the correct prompt, and a negative response by conditioning on a random, unrelated prompt. We show that using Direct Preference Optimization to learn from this paired data maximizes pointwise mutual information *under the base LLM* between prompts and model responses. Experiments with with 1-7B parameter Llama and Qwen instruct models show that MIPO achieves 3-16% gains (and 51% increase for Qwen2.5-1.5B-Instruct) on personalization compared to prompting baselines. Surprisingly, MIPO can also be useful in verifiable domains, such as math and multiple-choice question answering, yielding 1-20% gains *without any additional data or external supervision*. These results suggest a promising direction for self-improvement using intrinsic signals derived from contrastive data pairs.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- A Simple Framework for Contrastive Learning of Visual Representations
- IterAlign: Iterative Constitutional Alignment of Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
- STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving
- Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
- Efficient Exploration for LLMs
- Contrastive Learning as Goal-Conditioned Reinforcement Learning
- Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
- Contrastive Preference Learning: Learning from Human Feedback without RL
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- OLMES: A Standard for Language Model Evaluations
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- Large Language Models Cannot Self-Correct Reasoning Yet
- MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Aligning to Thousands of Preferences via System Message Generalization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks