Instruction Quality Matters: Refining Instructions for Effective Preference Learning
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: Preprint
Code: https://github.com/01choco/instruction-refinement
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Less is More: Improving LLM Alignment via Preference Data Selection
- Impact of Preference Noise on the Alignment Performance of Generative Language Models
- The Llama 3 Herd of Models
- Uncertainty-Penalized Direct Preference Optimization
- Constitutional AI: Harmlessness from AI Feedback
- From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
- Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
- DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in Large Language Models
- Alignment Data Map for Efficient Preference Data Selection and Diagnosis
- Red Teaming Language Models with Language Models
- Holistic Evaluation of Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
- WizardLM: Empowering large pre-trained language models to follow complex instructions
- UltraMedical: Building Specialized Generalists in Biomedicine
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering