ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
cs.LG, cs.AI, stat.ML
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/SewoongLab/abc-align
Terminology
Sources
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- On the Convergence of SGD with Biased Gradients
- PPI++: Efficient Prediction-Powered Inference
- PPI-SVRG: Unifying Prediction-Powered Inference and Variance Reduction for Semi-Supervised Optimization
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- Constitutional AI: Harmlessness from AI Feedback
- AutoEval Done Right: Using Synthetic Data for Model Evaluation
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- A Unifying Framework for Robust and Efficient Inference with Unstructured Data
- Quantifying the Gain in Weak-to-Strong Generalization
- Surrogate-Powered Inference: Regularization and Adaptivity
- A Unified Framework for Inference with General Missingness Patterns and Machine Learning Imputation
- Provably Robust DPO: Aligning Language Models with Noisy Feedback
- A Guide Through the Zoo of Biased SGD
- Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
- On Biased Stochastic Gradient Estimation
- The Llama 3 Herd of Models
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Prediction-Powered Inference Across Many Tasks for AI Evaluation & Social Science Research
- Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks