Self-Improvement as Coherence Optimization: A Theoretical Account
cs.LG, cs.AI, cs.CL
Submitted: 2026-01-20
Updated: 2026-09-23
Terminology
Sources
- Measuring Progress on Scalable Oversight for Large Language Models
- Avoiding Obfuscation with Prover-Estimator Debate
- HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
- AI safety via debate
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- On the generalization of language models from in-context learning and finetuning: a controlled study
- Scalable agent alignment via reward modeling: a research direction
- Composing Ensembles of Pre-trained Models via Iterative Consensus
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- Unsupervised Elicitation of Language Models
- Recursively Summarizing Books with Human Feedback
- An Explanation of In-context Learning as Implicit Bayesian Inference
- Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks