Beyond Surface Alignment: Grounding the Dynamics of Situational Understanding and Generative Control in LLMs
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: PhD Thesis submitted to UChicago (https://knowledge.uchicago.edu/records/5jrnc-nvp13). Update Branching Factor, AI Realtor, and BACo to the latest versions
Code: https://github.com/openai/openai-cookbook
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
- The Capacity for Moral Self-Correction in Large Language Models
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- Program Synthesis with Large Language Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Mirostat: A Neural Text Decoding Algorithm that Directly Controls Perplexity
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Theoretical guarantees on the best-of-n alignment policy
- Is Exploration or Optimization the Problem for Deep Reinforcement Learning?
- How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
- Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
- NatGen: Generative pre-training by "Naturalizing" source code
- Transfer Q Star: Principled Decoding for LLM Alignment
- Decoding Game: On Minimax Optimality of Heuristic Text Generation Strategies
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Universal Self-Consistency for Large Language Model Generation
- Teaching Large Language Models to Self-Debug
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- Reasoning with Exploration: An Entropy Perspective
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering