PA3: Policy-Aware Agent Alignment through Chain-of-Thought
cs.CL, cs.AI, cs.LG
Submitted: 2026-03-15
Updated: 2026-08-31
Terminology
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
- A Survey on LLM-as-a-Judge
- Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
- LongCoder: A Long-Range Pre-trained Language Model for Code Completion
- Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
- DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
- Training language models to follow instructions with human feedback
- Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis
- APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
- Measuring Faithfulness in Chain-of-Thought Reasoning
- ToolRL: Reward is All Tool Learning Needs
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ActionStudio: A Lightweight Framework for Data and Training of Large Action Models
- Large Language Models are Inconsistent and Biased Evaluators
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering