LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
cs.AI, cs.LG
Submitted: 2026-09-02
Updated: 2026-09-02
Terminology
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Concrete Problems in AI Safety
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Scaling Laws for Reward Model Overoptimization
- EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
- AI safety via debate
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Categorizing Variants of Goodhart's Law
- Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
- Feedback Loops With Language Models Drive In-Context Reward Hacking
- Spontaneous Reward Hacking in Iterative Self-Refinement
- Automatic Prompt Optimization with "Gradient Descent" and Beam Search
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection