Self-Designed Evaluators and Warm Memory for Long-Horizon Agents
cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Memory Reward Inflation in Self-Improving LLM Agents
- Constitutional AI: Harmlessness from AI Feedback
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- CodeT: Code Generation with Generated Tests
- From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape
- Teaching Large Language Models to Self-Debug
- GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Training Verifiers to Solve Math Word Problems
- TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
- TRAIL: Trace Reasoning and Agentic Issue Localization
- AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection