S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Project page: https://self-developing-agents.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies.
Terminology
Abstract
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use that experience to improve future decisions. We introduce S Gym, an interactive benchmark for evaluating LLM self-improvement through three coupled capabilities: Self-Testing, Self-Judging, and Self-Improvement. S cubed Gym separates permissive exploration from strict held-out evaluation and instantiates this protocol in seven text-based games with executable environment verifiers. We evaluate three pathways for incorporating interaction experience: direct History ICL, score-conditioned Summary Memory, and parameter Training. Our experiments reveal that self-improvement is neither automatic nor uniform. Context-level experience improves performance for several model--game pairs, but the most effective pathway depends strongly on the task structure: summaries are beneficial when experience can be compressed into reusable strategic rules, yet often underperform raw history when success depends on precise, state-contingent information. Parameter training produces substantial gains on some tasks, but also exhibits unstable improvement and severe negative transfer on others. These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies. S cubed Gym provides a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.
Sources
- Self-Improving LLM Agents at Test-Time
- TextWorld: A Learning Environment for Text-based Games
- FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
- Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents
- ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
- Reinforced Self-Training (ReST) for Language Modeling
- Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
- SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
- KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
- MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
- PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
- On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
- STaR: Bootstrapping Reasoning With Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering