SPADE: Self-Play in Adaptive Synthetic Executable Environments
cs.CL, cs.AI
Submitted: 2026-08-19
Updated: 2026-08-31
Comments: Work in progress. Project page: https://spade-rl.github.io ; Code: https://github.com/spade-rl/spade
Code: https://github.com/spade-rl/spade
Project page: https://spade-rl.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- How LLMs Distort Our Written Language
- Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
- Scaling Self-Play with Self-Guidance
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Dota 2 with Large Scale Deep Reinforcement Learning
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- OpenAI Gym
- PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- Towards Understanding Self-play for LLM Reasoning
- From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
- ACEBench: Who Wins the Match Point in Tool Usage?
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
- Self-Questioning Language Models
- Self-Evolving Curriculum for LLM Reasoning
- Scaling Agent Learning via Experience Synthesis
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Computer Environments Elicit General Agentic Intelligence in LLMs
- Anchored Self-Play for Code Repair
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering