J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-27
Updated: 2026-09-24
Code: https://github.com/EQ-bench/creative-writing-bench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision.
Terminology
Abstract
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.
Sources
- Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
- Supervising strong learners by amplifying weak experts
- Training Verifiers to Solve Math Word Problems
- DARC: Decoupled Asymmetric Reasoning Curriculum for LLM Evolution
- G-Zero: Self-Play for Open-Ended Generation from Zero Data
- Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
- Language Self-Play For Data-Free Training
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- A Survey on Self-Evolution of Large Language Models
- Gemma 4 Technical Report
- Recursively Summarizing Books with Human Feedback
- Qwen3 Technical Report
- Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
- Dr. Zero: Self-Evolving Search Agents without Training Data
- Instruction-Following Evaluation for Large Language Models
- Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks