INFUSER: Influence-Guided Self-Evolution Improves Reasoning
cs.LG, cs.AI, cs.CL, cs.GT, stat.ML
Submitted: 2026-06-08
Updated: 2026-09-25
Comments: 67 pages, 17 figures
Code: https://github.com/FFishy-git/INFUSER
License: http://creativecommons.org/licenses/by/4.0/
The gist: Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision.
Terminology
Abstract
Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers from a pool of unstructured, automatically collected documents, and a Solver that improves by training on them. The solver is trained with standard correctness rewards against the generator-provided answers, while the generator is rewarded by an optimizer-aware influence score that measures whether each proposed question would actually improve the solver on the target distribution. Because this continuous, noisy influence score is poorly served by standard GRPO, we propose DuGRPO, a dual-normalized variant of GRPO, for generator training. Together, these turn the document pool into an adaptive curriculum that favors questions useful to the current solver, not just hard ones. On Qwen3-8B-Base, INFUSER outperforms strong self-evolution baselines with over 20% relative improvement on Olympiad and SuperGPQA benchmarks, and an 8B INFUSER co-evolving generator outperforms a frozen 32B thinking generator on math and coding. Ablations confirm each design choice is necessary, and two extensions, applying INFUSER to an instruction-finetuned anchor and augmenting it with rule-verifiable RLVR data, further demonstrate the flexibility and generalizability of the framework. Code is available at https://github.com/FFishy-git/INFUSER.
Sources
- Olmo 3
- Scaling Self-Play with Self-Guidance
- Towards Understanding Self-play for LLM Reasoning
- Evaluating Large Language Models Trained on Code
- Self-Evolving Curriculum for LLM Reasoning
- Process Reinforcement through Implicit Rewards
- Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
- Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
- Skywork Open Reasoner 1 Technical Report
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Meta-SGD: Learning to Learn Quickly for Few-Shot Learning
- Efficacy of Language Model Self-Play in Non-Zero-Sum Games
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- Understanding R1-Zero-Like Training: A Critical Perspective
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks