Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?

arXiv:2603.03202 · cs.CL · Submitted 2026-03-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?".

Jane: As large language models advance their mathematical capabilities, there is a significant bottleneck in training and self-evolution due to the scarcity of challenging, high-quality problems.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Jane, I'm really pumped about this paper, "Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?". The whole idea of using code agents to create harder math problems sounds incredibly promising for training future AI.

Jane: It does sound exciting, Tom; the core thesis seems to be that LLMs hit a wall because there aren't enough tough math problems, and this paper suggests code execution can be a way for agents to generate those new challenges themselves.

Lu: I think what really catches my attention is how they set up this multi-agent framework to manage the complexity of evolving these problems, which is quite sophisticated.

Meng: From an engineering standpoint, I'm curious about how they handle the environment; if the code execution becomes a scalable environment for experimentation, what kind of computational resources are we talking about for these evolution runs?

Lalam: I see a potential here for improving our culture because if this helps LLMs generate truly novel reasoning data, it means our systems could move beyond simple pattern matching toward genuine structural exploration.

Tom: Exactly, Lalam! The paper introduces this multi-agent framework specifically to tackle the bottleneck in training and self-evolution of large language models when it comes to mathematical tasks.

Jane: So, what exactly is the main claim they are making about code agents in this context? What is the central argument driving this research?

Tom: They are investigating whether code agents can autonomously evolve existing math problems into more complex variations by using code execution as a scalable environment for mathematical experimentation. That's the core question they're aiming to answer.

Lu: The authors introduce three distinct stages handled by specialized agents to manage this long-horizon task of adapting the math problems, which is a clever way to structure the exploration process.

Meng: Three stages sound structured, but I wonder how that decomposition actually translates into reliable problem generation without getting stuck in a loop where one agent's output doesn't meet the others' criteria.

Jane: The framework decomposes the task into an Evolution Agent, a Solvability Verification Agent, and a Difficulty Verification Agent, each with specific roles to ensure quality control at every step.

Lalam: Having separate agents for evolution and verification sounds like a way to build in checks for logical validity right from the start of the generation process.

Tom: Right, Jane? The Evolution Agent is supposed to analyze an input problem's solution to find cognitive bottlenecks and then perform "free exploration based on the original problem to design a more challenging new problem," guided by Theory of Mind principles.

Lu: That Theory of Mind guidance is interesting; it suggests the agent isn't just randomly changing variables but actively trying to anticipate how a solver would think, injecting those "new Aha moments" to make the entry point harder.

Paper summary: Jane: And then we have the Solvability Verification Agent, which checks for flaws in both the problem statement and proposed solution steps, ensuring that a logically valid solution provides evidence that at least one solution path exists.

Meng: So they're building a system where one agent proposes, another checks if it works, and a third checks if it's actually hard enough to be useful? That sounds like layered validation.

Tom: Precisely! And finally, the Difficulty Verification Agent uses a five-point scoring mechanism to distinguish between "Artificial Complexity from Cognitive Depth," which specifically penalizes mere computational tedium and rewards adaptations that force a "deviation from rote application."

Lalam: That distinction between computational tedium and cognitive depth is really insightful; it suggests the system is designed to generate problems that truly test reasoning, not just brute force.

Jane: The authors then follow this up with how they use code execution during test-time exploration by employing multiple rollouts from the Evolution Agent until both verification agents’ criteria are satisfied.

Tom: They explicitly govern these agents to utilize code as a tool for empirical inquiry, allowing them to perform "rigorous empirical verification across diverse mathematical domains," which is where things get really hands-on.

Lu: The comprehensive Python sandbox they set up, containing libraries like SymPy and NetworkX, lets the agents perform symbolic computation and deterministic intermediate feedback to guide their evolution process.

Meng: That deterministic feedback loop through symbolic computation is crucial; it means the exploration isn't purely random guessing but guided by rigorous mathematical manipulation within that code environment.

Jane: The evaluation method they use looks at three main aspects: solvability, difficulty increase, and efficiency, which gives a balanced view of the system’s performance.

Tom: For solvability, they use an LLM-as-a-judge approach with a unified third-party model called GPT-five point two-High to keep things consistent.

Lu: And for difficulty escalation, they measure it by observing changes in accuracy and reasoning length across six solver models when comparing the original seed problems to the evolved ones, aiming for a "decrease in performance (i.e., Evolution-SR < Origin-SR)."

Meng: That performance decrease metric is what tells us if the problem actually became harder for current reasoning systems, which is a key measure of success.

Jane: Efficiency is quantified by metrics like the "average number of Evolution Agent rollouts" and analyzing the distribution of "rollout counts across different models," which helps them understand how much computational overhead this process requires.

Tom: The key finding they highlight is that code agents can synthesize new, solvable problems that are "structurally distinct from and more challenging than the originals." That’s a big statement about their capability.

Paper summary: Lalam: It really speaks to the potential for generating high-quality mathematical reasoning data if these agents can consistently produce problems that genuinely push the boundaries of existing solvers.

Jane: They pinpoint three key findings: first, "code-driven exploration helps discover hidden insights," second, "models can generate challenges beyond their own solving baselines," and third, "stronger difficulty enhancement requires nontrivial computational overhead."

Lu: That finding about nontrivial computational overhead is important because it sets a realistic expectation for when this method will be practical for widespread use in training.

Tom: So the conclusion they draw is that code execution serves as a "key exploration engine, enabling a shift from simple verification to deeper structural exploration," suggesting it's a "viable mechanism for synthesizing high-difficulty mathematical reasoning problems within scalable computational environments."

Meng: I think that implies we need robust infrastructure to support this kind of iterative, multi-agent mathematical synthesis if we want to actually deploy this on a large scale.

Jane: And they also flag some limitations: the work acknowledges that the relatively small scale of seed problems, just one hundred is due to the high computational cost of each evolution run.

Tom: Plus, while they show solvability and difficulty increase, they state that these generated problems "do not further verify whether these problems can improve model performance when used as training data." That's a crucial caveat.

Lalam: It shows humility in their research by not claiming the final step—that the generated problems will actually make models better—is solved yet.

Lu: And assessing problem quality itself remains labor-intensive, leading to human evaluation only on "sampled cases rather than the entire generated set." That's a practical hurdle for scaling up this approach.

Tom: So, while the concept of Code2Math is compelling because it moves beyond simple verification into structural exploration, the current state shows that logical consistency emerging as a primary bottleneck highlights a trade-off between reliability and computational overhead.

Jane: Overall, the paper presents an exciting demonstration of how agents can use code execution to navigate complex mathematical spaces autonomously to generate novel problems. It shows a pathway for creating synthetic reasoning data.

Lu: The implication is that we might see a new way of benchmarking LLM mathematical capabilities by using these self-evolving problems as the standard, rather than relying on static benchmarks.

Meng: If this works reliably and efficiently, it could significantly reduce the need for manually curated, expensive benchmark datasets for training advanced AI models.

Lalam: For our culture, seeing an AI system capable of autonomously synthesizing challenging math problems suggests we are moving toward systems that can truly invent novel research material instead of just re-solving known puzzles.

Tom: It’s a powerful mechanism for exploring mathematical reasoning, and this paper lays out a clear path forward for how code agents can start contributing to the creation of complex mathematical challenges on their own.

Conclusion: Tom: So we've been breaking down how these code agents are evolving math problems using their own code execution as an environment, and now we're at the end of this deep dive into "Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?"

Jane: Yeah, Tom, it really boils down to this idea: we’re looking at how AI agents can use programming tools to create new math puzzles that are genuinely harder than the ones they started with.

Lu: The authors are showing how this works by using a multi-agent system where one agent suggests an evolution, and then other agents rigorously check if the problem is still solvable and actually difficult enough to matter.

Meng: From my side, what I find most practical is that they use specific Python libraries like SymPy for symbolic computation, which means the agent isn't just guessing; it's doing actual math checks as it evolves.

Lalam: And from my perspective in the LLM space, this suggests that if we can give an AI the ability to autonomously design challenges, it opens up a whole new avenue for training models on complex reasoning skills.

Tom: Exactly! The big idea here is moving beyond just giving models answers and instead enabling them to actively create the very problems they need to master.

Jane: So when you look at the title, "Can Your Code Agent Evolve Math Problems Through Exploration?", it really asks if these agents can discover new mathematical territory on their own.

Lu: It points toward a future where AI doesn't just learn from existing data but actively contributes to the creation of that data through this iterative process of problem generation and verification.

Meng: I wonder how this translates into real-world impact, Tom; if these agents can create problems that truly stress other models, does that mean we can build more robust AI systems?

Lalam: I think the cultural implication is huge; if we see AI systems capable of generating novel and challenging reasoning material autonomously, it shifts our view of what machines can accomplish in terms of creative intellectual contribution.

Tom: That's a huge point, Lalam. It moves the conversation away from simply asking models questions to having them actively design the test itself.

Jane: And while the paper shows this is a viable way to explore new problem spaces, it also makes it clear that there's still quite a bit of work left before we can say these evolved problems automatically become training gold.

Lu: That's fair; they admitted that assessing the quality of these generated problems still requires significant human effort, which means the computational overhead for generating them is currently pretty high.

Meng: So, it’s like having a super-smart intern who can design complex math problems, but we still need senior researchers to audit their work before it goes into the main training pipeline.

Lalam: That kind of iterative refinement process, guided by code execution and multi-agent checks, seems like the most promising path forward for developing AI that can truly generate novel mathematical challenges.

Tom: It certainly sets a high bar for what we expect from these agentic systems going forward, showing us that deep structural exploration through code is a real mechanism.

Jane: And as we wrap up this segment, the question remains whether the computational cost of this exploration will ever be low enough to make it routine for training purposes.

Hong Kong University of Science and Technology · Tsinghua University · Zhejiang University · Nanjing Tech University · Shanghai Jiao Tong University · University of Michigan

cs.CL

Submitted: 2026-03-03

Updated: 2026-10-02

Comments: 38 pages

Code: https://github.com/TarferSoul/Code2Math

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: As large language models advance their mathematical capabilities, there is a significant bottleneck in training and self-evolution due to the scarcity of challenging, high-quality problems.

Key concepts

Evolution Agent
This agent analyzes a math problem's solution to find cognitive roadblocks. It uses 'Theory of Mind' principles to anticipate how a solver thinks and designs new problems that are more elusive, aiming for 'new Aha moments' that make the entry point harder.
Solvability Verification Agent
This agent rigorously checks both the problem statement and any proposed solution steps. It ensures that if a solution is logically valid, it proves at least one path to the answer exists, confirming the generated problem is actually solvable.
Difficulty Verification Agent
This agent scores problems based on 'Artificial Complexity from Cognitive Depth.' It rewards adaptations that force solvers away from simple rote application toward deeper reasoning, penalizing problems that are only computationally tedious.

Terminology

Summary

As large language models advance their mathematical capabilities, there is a significant bottleneck in training and self-evolution due to the scarcity of challenging, high-quality problems. This paper investigates whether code agents can autonomously evolve existing math problems into more complex variations by leveraging code execution as a scalable environment for mathematical experimentation.

How it works

The authors introduce a multi-agent framework designed to perform problem evolution while validating the solvability and increased difficulty of the generated problems. The system decomposes the long-horizon task of adapting mathematical problems into three distinct stages handled by specialized agents:

  1. The Evolution Agent, which analyzes the solution of the input problem to identify cognitive bottlenecks and performs free exploration based on the original problem to design a more challenging new problem. This agent is guided by Theory of Mind principles to anticipate solver reasoning paths and inject new Aha moments to make the entry point more elusive.

  2. The Solvability Verification Agent, which checks for flaws in both the problem statement and proposed solution steps, ensuring that a logically valid solution provides evidence that at least one solution path exists.

  3. The Difficulty Verification Agent, which assesses difficulty using a 5-point scoring mechanism to distinguish between Artificial Complexity from Cognitive Depth, penalizing mere computational tedium and rewarding adaptations that force a deviation from rote application.

Test-time Exploration through Code

The framework follows the test-time scaling paradigm by using multiple rollouts from the Evolution Agent until both verification agents’ criteria are satisfied. Agents are explicitly governed to utilize code as a tool for empirical inquiry, allowing them to perform rigorous empirical verification across diverse mathematical domains. The system is equipped with a comprehensive Python sandbox containing libraries such as SymPy, NetworkX, and itertools, enabling symbolic computation and deterministic intermediate feedback to guide evolution.

Evaluation Method

The evaluation assesses three primary aspects: solvability, difficulty increase, and efficiency. Solvability is determined using an LLM-as-a-judge approach with a unified third-party model (GPT-5.2-High) to ensure consistency. Difficulty escalation is measured by observing changes in accuracy and reasoning length across six solver models when comparing the original seed problems to the evolved ones, aiming for a "decrease in performance (i.e., Evolution-SR < Origin-SR). Efficiency is quantified by metrics such as the average number of Evolution Agent rollouts and analyzing the distribution of rollout counts across different models."

Key Findings

The experiments demonstrate that code agents can synthesize new, solvable problems that are structurally distinct from and more challenging than the originals. The authors identify three key findings:

  1. code-driven exploration helps discover hidden insights.

  2. models can generate challenges beyond their own solving baselines.

  3. stronger difficulty enhancement requires nontrivial computational overhead.

The paper concludes that code execution serves as a key exploration engine, enabling a shift from simple verification to deeper structural exploration, suggesting it is a viable mechanism for synthesizing high-difficulty mathematical reasoning problems within scalable computational environments. The process often requires multiple rollouts, with logical consistency emerging as a primary bottleneck, highlighting the trade-off between reliability and computational overhead.

Limitations

The work acknowledges several limitations, including the relatively small scale of seed problems (100), which is due to the high computational cost of each evolution run. Furthermore, while generated problems are shown to be solvable and more difficult, the authors state they do not further verify whether these problems can improve model performance when used as training data. Finally, assessing problem quality remains labor-intensive, leading to human evaluation only on sampled cases rather than the entire generated set.

Ethics Statements

The study addresses potential risks regarding automatically generated problems containing errors or inflated difficulty. Mitigation strategies include the use of both solvability and difficulty verification agents, as well as sampled human inspection of generated problems. The authors affirm that all computational experiments use publicly accessible APIs and that human annotation is conducted internally by the paper's authors, ensuring no external participant recruitment.

References

The paper cites numerous recent works spanning LLM reasoning, agentic systems (e.g., Smolagents), and mathematical problem adaptation techniques (e.g., MATH-Perturb, EvolMathEval). Key references include studies on Theory of Mind in LLMs and self-evolving benchmarks like DeepSeekmath-v2. The paper also references the test-time scaling paradigm as a core methodology for agentic exploration.

Appendix Details

The appendix provides detailed experimental setup information, including specific model roles, toolsets available in the Smolagents environment (SymPy, Z3, NetworkX), and prompt templates used for the Evolution Agent and Solvability Verification Agent. It also includes results from human audits confirming high agreement rates between GPT-5.2-High judgments and human assessments regarding validity and difficulty improvement on Olympiad-level problems.

Improvements for AI systems

Based on the scientific paper Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?, here are specific, high-impact improvements for AI systems and what those improved systems will be capable of:


) 1. Enhanced Mathematical Problem Synthesis and Generation (The Core Capability)

The system can autonomously evolve existing mathematical problems into novel, high-difficulty variations.

  • The agent will utilize a multi-agent framework (Evolution Agent, Solvability Verification Agent, Difficulty Verification Agent) to systematically modify seed problems.

  • It will leverage code execution tools (SymPy, Z3, NetworkX) to perform rigorous symbolic manipulation and constraint satisfaction testing during the evolution process.

  • The agent can be prompted with a Burden of Discovery mandate—forcing it to design problems where the entry point is deliberately obscured, requiring genuine mathematical insight rather than rote application of formulas.

) 2. Scalable Mathematical Experimentation (Test-Time Scaling)

The system can act as an autonomous computational laboratory for mathematical reasoning.

  • It can perform test-time scaling by generating multiple problem variations and running them through a verification pipeline simultaneously to efficiently explore the solution space.

  • It will be capable of empirical inquiry—running numerical simulations, probing tight inequality bounds, or searching for counterexamples to validate hypotheses in real-time.

) 3. Robust Problem Quality Assurance (Self-Correction and Validation)

The system can ensure generated problems are mathematically sound and genuinely challenging before being deployed as training data.

  • It will employ a dual verification system: one agent checks for logical consistency/solvability, and another assesses the Burden of Discovery (distinguishing genuine insight from mere computational tedium).

  • By using an external, high-capability LLM judge (like GPT-5.2) to verify solvability, the system ensures generated problems maintain high mathematical soundness.

) 4. Deep Insight into Model Capabilities (Capability Asymmetry Analysis)

The system can map the limits and strengths of different reasoning models in complex problem synthesis tasks.

  • It can identify capability asymmetries where certain LLM backbones (e.g., DeepSeek-Reasoner) are superior at introducing structural modifications that significantly degrade the performance of other solvers, suggesting targeted evolution strategies.

) 5. Efficient Exploration Strategy (Optimized Search for Difficulty)

The system will develop efficient heuristics for problem evolution, moving beyond simple text-based long-chain reasoning.

  • It will learn to prioritize exploration paths that yield high Aha moments (high difficulty scores) while minimizing computational overhead (rollout count), optimizing the trade-off between reliability and efficiency.

) Improved AI System Capabilities Summary:

The resulting AI system moves beyond simple question answering or text generation; it becomes an autonomous, self-improving mathematical researcher capable of:

  1. Synthesizing novel, high-difficulty mathematical challenges that test the limits of current LLM reasoning.

  2. Acting as a dynamic environment where code execution drives structured exploration to discover new mathematical patterns.

  3. Generating benchmark data (problems) that are rigorously validated for solvability and difficulty escalation, thereby creating a self-improving loop for training and evaluating other AI systems.

Sources

Related papers