An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems

summary

Video file (mp4)

The gist

The gist: The proposed Generalization–and–Perturbation (GAP) framework introduces a systematic methodology to assess LLMs’ mathematical-reasoning robustness by stress-testing them on

In short

The Generalization–and–Perturbation (GAP) framework systematically tests LLMs' mathematical reasoning by presenting problems that are mathematically equivalent but linguistically or parametrically varied. Results show that high scores on standard benchmarks can collapse when simple changes, like renaming symbols or altering numbers, are introduced. This reveals a weakness in models' true generalization beyond memorized patterns.

Key concepts

Generalization–and–Perturbation (GAP) Framework
A methodology used to stress-test LLMs by transforming original math problems into mathematically equivalent versions. It specifically tests how well a model maintains accuracy when the wording or numerical parameters change, revealing sensitivity to non-mathematical variations.
Tsurf and Tpara Super-families
These are two types of transformations applied to a problem. Tsurf involves surface renames that change symbol salience, while Tpara involves kernel rewrites that preserve the underlying proof steps but alter the scenario or numerical values. GAP uses both to test robustness.
Robustness Metric (Rb)
A measure quantifying how robust a model is to perturbations. It calculates a score based on correctness across different problem variants, specifically looking at the drop in accuracy when moving from an easy problem statement to a harder variant.

Terminology used across episodes

This episode discusses

The paper

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems · Read on arXiv

University of Illinois Urbana–Champaign · Stanford University

Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent variable roles, and (2) kernel rewrites, probing whether a high-level proof plan survives a change of mathematical setting. Compared with existing benchmarks, GAP has two key benefits: (1) novel, likely unseen variants mitigate data leakage, and (2) performance across transformation families enables failure diagnosis, each transformation testing a hypothesis about the cause of failure. We instantiate GAP on all 1,051 William Lowell Putnam Competition problems from 1938 to 2024, adding 5,255 unseen variants to form PutnamGAP, a 6,306-item competition-level mathematics corpus and the first public machine-readable dataset from the full Putnam archive. Using PutnamGAP, we evaluated 18 commercial and open-source models spanning sizes and providers. Accuracy drops across all models and variant families, most severely under kernel rewrites. This gap does not close with model strength, suggesting that even the strongest models' dominant weakness is transferring a proof plan to a changed mathematical setting, rather than handling surface changes. Further analysis provides finer failure diagnoses and potentially useful insights for improving LLM reasoning.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "An Investigation of Robustness of LLMs in Mathematical Reasoning".

Jane: The gist:

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about who wrote this and what they're calling their work. The authors are Yuren Hao, Xiang Wan, and Chengxiang Zhai from institutions like Illinois Urbana-Champaign and Stanford. They’re tackling a serious problem in the math side of AI.

Jane: It’s important to understand that the title of "An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems" tells us exactly what they did: they investigated robustness using mathematically equivalent transformations on advanced math problems.

Lu: What this investigation is really about is that current methods for testing these models are weak because existing benchmarks are either too small or don't have enough controlled variations to show true generalization #pg2.

Meng: So, the implication here is that we can finally move past just seeing high accuracy scores on GSM8K or MATH and start asking if those models can actually handle the kind of complex, multi-domain problems found in things like the Putnam competition.

Lalam: It suggests that our current testing methods are masking real weaknesses, which means what we measure right now might not be a reliable indicator of true reasoning ability #pg2.

Tom: Exactly. So, when you look at the setup for this paper, it’s about building a way to create an unbounded supply of test items so that those items don't end up polluting future training data.

Jane: Right. They are creating a dataset called PutnamGAP which takes every William Lowell Putnam Competition problem from one thousand nine hundred thirty-eight to two thousand twenty-four which is one thousand fifty-one originals, and expands each one into five variants—four surface renames and one kernel rewrite—giving them six thousand three hundred six stress-test questions #pg3 <ref:2508.08833#pg1>.

Lu: That scale is huge. They’re using this to test the model's capacity to generalize beyond memorized surface forms, which is what they call robustness #pg3.

Meng: So, if a model passes this rigorous test, it means it can handle changes in wording or parameters without losing its core mathematical reasoning skill.

Lalam: That’s a big deal because it shows that the high scores we see today might just be due to memorizing the exact phrasing of the problems they were trained on #pg3.

The paper's summary: Tom: So, what’s the actual summary of this work then? The authors introduce a framework called GAP to assess robustness by stress-testing models on problems that are mathematically equivalent but have different wording or parameters.

Jane: That GAP framework is their main contribution. They break down these transformations into two types: surface renames and kernel rewrites, which are designed to preserve the same proof steps while changing the scenario or parameters #pg3.

Lu: The authors run this whole thing through a two-stage check, involving fifteen rounds of O3 self-review plus a ten percent spot-check, and they found that all eighteen models they tested suffered from both simple renaming and step-based rewrites #pg3.

Meng: So the main point is that even top performers struggle when you introduce these cosmetic or structural changes because their high scores can collapse under those kinds of perturbations #pg3.

Lalam: It shows that even when a model gets a high score, it doesn't mean it has learned the actual underlying math structure, which is what we need to see in reasoning #pg3.

Tom: So they are showing us that these models have this fragility when things get slightly different, whether it’s just renaming symbols or changing the numbers involved in the problem.

Jane: They are essentially creating a diagnostic tool for analyzing and quantifying how robust an LLM’s mathematical reasoning capacity is at the level of competition problems #pg3.

Lu: The paper is focused on measuring how far a model can generalize beyond memorized surface forms, which is what they call robustness #pg3.

Meng: From an engineering standpoint, it confirms that simply increasing the size of the training data isn't enough; we need to change *how* we train them to handle these structural variations early on.

The paper's improvements: Tom: Now let’s talk about what the authors suggest we do next. They aren't just showing us a problem; they are proposing ways to actually improve these models based on this stress testing.

Jane: The first improvement they point to is using the GAP framework as a general diagnostic tool for analyzing and quantifying how robust an LLM’s mathematical reasoning capacity is at the level of competition problems #pg3.

Lu: They suggest that instead of just focusing on making models bigger, we should focus on curriculum fine-tuning where we explicitly randomize things like symbol identities and numeric parameters, rather than just enlarging the pre-training corpora #pg3.

Meng: That makes sense for practical application. Instead of just throwing more data at it, we should be training it in a way that forces it to handle those kinds of structural variations early on so it doesn't fail later when presented with novel phrasing #pg3.

Lalam: And from a risk perspective, this framework helps us see where the fragility is. If we can detect these types of sensitivity, production systems can be vulnerable to things like prompt injections that use mathematically innocuous renamings #pg3.

Tom: So the core improvement suggested is shifting from just scale-up training to training strategies that deliberately randomize inputs, forcing the model to learn structural invariance rather than just linguistic fluency.

Jane: They are showing us that improving robustness requires training strategies focused on structural invariance rather than just linguistic fluency #pg3.

Lu: And they also highlight a key finding: Descriptive Long renaming had the smallest impact overall, with drops being marginal and mostly not significant #pg3.

Conclusion: Tom: We’re wrapping up this discussion on "An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems." The main implication here is that we need a systematic method to test if math reasoning skills are genuine or just rote memorization #pg1.

Jane: The PutnamGAP dataset and the GAP framework give us a way to generate an unbounded supply of evaluation items, which helps limit future leakage into training corpora #pg2. This means we can stop relying on test scores that might be artificially inflated by data contamination.

Lu: They found that Descriptive Long renaming had the smallest impact overall, with drops being marginal, but the Kernel Variant—which keeps the math structure but changes constants and expressions—led to the sharpest decline across all eighteen models #pg3.

Meng: For us in development, this means we need to start thinking about curriculum fine-tuning that deliberately randomizes those parameters instead of just relying on sheer scale for improvement #pg3.

Lalam: Ultimately, this work suggests that improving robustness requires training strategies focused on structural invariance rather than just linguistic fluency #pg3. We need models that can reason regardless of how the symbols are presented.

Tom: So, to summarize, this paper shows us a path toward systematically measuring how far an LLM can generalize beyond memorized forms by stress-testing it with mathematically equivalent transformations on competition math problems #pg1. It’s a solid step in understanding the true limits of current reasoning AI.

More episodes

← Home