An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems

arXiv:2508.08833 · cs.CL, cs.AI, cs.LG · Submitted 2025-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "An Investigation of Robustness of LLMs in Mathematical Reasoning".

Jane: The gist:

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about who wrote this and what they're calling their work. The authors are Yuren Hao, Xiang Wan, and Chengxiang Zhai from institutions like Illinois Urbana-Champaign and Stanford. They’re tackling a serious problem in the math side of AI.

Jane: It’s important to understand that the title of "An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems" tells us exactly what they did: they investigated robustness using mathematically equivalent transformations on advanced math problems.

Lu: What this investigation is really about is that current methods for testing these models are weak because existing benchmarks are either too small or don't have enough controlled variations to show true generalization #pg2.

Meng: So, the implication here is that we can finally move past just seeing high accuracy scores on GSM8K or MATH and start asking if those models can actually handle the kind of complex, multi-domain problems found in things like the Putnam competition.

Lalam: It suggests that our current testing methods are masking real weaknesses, which means what we measure right now might not be a reliable indicator of true reasoning ability #pg2.

Tom: Exactly. So, when you look at the setup for this paper, it’s about building a way to create an unbounded supply of test items so that those items don't end up polluting future training data.

Jane: Right. They are creating a dataset called PutnamGAP which takes every William Lowell Putnam Competition problem from one thousand nine hundred thirty-eight to two thousand twenty-four which is one thousand fifty-one originals, and expands each one into five variants—four surface renames and one kernel rewrite—giving them six thousand three hundred six stress-test questions #pg3 <ref:2508.08833#pg1>.

Lu: That scale is huge. They’re using this to test the model's capacity to generalize beyond memorized surface forms, which is what they call robustness #pg3.

Meng: So, if a model passes this rigorous test, it means it can handle changes in wording or parameters without losing its core mathematical reasoning skill.

Lalam: That’s a big deal because it shows that the high scores we see today might just be due to memorizing the exact phrasing of the problems they were trained on #pg3.

The paper's summary: Tom: So, what’s the actual summary of this work then? The authors introduce a framework called GAP to assess robustness by stress-testing models on problems that are mathematically equivalent but have different wording or parameters.

Jane: That GAP framework is their main contribution. They break down these transformations into two types: surface renames and kernel rewrites, which are designed to preserve the same proof steps while changing the scenario or parameters #pg3.

Lu: The authors run this whole thing through a two-stage check, involving fifteen rounds of O3 self-review plus a ten percent spot-check, and they found that all eighteen models they tested suffered from both simple renaming and step-based rewrites #pg3.

Meng: So the main point is that even top performers struggle when you introduce these cosmetic or structural changes because their high scores can collapse under those kinds of perturbations #pg3.

Lalam: It shows that even when a model gets a high score, it doesn't mean it has learned the actual underlying math structure, which is what we need to see in reasoning #pg3.

Tom: So they are showing us that these models have this fragility when things get slightly different, whether it’s just renaming symbols or changing the numbers involved in the problem.

Jane: They are essentially creating a diagnostic tool for analyzing and quantifying how robust an LLM’s mathematical reasoning capacity is at the level of competition problems #pg3.

Lu: The paper is focused on measuring how far a model can generalize beyond memorized surface forms, which is what they call robustness #pg3.

Meng: From an engineering standpoint, it confirms that simply increasing the size of the training data isn't enough; we need to change *how* we train them to handle these structural variations early on.

The paper's improvements: Tom: Now let’s talk about what the authors suggest we do next. They aren't just showing us a problem; they are proposing ways to actually improve these models based on this stress testing.

Jane: The first improvement they point to is using the GAP framework as a general diagnostic tool for analyzing and quantifying how robust an LLM’s mathematical reasoning capacity is at the level of competition problems #pg3.

Lu: They suggest that instead of just focusing on making models bigger, we should focus on curriculum fine-tuning where we explicitly randomize things like symbol identities and numeric parameters, rather than just enlarging the pre-training corpora #pg3.

Meng: That makes sense for practical application. Instead of just throwing more data at it, we should be training it in a way that forces it to handle those kinds of structural variations early on so it doesn't fail later when presented with novel phrasing #pg3.

Lalam: And from a risk perspective, this framework helps us see where the fragility is. If we can detect these types of sensitivity, production systems can be vulnerable to things like prompt injections that use mathematically innocuous renamings #pg3.

Tom: So the core improvement suggested is shifting from just scale-up training to training strategies that deliberately randomize inputs, forcing the model to learn structural invariance rather than just linguistic fluency.

Jane: They are showing us that improving robustness requires training strategies focused on structural invariance rather than just linguistic fluency #pg3.

Lu: And they also highlight a key finding: Descriptive Long renaming had the smallest impact overall, with drops being marginal and mostly not significant #pg3.

Conclusion: Tom: We’re wrapping up this discussion on "An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems." The main implication here is that we need a systematic method to test if math reasoning skills are genuine or just rote memorization #pg1.

Jane: The PutnamGAP dataset and the GAP framework give us a way to generate an unbounded supply of evaluation items, which helps limit future leakage into training corpora #pg2. This means we can stop relying on test scores that might be artificially inflated by data contamination.

Lu: They found that Descriptive Long renaming had the smallest impact overall, with drops being marginal, but the Kernel Variant—which keeps the math structure but changes constants and expressions—led to the sharpest decline across all eighteen models #pg3.

Meng: For us in development, this means we need to start thinking about curriculum fine-tuning that deliberately randomizes those parameters instead of just relying on sheer scale for improvement #pg3.

Lalam: Ultimately, this work suggests that improving robustness requires training strategies focused on structural invariance rather than just linguistic fluency #pg3. We need models that can reason regardless of how the symbols are presented.

Tom: So, to summarize, this paper shows us a path toward systematically measuring how far an LLM can generalize beyond memorized forms by stress-testing it with mathematically equivalent transformations on competition math problems #pg1. It’s a solid step in understanding the true limits of current reasoning AI.

University of Illinois Urbana–Champaign · Stanford University

cs.CL, cs.AI, cs.LG

Submitted: 2025-08-12

Updated: 2026-10-07

Comments: 34 pages, 9 figures, accepted at ICLR 2026 workshop

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 83/100

The gist: The gist: The proposed Generalization–and–Perturbation (GAP) framework introduces a systematic methodology to assess LLMs’ mathematical-reasoning robustness by stress-testing them on

Key concepts

Generalization–and–Perturbation (GAP) Framework
A methodology used to stress-test LLMs by transforming original math problems into mathematically equivalent versions. It specifically tests how well a model maintains accuracy when the wording or numerical parameters change, revealing sensitivity to non-mathematical variations.
Tsurf and Tpara Super-families
These are two types of transformations applied to a problem. Tsurf involves surface renames that change symbol salience, while Tpara involves kernel rewrites that preserve the underlying proof steps but alter the scenario or numerical values. GAP uses both to test robustness.
Robustness Metric (Rb)
A measure quantifying how robust a model is to perturbations. It calculates a score based on correctness across different problem variants, specifically looking at the drop in accuracy when moving from an easy problem statement to a harder variant.

Terminology

Summary

The gist: The proposed Generalization–and–Perturbation (GAP) framework introduces a systematic methodology to assess LLMs’ mathematical-reasoning robustness by stress-testing them on mathematically equivalent but linguistically and parametrically varied advanced math problems, revealing sharp performance degradation across multiple models.

Motivation

Modern AI systems are increasingly entrusted with tasks that hinge on robust reasoning rather than pattern matching, making it important to precisely measure an LLM’s reasoning capacity and its ability to generalize beyond memorized textual surface forms. Existing math-reasoning benchmarks exhibit two critical weaknesses: (i) leakage-induced score inflation, since benchmark items rapidly seep into pre-training corpora, and (ii) limited robustness coverage, because today’s datasets are too small or lack controlled transformations that probe true generalization. Benchmark inflation through training leakage is a concern because public datasets have leaked into the web-scale corpora used to pre-train large language models (LLMs), artificially inflating test-time accuracy. Competition mathematics reveals the next robustness bottleneck because large language models now surpass 90% accuracy on widely-used benchmarks such as GSM8K and MATH, prompting claims of “near-human” numerical reasoning yet still falter on Olympiad-style or Putnam-level problems that intertwine multiple domains. Existing perturbation-based robustness benchmarks have begun to probe mathematical robustness by constructing perturbation-based benchmarks on top of GSM8K and related datasets. Beyond GSM8K, GSM8K MORE uses an ontology of perturbations to generate families of grade-school arithmetic variants (Hong et al., 2025), while Putnam-AXIOM introduces a smaller set of functional variations for university-level Putnam problems. Consequently, the existing perturbation benchmarks do not yet provide a largescale, systematically structured robustness test for competition-level, proof-style mathematics.

The Generalization–and–Perturbation (GAP) Framework

The GAP framework addresses both leakage and robustness by stress-testing the model on mathematically equivalent versions of the same problem. Robustness is defined as the expected accuracy when a problem x is transformed by a family T of equivalence-preserving operators. This framework partitions T into two disjoint transformation super-families: Tsurf (surface renames that alter symbol salience) and Tpara (kernel rewrites that preserve the same proof steps while changing the scenario and parameters). GAP serves as a general diagnostic evaluation methodology for analyzing and quantifying the robustness of an LLM’s mathematical reasoning capacity at the level of competition problems.

Evaluation Model and Metrics

The evaluation model starts from a curated set of N canonical items P = ⟨xi, yi, πi⟩ where xi is a problem statement, yi is its reference answer(s), and πi an unreleased expert solution path used internally for safe variant generation. The model interface receives a prompt x and returns ŷ = fθ(x), which an automatic checker maps to a binary label z = grade(ˆy, y) ∈ 0, 1. The penalty robustness Rb is defined as Rb(e, h) = 1/N Σ exp− β dbj where dbj is the soft-saturated version of the drop dj.

PutnamGAP Dataset

The PutnamGAP dataset comprises all Putnam Problems 1938–2024 (N = 1,051 items after deduplication), expanding each item into five mathematically equivalent variants, totaling 6,306 items. The part distribution is balanced (527 A vs. 524 B), and the difficulty proxy uses indices 1−2 as Easy, 3−4 as Medium, and 5−6 as Hard.

Experimental Results

Across 18 models, all of them suffer from both simple renaming and step-based rewrites. OpenAI’s O3 scores 51.5% on original statements but loses 4.7 pp under surface renames and 12.9 pp under parametric rewrites. The results show that high leaderboard scores can collapse when cosmetic or structural perturbations are applied—precisely the effect that data leakage masks.

Robustness Breakdown

The analysis of transformation-wise breakdown shows that Descriptive Long (DL) renaming has the smallest impact overall, with drops being marginal and mostly not significant. Kernel Variant (KV)—which keeps each question’s mathematical structure but replaces constants and expressions with different values—led to the sharpest decline overall.

Improvements for AI systems

  1. Robustness evaluation via GAP framework: AI systems can be stress-tested on mathematically equivalent but with linguistic and parametric variation problems to measure their sensitivity to non-mathematical perturbations. This enables a more accurate evaluation of reasoning capabilities by quantifying how far a model can generalize beyond memorized forms, as captured by the metric Robustness Metric. Let e, h ∈ [0, 1] N denote per-item correctness on the easy (original) and hard (variant) sets.

  2. Leakage mitigation in benchmarking: The creation of PutnamGAP allows for an infinite stream of unseen test items, which mitigates future contamination by limiting where benchmark items are used in training corpora, thereby ensuring that a leaderboard score no longer guarantees genuine reasoning ability.

  3. Targeted fine-tuning strategies: AI models can be improved via curriculum fine-tuning that explicitly randomizes (i) symbol identities and (ii) numeric parameters, instead of simply enlarging pre-training corpora, which the paper suggests is necessary to improve robustness for formal reasoning domains.

  4. Security risk detection in production: Surface-level fragility allows for the detection of risks, as production systems can be prompt-injected with mathematically innocuous renamings—highlighting the need to integrate robustness checks into red-team pipelines.

Abstract

Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent variable roles, and (2) kernel rewrites, probing whether a high-level proof plan survives a change of mathematical setting. Compared with existing benchmarks, GAP has two key benefits: (1) novel, likely unseen variants mitigate data leakage, and (2) performance across transformation families enables failure diagnosis, each transformation testing a hypothesis about the cause of failure. We instantiate GAP on all 1,051 William Lowell Putnam Competition problems from 1938 to 2024, adding 5,255 unseen variants to form PutnamGAP, a 6,306-item competition-level mathematics corpus and the first public machine-readable dataset from the full Putnam archive. Using PutnamGAP, we evaluated 18 commercial and open-source models spanning sizes and providers. Accuracy drops across all models and variant families, most severely under kernel rewrites. This gap does not close with model strength, suggesting that even the strongest models' dominant weakness is transferring a proof plan to a changed mathematical setting, rather than handling surface changes. Further analysis provides finer failure diagnoses and potentially useful insights for improving LLM reasoning.

Sources

Related papers