Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

arXiv:2606.29985 · cs.CL · Submitted 2026-06-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning".

Jane: The paper investigates a critical gap in measuring diversity within Large Language Model (LLM) mathematical reasoning: conventional metrics capture only surface-level variation,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: The paper starts by introducing this concept of approach-level diversity as a way to fix that measurement problem. They define it as variation in underlying solution strategies across correct solutions for the same mathematical problem, going beyond just surface differences.

Jane: That makes sense, Tom; so instead of just counting how many different ways someone wrote down the answer, they are looking at the actual mental steps involved in getting there.

Lu: The authors establish this by breaking down approaches into three dimensions: mathematical tools used, how the problem is set up structurally, and the representational viewpoint taken.

Meng: That level of categorization sounds really useful for understanding *why* one solution works where another does not.

Lalam: It seems like this framework gives us a clearer lens to see what kind of reasoning capabilities an AI actually has, which is pretty valuable for how we build these things.

The paper's summary: Tom: Moving into the summary of "Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning," they really show how old metrics fail when comparing different reasoning strategies.

Jane: They provide a concrete example using a quadratic equation problem where conventional metrics like N-gram distance show high diversity for solutions that are actually very similar in their method, which is pretty confusing.

Lu: The paper points out that prior diversity measures, like N-gram distance and Self-BLEU, are susceptible to paraphrastic variation in math reasoning where changes in wording or layout can inflate the apparent diversity among solutions with the same reasoning approach.

Meng: So they're showing that these metrics can be tricked into thinking there’s more variety when there really isn't any fundamental shift in how the math is done.

Lalam: That confirms what we suspected; those surface-level metrics aren't reliable proxies for the deep structural differences in reasoning that matter.

The paper's improvements: Tom: Now, let’s talk about what the authors suggest to actually fix this situation. They propose using a human-calibrated LLM judge framework to formally define and validate approach-level diversity as strategy-level variation among correct solutions.

Jane: That sounds like a solid way to bridge the gap because it brings in human judgment, which is what we want our evaluations to reflect.

Lu: This framework involves filtering problems so they admit multiple approaches and then using an LLM judge, specifically GPT-five point two in their setup, to confirm if those different solutions are truly distinct in their strategy.

Meng: From an engineering standpoint, having a standardized way to label these approaches is crucial for scalable evaluation.

Lalam: This formalization of approach-level diversity as strategy-level variation really helps us move past just looking at text and start assessing the actual reasoning mechanism of the AI.

Conclusion: Tom: So, to wrap up on "Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning," the core message is that we need new ways to measure diversity that capture underlying strategy variation rather than just surface phrasing.

Jane: They conclude that this approach helps us see where existing diversity metrics fall short and points toward how we can better evaluate the reasoning capabilities of LLMs on mathematical tasks.

Lu: They also explore how this matters for Reinforcement Learning with Verifiable Rewards, showing that optimizing surface proxies doesn't guarantee approach-level diversity is preserved.

Meng: It seems like the practical implication is that if we want AI systems to be genuinely flexible in problem-solving, we have to train them to explore different strategies directly, not just reward surface variation.

Lalam: I think this whole paper gives us a strong direction for future work, suggesting we need objectives that actually preserve the structure of approach-level diversity while keeping things robust under training optimization.

Sangmook Lee, Minbeom Kim, Jeonghye Kim, Dohyung Kim, Sojeong Rhee, Kyomin Jung

Seoul National University · KAIST

cs.CL

Submitted: 2026-06-29

Updated: 2026-09-29

Comments: Accepted to EMNLP 2026 and ICML 2026 Ai4Math Workshop (Honorable Mention). Code at https://github.com/helmsman12/LLM_reasoning_diversity

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: The paper investigates a critical gap in measuring diversity within Large Language Model (LLM) mathematical reasoning: conventional metrics capture only surface-level variation, failing to

Key concepts

Approach-Level Diversity
Variation in the underlying solution strategies for the same mathematical problem. This goes beyond just how a solution is written down on the surface and looks at the actual mental steps involved in reaching the correct answer.
Surface-Level Metrics
Conventional metrics like N-gram distance or Self-BLEU that measure diversity. The paper shows these are unreliable because they can be tricked by paraphrasing, inflating diversity scores even when solutions use the same fundamental reasoning method.
Strategy Dimension
A way to categorize approaches into three dimensions: mathematical tools used, how the problem is set up structurally, and the representational viewpoint taken. This framework helps understand why one solution works while another fails.
Human-Calibrated LLM Judge Framework
A proposed method to fix measurement problems by using an LLM judge to formally define and validate approach-level diversity as strategy variation. This brings in human judgment to assess if different solutions are truly distinct in their reasoning strategy.

Terminology

Summary

The paper investigates a critical gap in measuring diversity within Large Language Model (LLM) mathematical reasoning: conventional metrics capture only surface-level variation, failing to distinguish between genuine differences in problem-solving strategies. This research introduces the concept of approach-level diversity—variation in underlying solution strategies—and demonstrates that existing diversity measures and RL post-training methods often conflate superficial changes in wording or notation with fundamental shifts in reasoning mechanisms. Understanding this mismatch is vital for developing LLMs that exhibit genuinely diverse, human-like strategic flexibility, which is crucial for downstream tasks like verifier-based selection and multi-agent collaboration.

Defining Approach-Level Diversity and its Limitations

The core contribution of the paper is the introduction of approach-level diversity as a distinct axis of LLM reasoning behavior: variation in the underlying solution strategies used to arrive at the correct answer, beyond differences in wording, notation, or exposition. The authors define different approaches based on three dimensions: Mathematical tools — the techniques invoked (e.g., algebraic vs. geometric), Structural definitions — how the problem is set up, and Representational viewpoint — the perspective taken. They establish this concept by constructing a human-calibrated reference set, finding that prior metrics like N-gram distance and cosine similarity are susceptible to paraphrastic variation in math reasoning, where changes in wording, layout, or symbolic expression can inflate the apparent diversity among solutions with the same reasoning approach.

Evaluating Conventional Metrics Against Human Judgments

The study systematically evaluates five baseline diversity metrics—N-gram distance (lexical), Self-BLEU (lexical), Cosine Distance (semantic), Distinct-Equations (symbolic), and Reasoning Path Decomposition (RPD)—against human judgments of approach diversity. The results reveal a systematic divergence: "Conventional metrics assign a higher diversity score to a pair of solutions that follow the same approach than to a pair that uses different mathematical approaches, illustrating a mismatch between surface-level variation and approach-level diversity. This failure is traced to two primary confounds: shared scaffolding (common setup steps) and approach-preserving paraphrasing" (changes in wording or notation that leave the underlying strategy unchanged).

Scaling Evaluation with an LLM Judge Framework

To move beyond surface metrics, the authors design a scalable evaluation framework centered around a human-calibrated LLM judge. This framework addresses three requirements: approach-feasible evaluation, scalable and interpretable labeling, and human-calibrated decision boundaries. They filter problems to ensure they admit multiple approaches by prompting models to generate distinct solutions and using a GPT-5.2 judge to confirm their mutual distinction. The LLM judge is shown to align well with human annotations, achieving 85.0% agreement with the human reference labels, formalizing approach-level diversity as strategy-level variation among correct solutions.

Impact on Diversity-Aware RLVR Methods

The paper examines how diversity-aware Reinforcement Learning with Verifiable Rewards (RLVR) methods, such as DQO and DIVER, handle this gap. They find that these methods typically optimize surface proxies—like lexical overlap or embedding distance—which does not imply that approach-level diversity is preserved. Specifically, optimizing these targets causes the policy to exploit judge-specific preferences rather than broaden its approaches, leading to a situation where training increases surface variation within a narrower set of approaches, making outputs look more diverse even as approach-level diversity declines.

Utility and Limits in Training and Inference

The research explores the practical utility of approach-level diversity. They find that Candidate sets containing distinct approaches yield larger gains under test-time scaling, suggesting that approach-diverse candidate sets can improve inferencetime performance. However, when directly optimized for the LLM judge's approach-diversity signal during training, it fails to broaden strategies. Furthermore, they demonstrate that preserving target metrics does not guarantee strategy preservation: DIVER maintains textual or equation-level diversity but still loses approach coverage. Finally, while approach-seeking supervised fine-tuning can improve accuracy in sequential settings, its gains are contingent on the alternative approaches being accessible to the target model, meaning it only helps when the necessary strategies are reachable. The paper concludes by identifying an open problem: designing approach-level diversity signals that capture genuine strategy-level differences and remain robust to reward hacking when optimized against.

Future Directions

The authors suggest two primary avenues for future research. First, useful diversity is likely to become increasingly domainspecific, necessitating domain-aware definitions and evaluations of approach diversity as LLMs move into complex tasks like scientific discovery. Second, approach-level diversity can be further studied as a training objective. The key challenge remains "designing objectives that preserve the human-relevant structure of approach-level diversity while remaining robust under direct optimization—a step toward training methods that encourage genuinely distinct reasoning rather than surface-level variation."

Improvements for AI systems

Here are specific, high-impact improvements for AI systems based on the findings of this research:


  1. Find a way to measure and optimize approach-level diversity rather than relying solely on surface-level metrics (like N-gram overlap or cosine similarity).

  2. Develop a new RL objective function that directly rewards policies for exploring genuinely different mathematical strategies across correct solutions, moving beyond proxies like embedding distance or equation counts.

  3. Implement a Approach-Level Diversity evaluation framework using an LLM judge to formally categorize and measure the underlying reasoning mechanisms of generated solutions.

This improved AI system can achieve the following:

  1. Find a way to measure and optimize approach-level diversity rather than relying solely on surface-level metrics (like N-gram overlap or cosine similarity).

  2. Develop a new RL objective function that directly rewards policies for exploring genuinely different mathematical strategies across correct solutions, moving beyond proxies like embedding distance or equation counts.

  3. Implement an Approach-Level Diversity evaluation framework using an LLM judge to formally categorize and measure the underlying reasoning mechanisms of generated solutions.

This improved system will enable:

These improvements will allow the AI system to:

Abstract

Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surface-level variation rather than differences in how a problem is solved. We address this gap by introducing approach-level diversity: variation in strategies across correct solutions to the same problem. Using a human-calibrated LLM judge framework, we show that prior diversity measures are unreliable proxies for approach-level diversity, and this mismatch carries over to diversity-aware RLVR, where target metrics are preserved while approach-level diversity declines. Investigating when approach-level diversity helps and whether it can be directly induced, we find that approach-diverse candidate sets improve test-time scaling. However, optimizing an LLM judge diversity reward during training causes the policy to exploit judge-specific preferences rather than broaden its approaches, leaving direct optimization of approach-level diversity as an open problem. Together, our work introduces the notion of approach-level diversity and uncovers a systematic divergence between surface- and approach-level signals, marking a step toward LLMs that reason in genuinely diverse, human-like ways.

Sources

Related papers