From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
Si'an Xie, Jiaxun Liu, Biao Yang, Wei Yuan, Fan Yang, Tingting Gao, Ming Wu
Beijing University of Posts and Telecommunications · Peking University · Kuaishou Technology
cs.CL, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper introduces MPAR-Bench, a bilingual (English–Chinese) benchmark designed to evaluate "reasoning breadth" in large language models (LLMs), as opposed to the extensively studied "reasoning
Terminology
Summary
The paper introduces MPAR-Bench, a bilingual (English–Chinese) benchmark designed to evaluate reasoning breadth
in large language models (LLMs), as opposed to the extensively studied reasoning depth.
The authors argue that while modern LLMs have achieved remarkable success in step-by-step linear reasoning through techniques such as Chain-of-Thought prompting, reinforcement learning, and supervised finetuning, a complementary capability remains largely unexplored: the ability to aggregate dispersed semantic signals and perform abstract conceptual convergence.
The paper states: An important question remains unresolved: do current LLMs possess the non-linear, cross-domain associative reasoning abilities?
The authors note that existing benchmarks "primarily emphasize reasoning 'depth,' evaluating step-by-step logical deduction and procedural reasoning. In contrast, the evaluation of reasoning 'breadth'—the ability to aggregate dispersed semantic signals and perform abstract conceptual convergence—remains largely unexplored."
MPAR-Bench is inspired by the cooperative board game Just One, where players give single-word hints to help a guesser infer a hidden target word, while direct synonyms, translations, homophones, and duplicate clues are forbidden. The paper explains: "These constraints are what make the game a clean instrument for breadth. Rather than rewarding the most obvious lexical association, they force clue writers to approach the target from distinct, indirect angles; the guesser must then integrate fragmented, non-overlapping signals rather than pattern-match a single cue."
The task is formally defined as: "Given a clue set C = c1, c2,..., cn, the task is to recover a target y such that each clue ci contributes an independently informative semantic relation to y. Reasoning breadth, in this setting, is the ability to integrate multiple semantically distinct and non-redundant clues into a single coherent answer."
The benchmark contains 1,000 items (500 English, 500 Chinese) constructed through a multi-stage pipeline:
-
Answer space: Target words are drawn from public word lists—RAT-derived vocabulary and Just One word cards.
-
Multi-agent clue generation: "LLM-based agents iteratively propose new clues that remain semantically relevant to the target while minimizing redundancy with existing clues. Each agent is assigned a distinct association angle to encourage coverage across semantic directions. A judge agent then removes clues that are the answer itself, direct synonyms/translations/homophones/morphological variants, exact or near-duplicates of accepted clues, or genuinely low quality."
-
Embedding-based diversity filtering: Using Qwen3Embedding-8B, the pipeline scores clue–answer and clue–clue similarity,
discarding clues that are either trivially close to the answer or too weakly related to be informative, and clue pairs that are near-duplicates.
-
Answer uniqueness verification: Human verification on 250 randomly sampled items showed
92.8% of the items were judged as having a unique answer
(95% Wilson CI: [88.6%, 95.4%]). -
Bilingual design:
The English subset emphasizes lexical and abstract associations; the Chinese subset additionally incorporates idioms, character-level and pictographic properties, and contemporary cultural memes.
Crucially, the authors emphasize: only the answer space is drawn from public word lists, whereas every clue set is generated from scratch,
substantially reducing memorization risk compared to public RAT items.
The paper proposes a coarse-to-fine evaluation framework with three levels:
-
Accuracy: Exact-match accuracy where
A prediction is considered correct if and only if it exactly matches the ground-truth answer.
-
Word-based similarity: Two metrics:
-
ANLS (Average Normalized Levenshtein Similarity): Computed as ANLS(ŷ, y) = 1 − d lev(ŷ, y) / max(ŷ, y)
-
Word embedding similarity: Using fastText embeddings, computed as cosine similarity between prediction and ground truth embeddings.
- Reasoning trace evaluation: "We assess the validity of intermediate processes via reasoning trace evaluation, decomposed into two dimensions: logical verification and factual verification. Logical verification examines whether reasoning trajectories follow coherent inferential steps from clues to predictions, while factual verification checks whether intermediate claims are factually grounded.
Human review of 300 reasoning traces confirmed
high consistency (98.7% on factual verification, 94.7% on logical verification) between human and LLM judgement."
The paper introduces four perturbation types to test robustness:
-
Clue Masking:
Randomly masking clues to evaluate model reasoning ability under information deficiency.
-
Order Shuffling:
Shuffling the clues order to find whether the model's reasoning process is sensitive to order.
-
Distractors:
Injecting semantically misleading or irrelevant cue words to test model resistance to noisy contexts and spurious correlations.
-
Multi-step Inferring:
Increasing the associative semantic distance between clues and the mystery word, forcing models to generate intermediate latent connections rather than relying on direct surface co-occurrence.
Models evaluated: GPT-5.2, Gemini-3.1pro, Gemini-3flash, Sonnet-4.5, Qwen3-max, Kimi-k2, Deepseek-v3.2, Seed-2-pro, plus locally deployed Qwen3 models (0.6B to 32B parameters).
Standard setting results (thinking mode):
-
Gemini-3.1pro leads on both English (86.8%) and Chinese (72.2%)
-
GPT-5.2 achieves 77.6% English, 64.4% Chinese
-
Sonnet-4.5 achieves 79.0% English, 67.4% Chinese
Standard setting results (non-thinking mode):
-
Sonnet-4.5 achieves highest accuracy: 70.4% English, 68.8% Chinese
-
Gemini-3flash: 70.0% English, 67.0% Chinese
Enhanced setting results: Perturbations reduce accuracy by 9–18 points in English and 5–12 points in Chinese.
For example, in thinking mode, Deepseek-v3.2 shows the largest English decline, while Kimi-k2 drops most in Chinese.
Thinking vs. Non-Thinking: "Thinking mode consistently improves most indicators, but the magnitude of the gain is markedly larger on English than on Chinese... thinking lifts English accuracy by a substantially wider margin and produces clear, stable gains for every model, whereas its effect on Chinese is much smaller and model-dependent (Sonnet-4.5 even shows a slight regression)."
1. Reasoning breadth remains far from solved: The best models reach 86.8%/72.2% accuracy in English/Chinese, with perturbations causing 9–18/5–12 point drops.
2. Greater reasoning depth does not automatically confer breadth: Thinking mode improves standard-setting accuracy but does not consistently reduce perturbation sensitivity, and case-level analysis reveals that extended reasoning can override correct answers through overthinking.
3. Overthinking is a notable failure pattern: Models initially arrive at the correct answer but subsequently override it during extended reasoning, often drifting toward a semantically related but incorrect concept.
For example, given the answer word Philosophy, the model outputs Plato, over-focusing on a representative entity implied by the clues rather than the academic discipline itself.
This is particularly pronounced in Qwen3-max and Kimi-k2.
4. Information gain curve: Using Seed-2-pro, the paper shows "accuracy consistently improves as the number of clue words increases... However, the marginal gain progressively slows down, indicating that while models benefit from richer semantic context, they saturate beyond a certain evidence threshold."
5. Scaling laws: Except for Qwen3-32B, accuracy consistently improves with increasing model size.
The exception is attributed to overthinking: Case-level inspection finds that Qwen3-32B suffers from overthinking, causing it to reject correct answers during an extended reasoning process.
6. Feedback mechanism: Iterative semantic feedback (ANLS and embedding similarity) shows "LLMs are not fully sensitive to word embedding similarity as a guidance signal... the revision trajectory remains weakly aligned with semantic indicators, suggesting that models do not naturally exploit surface-level semantic proximity for iterative refinement."
7. Structured reasoning skill: A three-step prompting intervention (examine all clues, prioritize specific concepts, reverse verification) yields only marginal gains: +1.0pp
on English accuracy and +3.2pp
on Chinese accuracy, suggesting prompt-level interventions alone offer limited leverage on the core challenge of multi-source evidence integration.
8. Reasoning trace error analysis: Across all models, logical error rates substantially exceed factual error rates.
In English thinking standard setting, logical error rates range from 20.61% (Gemini-3.1pro) to 43.60% (Deepseek-v3.2), while factual error rates are considerably lower (11.52%–15.77%). This indicates that the primary failure in multi-point associative reasoning is not incorrect factual knowledge but rather invalid inferential jumps in constructing the reasoning chain from clues to the answer.
9. Perturbation type analysis: "Order Shuffling consistently yields the highest accuracy across all models, some of which are even higher than Standard setting. In contrast, Clue Masking causes the most severe degradation in English, with an average drop of 20.0% from Order Shuffling. Distractor Injection also substantially reduces performance... indicating that spurious semantic correlations introduced by irrelevant clue words can effectively derail the model's reasoning trajectory."
The paper concludes: "We introduced MPAR-Bench, a bilingual benchmark that evaluates multi-point associative reasoning—reasoning breadth—in LLMs through 1,000 boardgame-rule-based questions and a coarse-to-fine evaluation protocol spanning accuracy, ANLS, embedding similarity, and reasoning-trace verification, complemented by a four-axis perturbation suite."
The three main findings are: "(1) reasoning breadth remains far from solved: the best models reach 86.8%/72.2% accuracy in English/Chinese, with perturbations causing 9–18/5–12 point drops. (2) greater reasoning depth does not automatically confer breadth: thinking mode improves standard-setting accuracy but does not consistently reduce perturbation sensitivity, and case-level analysis reveals that extended reasoning can override correct answers through overthinking. (3) improving breadth appears challenging: scaling model size, adding reasoning strategies, and iterative feedback each bring only partial gains, suggesting that reasoning breadth may be a capability that current training paradigms do not naturally optimize for."
The authors release MPAR-Bench to encourage the community to move beyond depth-oriented evaluation and toward a more complete picture of reasoning—one that values breadth as much as depth.
Improvements for AI systems
Based on this paper, here are specific improvements to AI systems and what the improved systems can do:
Improvement: Implement a confidence-checking layer that monitors reasoning trajectories for late-stage answer overrides. When a model initially produces a high-confidence answer and then revises it during extended reasoning, the system should compare the semantic distance between the initial and revised answers. If the revision drifts toward a related-but-incorrect concept (e.g., Philosophy
→ Plato
), the system should flag this as a potential overthinking failure and either revert to the initial answer or require explicit justification for the change.
What it can do: Prevent accuracy degradation in tasks requiring associative reasoning, especially in models like Qwen3-32B and Kimi-k2 that show overthinking patterns. The system would maintain correct answers that would otherwise be lost during extended reasoning.
Improvement: Train models with explicit order-shuffling augmentation during fine-tuning, where each training example is presented in multiple random clue orders. Additionally, add a permutation-invariant attention mechanism that aggregates clue information symmetrically, ensuring the model's internal representation of the clue set does not depend on presentation order.
Improvement: Add a pre-processing module that scores each input clue for semantic coherence with the overall clue set before reasoning begins. Clues that are statistical outliers (low average similarity to other clues) would be down-weighted or explicitly labeled as potential distractors. This module could use embedding-based clustering to identify and isolate spurious signals.
Improvement: Instead of a single reasoning chain, generate multiple parallel reasoning paths, each starting from a different subset or ordering of clues. Use a voting mechanism where the final answer must be supported by at least two independent reasoning paths. This ensemble approach directly targets the reasoning breadth
gap by forcing the model to integrate evidence from multiple angles rather than committing to a single linear trajectory.
Improvement: During iterative refinement, explicitly incorporate embedding-based similarity scores between the current prediction and the clue set as a training signal. The model should be trained to recognize when its current answer has low semantic alignment with the clues and use this as a trigger for revision, rather than relying solely on self-generated feedback.
Improvement: Create a shared reasoning layer that operates on abstract semantic representations independent of language. Train this layer on both English and Chinese associative reasoning tasks, with explicit cross-lingual alignment loss to ensure that reasoning strategies learned in one language transfer to the other.
Improvement: Implement a mechanism that tracks the marginal information gain as additional clues are processed. When the model's confidence improvement from new clues falls below a threshold (indicating saturation), it should finalize its answer rather than continuing to process potentially confusing additional evidence.
Improvement: Add a dedicated verification module that checks each intermediate reasoning step for logical validity (not just factual correctness). This module would flag steps where the model makes inferential jumps without sufficient evidence from the clues, and either request re-reasoning or reduce confidence in the final answer.
Sources
- Training Verifiers to Solve Math Word Problems
- Qwen3 Technical Report
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Kimi K2: Open Agentic Intelligence
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- Modeling Associative Reasoning Processes
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- CREATE: Testing LLMs for Associative Creativity
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Scaling Laws for Neural Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering