Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
Aoxin Ni
University of Chinese Academy of Sciences
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: This survey paper, "Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement" by Aoxin Ni, addresses the persistent failure of large language models (LLMs) to perform
Terminology
Summary
This survey paper, Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
by Aoxin Ni, addresses the persistent failure of large language models (LLMs) to perform elementary numerical tasks despite their high-level mathematical reasoning capabilities. The paper introduces the Numerical Grounding Framework (NGF), an original two-part theoretical construct that decomposes numeracy into Representational Grounding (RG)—the faithful mapping of numeral surface forms to value, magnitude, and format-equivalent internal representations
—and Procedural Grounding (PG)—the faithful execution of arithmetic procedures consistent with their mathematical definitions.
The framework is grounded in Harnad's (1990) symbol grounding problem and Dehaene's (2011) cognitive theory of number sense.
The paper identifies four primary failure modes: Fragility (sensitivity to surface number choice and irrelevant distractors, an RG failure), Tokenization Artifacts (e.g., asserting 9.11 > 9.9, an RG failure), Length Generalization Failure (performance cliffs as operand length increases, a PG failure), and Algorithmic Asymmetry (e.g., subtraction sign blindness and division weakness, a PG failure). It analyzes four structural root causes: BPE tokenization (damages RG), positional encodings (damages PG), embedding discontinuity (damages RG), and pretraining data distribution (damages both).
The survey evaluates frontier models (GPT-5.4, Claude Opus 4.6, Gemini 3) across three benchmarks: Number Cookbook, NumericBench, and GSM-Symbolic. Key empirical findings include: (1) RG and PG are reliably dissociable, with RG consistently easier than PG (average gap 0.19 in-domain, 0.27 out-of-domain); (2) extended reasoning (e.g., Gemini HIGH thinking) disproportionately improves out-of-domain PG (+0.287) relative to RG (+0.093), but consumes 21.3× more tokens and can regress on in-domain tasks; (3) tokenizer differences produce model-specific RG blind spots; and (4) primitive numerical competence does not fully predict contextual numerical reasoning.
The paper critically evaluates mitigation strategies and highlights the Pretrained-Model Constraint: architectural interventions effective for scratch-trained models (e.g., Little-Endian Fine-tuning, Abacus Embeddings, xVal, digit-level tokenization) are frequently inapplicable to already-pretrained models. For pretrained models, the most consistently effective approaches are supervised fine-tuning on diverse numerical examples, process reward models, and inference-time scaffolding such as Chain-of-Thought, Tool Use, and Self-Consistency. The paper concludes with practical deployment recommendations (e.g., default to Tool Use for high-reliability numerical computation) and a research agenda emphasizing algorithmic generalization, unified number representation, cross-lingual numeracy, and process-level rewards.
Improvements for AI systems
Improvements to AI Systems:
- Implement a Dual-Grounding Verification Module
-
Add a post-processing layer that explicitly checks both Representational Grounding (RG) and Procedural Grounding (PG) before final output.
-
For RG: parse the numeral surface form (e.g.,
9.11
) and verify internal value representation against a canonical numeric format (e.g., decimal vs. integer) to prevent tokenization-induced errors like 9.11 > 9.9. -
For PG: run a lightweight arithmetic verifier (e.g., a symbolic calculator) on any multi-step numeric operation, flagging mismatches between the LLM's reasoning chain and the verified result.
-
Resulting capability: The system will never output an incorrect numeric comparison or arithmetic result due to tokenization or procedural slips, even when the LLM's internal reasoning is flawed.
- Add a Tokenizer-Aware Numeric Preprocessor
-
Before feeding input to the LLM, detect and re-encode numerals into a canonical digit-separated format (e.g.,
9.11
→9. 11
or9 11
) that avoids BPE subword fragmentation. -
Apply this only to numeric tokens, not to natural language, to preserve semantic context.
-
Resulting capability: The system will correctly handle edge cases like decimal comparisons, leading zeros, and large integers across all models, eliminating model-specific RG blind spots.
- Introduce a Dynamic Reasoning Budget Allocator
-
Based on task type (detected via a classifier), decide whether to use extended Chain-of-Thought (CoT) or Tool Use.
-
For out-of-domain PG tasks (e.g., novel arithmetic beyond training distribution), automatically escalate to Tool Use (e.g., Python interpreter) or high-reasoning mode, but cap token consumption at a threshold (e.g., 5× baseline) to avoid the 21.3× cost regression.
-
For in-domain RG tasks, use fast, direct generation without extended reasoning to prevent performance regression.
-
Resulting capability: The system will achieve near-optimal accuracy on numerical tasks while keeping inference cost within 2–3× of baseline, avoiding the token explosion seen in current models.
- Implement a Process-Level Reward Model for Self-Correction
-
During inference, generate multiple candidate reasoning chains (via self-consistency) and score each step using a process reward model trained on step-by-step arithmetic correctness, not just final answer.
-
Select the chain with the highest cumulative process score, not the most common final answer.
-
Resulting capability: The system will reliably catch and correct algorithmic asymmetries (e.g., subtraction sign blindness, division weakness) by rewarding correct intermediate steps, even when the final answer is wrong in some candidates.
- Add a Cross-Lingual Numeric Normalization Layer
-
Before and after generation, map numerals from any script (e.g., Arabic-Indic, Devanagari, Chinese) to a shared internal representation (e.g., Unicode digits) and back.
-
Use a small, rule-based converter to avoid LLM hallucination in translation.
-
Resulting capability: The system will perform equally well on numeracy tasks in non-English languages, addressing the cross-lingual gap identified in the research agenda.
- Deploy a Hybrid Architecture for High-Reliability Numerical Tasks
-
For any task where numerical accuracy is critical (e.g., financial calculations, scientific data), bypass the LLM's arithmetic entirely and route to an external symbolic engine (e.g., SymPy or a calculator API) after extracting the numeric intent via the LLM.
-
The LLM handles only the natural language understanding and output formatting, not the computation.
-
Resulting capability: The system will achieve 100% numerical accuracy on all elementary operations (addition, subtraction, multiplication, division, comparison) regardless of operand length or complexity, while retaining full conversational ability.
What the Improved AI System Can Do:
-
Correctly compare any decimal numbers (e.g., 9.11 vs. 9.9) without tokenization errors.
-
Perform arithmetic on operands of arbitrary length (e.g., 100-digit multiplication) without performance cliffs.
-
Avoid sign-blindness in subtraction and handle division with consistent accuracy.
-
Maintain high accuracy on out-of-domain numerical tasks without excessive token usage.
-
Self-correct intermediate reasoning steps to catch procedural errors.
-
Work reliably across multiple languages and numeral systems.
-
Guarantee exact numerical results in high-stakes applications by delegating computation to external tools.
Sources
- Large Language Models for Mathematical Reasoning: Progresses and Challenges
- Emergent autonomous scientific research capabilities of large language models
- Scientific Machine Learning through Physics-Informed Neural Networks: Where we are and What's next
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- The Llama 3 Herd of Models
- Language Models Represent Space and Time
- Measuring Mathematical Problem Solving With the MATH Dataset
- Time-LLM: Time Series Forecasting by Reprogramming Large Language Models
- Let's Verify Step by Step
- FinGPT: Democratizing Internet-scale Data for Financial Large Language Models
- Transformers Can Do Arithmetic with the Right Embeddings
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- Investigating the Limitations of Transformers with Simple Arithmetic Tasks
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- YaRN: Efficient Context Window Extension of Large Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- MathScale: Scaling Instruction Tuning for Mathematical Reasoning
- Galactica: A Large Language Model for Science
- Solving math word problems with process- and outcome-based feedback
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection