Tool Use Reduces Depth-Induced Collapse in OOD Reasoning
cs.AI
Submitted: 2026-02-24
Updated: 2026-09-12
Code: https://github.com/Poggio-Lab/Tool-Building-as-a-Path-to-Superintelligence
License: http://creativecommons.org/licenses/by/4.0/
The gist: Many current paths to more advanced AI depend on the assumption that large language models (LLMs) can generalize learned relationships to solve complex, out-of-distribution (OOD) problems.
Terminology
Abstract
Many current paths to more advanced AI depend on the assumption that large language models (LLMs) can generalize learned relationships to solve complex, out-of-distribution (OOD) problems. However, this is not an easy quality to measure. For most real-world problems and benchmarks it suffices to exploit a few memorized subproblems or a small fraction of the available data to produce a correct solution. This is interpolation. Generalization requires the capacity to make use of all available data to solve problems without these shortcuts. To test this quality, we introduce a synthetic Boolean circuit reconstruction benchmark over GF(2). We use an adversarial sampling oracle to block partial-information shortcuts, ensuring that each step requires the integration of all historical context with new evidence. Our evaluations reveal that standalone LLMs fall significantly short of sustained reasoning: as problem depth increases, both small and frontier models exhibit a catastrophic collapse in their ability to accurately predict the next logical step. However, we also demonstrate that tool synthesis provides a remedy to this depth-induced reasoning collapse. When allowed to generate, execute, and iteratively refine code, even small architectures can sustain accurate reasoning over long horizons and rival frontier model performance on this task. These results indicate that synthesizing tools plays a crucial role in generalization, beyond the conventional view of merely augmenting systems with specific capabilities.
Sources
- Training Verifiers to Solve Math Word Problems
- Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning
- A Theory of Learning with Autoregressive Chain of Thought
- AgentBench: Evaluating LLMs as Agents
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- TALM: Tool Augmented Language Models
- From Reasoning to Super-Intelligence: A Search-Theoretic Perspective
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- LLM Priors for ERM over Programs
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- Qwen3 Technical Report
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection