DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
cs.CL, cs.LG
Submitted: 2026-01-11
Updated: 2026-09-01
Code: https://github.com/project-numina/aimo-progress-prize
Project page: https://dipta007.github.io/DAGGER
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot
- Language Models are Multilingual Chain-of-Thought Reasoners
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection
- MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
- Are NLP Models really able to Solve Simple Math Word Problems?
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning
- Qwen3 Technical Report
- Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering