Thinking effort aligns between humans and reasoning models in abductive reasoning
cs.CL, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 14 pages, 5 figures. To appear in Findings of EMNLP 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Abductive Reasoning with Syllogistic Forms in Large Language Models
- Abductive Commonsense Reasoning
- Language Models are Few-Shot Learners
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V3 Technical Report
- Training Large Language Models to Reason in a Continuous Latent Space
- Distilling the Knowledge in a Neural Network
- The Curious Case of Neural Text Degeneration
- Large Language Models are Zero-Shot Reasoners
- Dissociating language and thought in large language models
- Frontier LLMs Still Struggle with Simple Reasoning Tasks
- s1: Simple test-time scaling
- gpt-oss-120b & gpt-oss-20b Model Card
- Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
- Wait, Wait, Wait... Why Do Reasoning Models Loop?
- Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering