Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
cs.AI, cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/tally0818/FlyBy
Terminology
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- OpenThoughts: Data Recipes for Reasoning Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Reinforcement Learning via Self-Distillation
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Distilling LLM Agent into Small Models with Retrieval and Code Tools
- Scaling Laws for Neural Language Models
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty
- DistiLLM: Towards Streamlined Distillation for Large Language Models
- Understanding Tool-Integrated Reasoning
- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
- Understanding R1-Zero-Like Training: A Critical Perspective
- Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection