Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA
cs.CL
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Concrete Problems in AI Safety
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
- OpenAI o1 System Card
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Let's Verify Step by Step
- ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
- Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Defining and Characterizing Reward Hacking
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Solving math word problems with process- and outcome-based feedback
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Qwen3 Technical Report
- TTCS: Test-Time Curriculum Synthesis for Self-Evolving
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering