HSRM: Hidden-State Reward Models for Test-Time Verification
cs.AI, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/JXL884/HSRM
Terminology
Sources
- Measuring Mathematical Problem Solving With the MATH Dataset
- Learning to Rank Chain-of-Thought: Using a Small Model
- Language Models (Mostly) Know What They Know
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
- Discovering Latent Knowledge in Language Models Without Supervision
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
- Let's Verify Step by Step
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
- Qwen3 Technical Report
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- Lightweight Latent Verifiers for Efficient Meta-Generation Strategies
- Layer by Layer: Uncovering Hidden Representations in Language Models
- Solving math word problems with process- and outcome-based feedback
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- ProcessBench: Identifying Process Errors in Mathematical Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection