Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/huggingface/Math-Verify
Terminology
Sources
- Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
- Recursive Agent Optimization
- FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Self-Refine: Iterative Refinement with Self-Feedback
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
- ATLAS: Agentic Test-time Learning-to-Allocate Scaling
- Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
- THREAD: Thinking Deeper with Recursive Spawning
- The Era of Agentic Organization: Learning to Organize with Language Models
- MEMENTO: Teaching LLMs to Manage Their Own Context
- Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
- SPIRAL: Learning to Search and Aggregate
- Scaling Long-Horizon LLM Agent via Context-Folding
- ReAct: Synergizing Reasoning and Acting in Language Models
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks