Efficient Reasoning Exploration via State-Conditioned Latent Steering with Progress Guidance
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/rattlesnakey/SPS
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- OpenAI GPT-5 System Card
- Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
- Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
- MMFormalizer: Multimodal Autoformalization in the Wild
- Qwen3 Technical Report
- Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
- Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
- Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
- Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
- Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- First Return, Entropy-Eliciting Explore
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering