Identifying and Transferring Reasoning-Critical Neurons: Improving LLM Inference Reliability via Activation Steering
cs.CL
Submitted: 2026-01-27
Updated: 2026-09-12
Comments: Accepted to EMNLP 2025 main conference
Code: https://github.com/cat-sk/AdaRAS
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program Synthesis with Large Language Models
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
- Training Verifiers to Solve Math Word Problems
- Expanded Gating Ranges Improve Activation Functions
- OpenAI o1 System Card
- Process Reinforcement through Implicit Rewards
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
- START: Self-taught Reasoner with Tools
- OpenThoughts: Data Recipes for Reasoning Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Inspecting and Editing Knowledge Representations in Language Models
- Self-critiquing models for assisting human evaluators
- Investigating Bias Representations in Llama 2 Chat via Activation Steering
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering