The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts
cs.CL
Submitted: 2026-09-21
Updated: 2026-10-02
Comments: 19 pages, 7 figures
Code: https://github.com/V1centNevwake/answer-basin-representation
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- The Geometry of Categorical and Hierarchical Concepts in Large Language Models
- The Information Geometry of Softmax: Probing and Steering
- Steering Language Models With Activation Engineering
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering