Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations
cs.CL, cs.LG
Submitted: 2026-07-20
Updated: 2026-09-18
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
- The Llama 3 Herd of Models
- Qwen2.5-Coder Technical Report
- Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
- Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
- Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
- DeepSeek-V3 Technical Report
- Qwen3.5-Omni Technical Report
- Mechanistically Eliciting Latent Behaviors in Language Models
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Automatically Interpreting Millions of Features in Large Language Models
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
- Steering Language Models With Activation Engineering
- Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
- Representation Engineering: A Top-Down Approach to AI Transparency
- MPNet: Masked and Permuted Pre-training for Language Understanding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering