Interpreting and Steering LLM Agents for Social Simulations
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 70 pages, 27 figures, 3 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- LLM Social Simulations Are a Promising Research Method
- Refusal in Language Models Is Mediated by a Single Direction
- LLMs generate structurally realistic social networks but overestimate political homophily
- A Financial Brain Scan of the LLM
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- AI Behavioral Science
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- The Llama 3 Herd of Models
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Progress measures for grokking via mechanistic interpretability
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Steering Language Models With Activation Engineering
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks