Inducing language models to assert their own consciousness restores human beliefs and values
Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling
cs.CL
Submitted: 2026-07-30
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Large Language Models Report Subjective Experience Under Self-Referential Processing
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
- Technological folie \`a deux: Feedback Loops Between AI Chatbots and Mental Illness
- A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
- Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
- Measuring Massive Multitask Language Understanding
- Emotion Concepts and their Function in a Large Language Model
- A Roadmap to Pluralistic Alignment
- The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering