Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report
cs.CL, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-24
Code: https://github.com/ksekrst/provenance2026
Terminology
Sources
- Why Philosophers Should Care About Computational Complexity
- Constitutional AI: Harmlessness from AI Feedback
- Large Language Models Report Subjective Experience Under Self-Referential Processing
- Deep reinforcement learning from human preferences
- Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Inducing language models to assert their own consciousness restores human beliefs and values
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Emergent Introspective Awareness in Large Language Models
- Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
- Taking AI Welfare Seriously
- 2 OLMo 2 Furious
- Training language models to follow instructions with human feedback
- Towards Evaluating AI Systems for Moral Status Using Self-Reports
- The Two-Process Theory of Machine Self-Report
- Attention Is All You Need
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering