Thinking beyond the anthropomorphic paradigm benefits LLM research
cs.CL
Submitted: 2025-02-13
Updated: 2026-09-14
Terminology
Sources
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models
- Language Models as Agent Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Mechanistic Interpretability for AI Safety -- A Review
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- ChatBench: From Static Benchmarks to Human-AI Evaluation
- From tools to thieves: Measuring and understanding public perceptions of AI through crowdsourced metaphors
- Beyond Personhood: Agency, Accountability, and the Limits of Anthropomorphic Ethical Analysis
- The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms
- An Empirical Categorization of Prompting Techniques for Large Language Models: A Practitioner's Guide
- Cocoa: Co-Planning and Co-Execution with AI Agents
- Improving alignment of dialogue agents via targeted human judgements
- Alignment faking in large language models
- The R-U-A-Robot Dataset: Helping Avoid Chatbot Deception by Detecting User Questions About Human or Non-Human Identity
- Training Large Language Models to Reason in a Continuous Latent Space
- Instruction Following without Instruction Tuning
- Characterizing and modeling harms from interactions with design patterns in AI interfaces
- Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
- Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering