AwarenessBench: Assessing Cognitive Capabilities of Language Models
cs.CL, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- GPT-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Tell me about yourself: LLMs are aware of their learned behaviors
- Centaur: a foundation model of human cognition
- Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
- Pragmatic Reasoning improves LLM Code Generation
- Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks
- ToMBench: Benchmarking Theory of Mind in Large Language Models
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Dynamic Planning with a LLM
- Self-Recognition in Language Models
- Misusing Tools in Large Language Models With Visual Adversarial Examples
- Alignment faking in large language models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models
- On the link between conscious function and general intelligence in humans and machines
- "A Woman is More Culturally Knowledgeable than A Man?": The Effect of Personas on Cultural Norm Interpretation in LLMs
- Theory of Mind for Multi-Agent Collaboration via Large Language Models
- Knowledge Boundary of Large Language Models: A Survey
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering