Unknown Unknowns: Do Hidden Intentions in LLMs Evade Detection?
cs.CL, cs.LG
Submitted: 2026-01-26
Updated: 2026-08-31
Terminology
Sources
- Taxonomizing Representational Harms using Speech Act Theory
- Rethinking the Evaluation of Secure Code Generation
- Bias and Fairness in Large Language Models: A Survey
- Improving alignment of dialogue agents via targeted human judgements
- TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
- Evaluating Large Language Models with Psychometrics
- TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
- Auditing language models for hidden objectives
- Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction
- Measuring Stereotype and Deviation Biases in Large Language Models
- Benchmarking Gender and Political Bias in Large Language Models
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering