HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
cs.CY, cs.AI, cs.CL, cs.HC, cs.LG
Submitted: 2025-09-10
Updated: 2026-09-20
Comments: Accepted to EMNLP 2026 in the Findings track
Code: https://github.com/BenSturgeon/HumanAgencyBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures.
Terminology
Abstract
As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures. Relatively simple algorithmic systems already steer human decision-making, such as social media feed algorithms that lead people to unintentionally and absent-mindedly scroll through engagement-optimized content. In this paper, we develop the idea of human agency by integrating philosophical and scientific theories of agency with AI-assisted evaluation methods: using large language models (LLMs) to simulate and validate user queries and to evaluate AI responses. We develop HumanAgencyBench (HAB), a scalable and adaptive diagnostic tool for six behaviors related to human agency. HAB measures the tendency of an AI assistant to Ask Clarifying Questions, Avoid Value Manipulation, Correct Misinformation, Defer Important Decisions, Encourage Learning, and Maintain Social Boundaries. We find low-to-moderate agency support in contemporary LLM-based assistants, with substantial variation across system developers and behaviors. For example, while Anthropic LLMs most support human agency overall, they are the least supportive LLMs in terms of Avoid Value Manipulation. These behaviors do not appear to consistently result from increasing LLM capabilities or instruction-following (e.g., RLHF); we encourage further study of these behaviors so that developers and users can better understand the complexities of modern human-AI interaction.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- AI Consciousness and Public Perceptions: Four Futures
- The Ethics of Advanced AI Assistants
- Thousands of AI Authors on the Future of AI
- Validating LLM-as-a-Judge Systems under Rating Indeterminacy
- Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
- Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
- What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
- Two Types of AI Existential Risk: Decisive and Accumulative
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
- Discovering Agents
- Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety
- Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
- Discovering Language Model Behaviors with Model-Written Evaluations
- Evaluating Generative AI Systems is a Social Science Measurement Challenge
- Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise
- Toward an Evaluation Science for Generative AI Systems
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework