Evaluating the Social Impact of Generative AI Systems in Systems and Society
cs.CY, cs.AI
Submitted: 2023-06-09
Updated: 2024-06-28
Comments: This version has been removed by arXiv administrators as the submitter did not have the right to agree to the license at the time of submission
Journal ref: The Oxford Handbook of the Foundations and Regulation of Generative AI, 18 December 2025
DOI: 10.1093/oxfordhb/9780198940272.013.0025
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Generative AI systems across modalities, ranging from text (including code), image, audio, and video, have broad social impacts, but there is no official standard for means of evaluating those
Terminology
Abstract
Generative AI systems across modalities, ranging from text (including code), image, audio, and video, have broad social impacts, but there is no official standard for means of evaluating those impacts or for which impacts should be evaluated. In this paper, we present a guide that moves toward a standard approach in evaluating a base generative AI system for any modality in two overarching categories: what can be evaluated in a base system independent of context and what can be evaluated in a societal context. Importantly, this refers to base systems that have no predetermined application or deployment context, including a model itself, as well as system components, such as training data. Our framework for a base system defines seven categories of social impact: bias, stereotypes, and representational harms; cultural values and sensitive content; disparate performance; privacy and data protection; financial costs; environmental costs; and data and content moderation labor costs. Suggested methods for evaluation apply to listed generative modalities and analyses of the limitations of existing evaluations serve as a starting point for necessary investment in future evaluations. We offer five overarching categories for what can be evaluated in a broader societal context, each with its own subcategories: trustworthiness and autonomy; inequality, marginalization, and violence; concentration of authority; labor and creativity; and ecosystem and environment. Each subcategory includes recommendations for mitigating harm.
Sources
- The De-democratization of AI: Deep Learning and the Compute Divide in Artificial Intelligence Research
- Constitutional AI: Harmlessness from AI Feedback
- Actionable Guidance for High-Consequence AI Risk Management: Towards Standards Addressing AI Catastrophic Risks
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- On Hate Scaling Laws For Data-Swamps
- Into the LAIONs Den: Investigating Hate in Multimodal Datasets
- Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- Quantifying Memorization Across Neural Language Models
- Combating Misinformation in the Age of LLMs: Opportunities and Challenges
- Evaluating Large Language Models Trained on Code
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Recourse for reclamation: Chatting with generative language models
- But Who Protects the Moderators? The Case of Crowdsourced Image Moderation
- Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
- Energy Consumption of Deep Generative Audio Models
- Do Membership Inference Attacks Work on Large Language Models?
- What's In My Big Data?
- Ethics, Rules of Engagement, and AI: Neural Narrative Mapping Using Large Transformer Language Models
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework