Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure
cs.CL
Submitted: 2026-04-28
Updated: 2026-08-31
Comments: Accepted at EMNLP Findings 2026
Code: https://github.com/conversationai/conversationai.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers' dependence on a tool they did not control.
Terminology
Abstract
Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers' dependence on a tool they did not control. Drawing on this case, we argue that a research field must build and govern its own measurement infrastructure rather than borrow it. Surveying 241 papers that use or study Perspective, we show what depending on it cost the research community: claims reaching past what the tool could support, results that shifted when its model was silently retrained, and disparities researchers could measure but not explain. These failures were amplified throughout the LLM lifecycle, where Perspective supplied the labels, filtered the corpora, and graded the systems trained on each, rewarding errors rather than catching them. To keep the instrument open to study after shutdown, we release Perspective scores for 5.9 million text snippets from 77 datasets. We further specify ten requirements for measurement infrastructure a field owns, and argue that what blocks such infrastructure is not technical capability but the value the field places on infrastructure work.
Sources
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- The Accountability Paradox: How Platform API Restrictions Undermine AI Transparency Mandates
- Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
- Characterizing Twitter Users Who Engage in Adversarial Interactions against Political Candidates
- Towards Measuring Adversarial Twitter Interactions against Candidates in the US Midterm Elections
- Holistic Evaluation of Language Models
- Red Teaming Language Models with Language Models
- On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research
- The Enforcement and Feasibility of Hate Speech Moderation
- Elite political incivility is rising across democracies
- Evaluating Generative AI Systems is a Social Science Measurement Challenge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering