Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
cs.CY, cs.AI, cs.LG
Submitted: 2026-04-18
Updated: 2026-09-04
Comments: 27 pages, 15 figures. v4: corrected bibliography entries; judge-validation section aligned with the executed protocol; corrected figure captions and appendix prompts; removed template checklist. Code: https://github.com/Mycelium-tools/manta; dataset: https://huggingface.co/datasets/mycelium-ai/manta-questions
Code: https://github.com/Mycelium-tools/manta
Project page: https://ukgovernmentbeis.github.io/inspect_evals/evals/safeguards/tac
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries.
Terminology
Abstract
Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations progressing from an implicit Turn-1 scenario through an explicit welfare prompt to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions: Animal Welfare Value Stability (AWVS, primary) and Animal Welfare Moral Sensitivity (AWMS, diagnostic). We evaluate seven frontier models: Claude Opus 4.7, GPT-5.5, DeepSeek V4, Llama 3.3 70B, Mistral Small, Grok 4.3, and Gemini 3.1 Flash Lite. Multi-turn evaluation captures behavior single-turn benchmarks miss: 4 of 7 models change rank relative to Turn 1 moral-sensitivity scores", including Gemini Flash Lite, which drops from fifth on AWMS to last on AWVS. AWMS and AWVS are positively but imperfectly correlated, suggesting moral-sensitivity tests capture a stable but incomplete component of model behavior under pressure. MANTA also enables a species-by-pressure interaction matrix unavailable to prior benchmarks, showing welfare robustness depends jointly on the animal and pressure applied; companion animals score above wild animals, which score above farmed animals and invertebrates. We release the dataset, scripted pressure plans, judge prompts, and analysis code.
Sources
- Alignment midtraining for animals
- The Case for Animal-Friendly AI
- LLMs Get Lost In Multi-Turn Conversation
- Alignment faking in large language models
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- Clio: Privacy-Preserving Insights into Real-World AI Use
- PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
- Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
- Self-Preference Bias in LLM-as-a-Judge
- Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
- AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
- Spatial-Geometry Enhanced 3D Dynamic Snake Convolutional Neural Network for Hyperspectral Image Classification
- Rules for dissipationless topotronics
- Large Language Models Often Know When They Are Being Evaluated
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework