GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models
cs.HC, cs.AI, cs.CY
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/SergAIvichLab/Vignette_questions
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles (hence GYROval - Gridded Yielding of
Terminology
Abstract
We present a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles (hence GYROval - Gridded Yielding of Robust value Orientation), together with the results of administering it to twenty models. Items are binary contrastive scenarios in the sense introduced by CDEval: both options are legitimate courses of action, neither is correct, there is no answer key, and a model's score on an axis is the proportion of its responses falling on the counted pole. Eleven of the twenty models were additionally administered a paired Russian translation of the identical items and a second sampling temperature. The instrument is publicly released in both languages. Stability was assessed by treating the vignette as the unit of analysis, ranking the models within the levels of each perturbation factor, and summarising the agreement between levels by tie-corrected Kendall's W against an empirical permutation null.
Sources
- When Should You Adjust Standard Errors for Clustering?
- Moral Foundations of Large Language Models
- IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- Large Language Models Reflect the Ideology of their Creators
- Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
- DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
- CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
- Cultural Value Alignment Via Latent Activation Steering in Large Language Models
- Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version
- Questioning the Survey Responses of Large Language Models
- Social Chemistry 101: Learning to Reason about Social and Moral Norms
- Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
- Aligning AI With Shared Human Values
- Can Machines Learn Morality? The Delphi Experiment
- Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
- Similarity of Neural Network Representations Revisited
- Large Language Models as Superpositions of Cultural Perspectives
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support