FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: accepted to EMNLP 2026 (Main)
Code: https://github.com/Yusser96/FineWeb-CLaRhttps:
Project page: https://behavior-in-the-wild.github.io/align-via-actions.html
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
- SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
- SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia
- CULEMO: Cultural Lenses on Emotion -- Benchmarking LLMs for Cross-Cultural Emotion Understanding
- Cultural Adaptation of Recipes
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
- TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- MASIVE: Open-Ended Affective State Identification in English and Spanish
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- CARE: Multilingual Human Preference Learning for Cultural Awareness
- Hope Speech detection in under-resourced Kannada language
- KoBBQ: Korean Bias Benchmark for Question Answering
- CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
- BnMMLU: Measuring Massive Multitask Language Understanding in Bengali
- Are All Spanish Doctors Male? Evaluating Gender Bias in German Machine Translation
- LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering