The Role of Dataset Linguistic Structure in the Cultural Awareness of Large Language Models
cs.CL
Submitted: 2026-02-01
Updated: 2026-09-21
Code: https://github.com/levelevel/AozoraTxt
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- ArabicaQA: A Comprehensive Dataset for Arabic Question Answering
- Investigating Cultural Alignment of Large Language Models
- Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations
- CIDAR: Culturally Relevant Instruction Dataset For Arabic
- COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
- Massively Multi-Cultural Knowledge Acquisition & LM Benchmarking
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- NativQA: Multilingual Culturally-Aligned Natural Query for LLMs
- C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
- On the Diversity of Synthetic Data and its Impact on Training Large Language Models
- Mistral 7B
- CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
- LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
- Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
- Improving Data Efficiency via Curating LLM-Driven Rating Systems
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
- Commonsense Reasoning in Arab Culture
- Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering