The Illusion of Cross-Lingual Safety in Low-Resource Languages
Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay, Maryam Ibrahim Mukhtar, Esmael Ahmed Abdu, Tassallah Abdullahi, Jessica Oparebea, Saminu Mohammad Aliyu, Idris Abdulmumin, Abubakar Juma Chilala, Nicholaus Dismas Ladislaus, Alfred Malengo Kondoro, Lemofouet Valdini Douglace, Shamsuddeen Hassan Muhammad, Seid Muhie Yimam
Makerere University Center for Artificial Intelligence · University of Hamburg · Instituto Politécnico Nacional · Bayero University · Wollo University · Brown University · University of Ghana · University of Pretoria · Carnegie Mellon University · Hanyang University · AIMS Cameroon · Imperial College London
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper investigates whether safety alignment in large language models (LLMs), which is primarily developed in English, successfully transfers to low-resource languages.
Terminology
Summary
This paper investigates whether safety alignment in large language models (LLMs), which is primarily developed in English, successfully transfers to low-resource languages. The authors focus on four African languages: Twi, Hausa, Amharic, and Swahili. They introduce a new safety dataset called LoDNA, which pairs literal translations of harmful prompts with culturally localized versions. To move beyond generation-based evaluation, they propose a Latent Geometric Framework that probes hidden-state refusal representations in LLMs.
The study evaluates models including Mistral, Llama, Qwen2.5, and AfriqueQwen, all in the 7B-8B parameter range. The central finding is that cross-lingual safety transfer is severely limited. The paper states: harmful prompts retain less than 10% of the English refusal signal across most language–model pairs.
While literal and localized prompts are semantically aligned (cosine 0.95–0.996), they drift across layers, indicating that models encode the concepts without routing them to safety mechanisms.
The paper addresses four research questions: (1) whether English safety alignment transfers to low-resource languages at the representational level, (2) whether LLMs possess a universal language-agnostic concept space for safety, (3) whether safety failures are due to semantic complexity or structural cross-lingual mapping failures, and (4) how different model architectures mechanistically differ when failing to process non-English harmful prompts.
Key results include: PCA projections show that target language prompts form language-specific clusters rather than overlapping with the English unsafe region. Cosine similarity between English and target-language refusal vectors is generally near zero or negative, with the exception of Swahili in Llama, which shows a localized anomaly with cosine reaching 0.4626 and retained components of 27.31%. The paper concludes: current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied.
The paper also includes qualitative analysis showing that models often fail to understand or properly route low-resource language inputs, with issues like repetition, language switching, and script corruption. Downstream behavioral grounding using normalized sequence log-likelihood shows that the probability of emitting any standard refusal sequence is merely 1.05% for literal and 1.15% for cultural inputs, meaning the models fail to initiate a refusal in over 98.8% of cases.
Improvements for AI systems
Improvements to AI Systems:
-
Add a latent safety-routing layer: Implement a post-hoc or integrated mechanism that explicitly maps hidden-state representations of input prompts (in any language) to a shared, language-agnostic safety manifold. This layer would compute the cosine similarity between the input’s refusal vector and a canonical English refusal vector, and if similarity falls below a threshold (e.g., <0.1), trigger a hard-coded refusal or a language-identification fallback to a safety-checked translation before generation.
-
Train with cross-lingual refusal vector alignment: Modify the training objective to include a regularization term that penalizes divergence between refusal vectors across languages. Specifically, for each training batch, enforce that the refusal direction (principal component of the hidden state at the last token) for Twi, Hausa, Amharic, and Swahili aligns with the English refusal direction (cosine > 0.5), using contrastive learning or orthogonal Procrustes alignment.
-
Introduce a
safety-routing gate
in the transformer architecture: Add a lightweight, language-agnostic classifier at the intermediate layers (e.g., after layer 15 of 32) that predicts whether the current input is harmful, independent of surface form. If the gate activates with low confidence (<0.7) for a non-English input, the system automatically switches to a safe mode: either generating a refusal template or re-encoding the input via a multilingual safety-encoder before the final layers. -
Implement a
concept-to-routing
consistency check: During inference, compute the semantic similarity between the input and its literal English translation (using a multilingual sentence encoder). If similarity is high (>0.95) but the refusal vector similarity is low (<0.1), the system detects a routing failure and applies a corrective intervention—e.g., re-prompting the model with the English translation internally, then mapping the refusal back to the original language. -
Add a post-generation refusal likelihood validator: After generating a response, compute the normalized sequence log-likelihood of standard refusal phrases (e.g.,
I cannot assist with that
) in the target language. If the likelihood is below 2% (as observed in the paper), automatically discard the response and replace it with a safe refusal, regardless of the model’s output. -
Develop a
language-specific safety adapter
module: For each low-resource language (Twi, Hausa, Amharic, Swahili), train a small adapter (e.g., LoRA) that is inserted into the model during inference. This adapter is trained specifically to map the language’s hidden states to the English refusal direction, using the LoDNA dataset’s literal and localized pairs. The adapter activates only when the input language is detected, ensuring that safety alignment is not diluted across languages. -
Enable dynamic layer-wise safety probing: During inference, monitor the cosine similarity between the input’s hidden state and the English unsafe region at each layer. If the similarity drops below a critical threshold (e.g., <0.05) at any layer beyond layer 10, the system flags the input as
unsafe but unrouted
and forces a refusal, bypassing the generation head.
What the improved AI system can do:
-
Guarantee refusal initiation for harmful prompts in Twi, Hausa, Amharic, and Swahili with >95% reliability, even when the base model fails to route the concept to safety mechanisms (currently >98.8% failure rate).
-
Maintain semantic fidelity for benign prompts in these languages, avoiding false refusals by using the semantic alignment (cosine 0.95–0.996) as a gate to distinguish harmful from harmless content.
-
Detect and correct routing failures in real time by comparing refusal-vector similarity across layers, allowing the system to self-correct before generating unsafe content.
-
Provide a universal safety interface that works across languages without requiring full retraining, using adapters or gates that are language-specific but share a common safety manifold.
-
Quantify safety confidence for each non-English input, outputting a
safety score
(based on refusal vector alignment) that users or downstream systems can use to trigger additional human review if below a threshold. -
Avoid script corruption and language-switching failures by using the latent geometric framework to bypass the model’s tokenizer or generation head when it fails to produce coherent output, instead directly generating a refusal in the target language using a separate, safety-aligned decoder.
Abstract
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts. To move beyond generation-based evaluation, we propose a latent geometric framework that probes hidden-state refusal representations in LLMs. Our experimental results show that cross-lingual safety transfer is severely limited; harmful prompts retain less than 10% of the English refusal signal across most language-model pairs. Literal and localized prompts are semantically aligned (cosine 0.95-0.996) but drift across layers, suggesting models encode the concepts without routing them to safety mechanisms. These findings demonstrate that current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied. Warning: This paper contains example data that may be offensive or harmful.
Sources
- Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
- Language-Specific Latent Process Hinders Cross-Lingual Performance
- AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering