Low-Resource Safety Failures Are Action Failures, Not Representation Failures

arXiv:2606.01196 · cs.CL, cs.AI · Submitted 2026-05-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Low-Resource Safety Failures Are Action Failures, Not Representation Failures".

Jane: The paper investigates why safety alignment learned in high-resource languages fails to transfer effectively to low-resource languages,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Hey Jane. So we’re looking at this paper today from arXiv. It’s called "Low-Resource Safety Failures Are Action Failures, Not Representation Failures," and it looks like they’re tackling a real problem in making AI safe across different languages.

Jane: That sounds really specific, Tom. Basically, the issue is that safety alignment learned in rich languages just doesn't carry over well when you translate those safety checks into languages that aren't well-represented in the training data.

Lu: Exactly. The core idea they’re pushing is that the problem isn't a missing concept of what’s harmful, but rather how the model decides to refuse based on what it sees in a new language.

Meng: So, if I'm getting it right, they’re suggesting that the underlying safety understanding is actually there, even in those low-resource languages.

Tom: That's right, Meng. The paper looks at models like Qwen2 point 5-7B and Llama-three point one-8B across twenty-three different languages and finds that the harmfulness direction extracted from high-resource data still separates harmful prompts from harmless ones almost as well as it does in English or French <ref:2606.01196#pg0,the harmfulness direction extracted from high-resource>.

Jane: But even though that representation is present, the actual refusal rate drops quite a bit. They found that harmful refusal goes down from eighty-seven point nine percent to just forty-three point nine percent when the prompts are translated into languages like Swahili or Burmese, which is a big gap for real-world deployment <ref:2606.01196#pg0>.

Lu: It points to a failure in the calibration of the safety decision, not a failure in the representation itself. They observed that low-resource prompts produce lower projections onto that harmfulness direction, so the model’s internal threshold for refusal doesn't trigger reliably <ref:2606.01196#pg1>.

Tom: So what do they propose to fix this calibration issue without retraining the whole thing? That’s the big question for this paper, right?

Jane: They introduce something called a few-shot latent gate that uses that existing high-resource signal and recalibrates the decision threshold using just a few examples in the target language. It avoids needing to learn an entirely new harmfulness direction from scratch <ref:2606.01196#pg1>.

Meng: From an engineering standpoint, reusing the signal makes sense if we don't want to waste massive amounts of data just to define a new safety vector for every language. How does that actually work in practice?

Title and authors: Lu: The gate acts as a switch between steering toward refusal and looking at the harmfulness direction ablation, which substantially raises what they call mean refusal selectivity from thirty-three point six up to fifty-four point five for their strongest adapted baseline <ref:2606.01196#pg2>.

Tom: Fifty-four point five is a noticeable jump in selectivity, which means the model starts refusing much more often than it was before this intervention <ref:2606.01196#pg2>. So, what’s the final evaluation of this few-shot latent gate?

Jane: They tested it using a few-shot threshold calibration method where they picked the threshold to maximize the F1 score on both harmful and harmless training examples <ref:2606.01196#pg5>. The results show that when using zero-shot HRL routing, they raise the mean delta from fifteen point three up to thirty-eight point six on Qwen and up to seventy-five point six on Gemma <ref:2606.01196#pg8>.

Meng: It sounds like this method is very selective, achieving the best delta in seventeen out of eighteen language pairs <ref:2606.01196#pg8>. That level of performance across different models is interesting for practical deployment.

Lu: What I find really compelling is that they show this works well even when you combine it with a thirty-two low-resource language setup, pushing the mean delta from thirty-eight point six to forty-two point four for Qwen and Gemma <ref:2606.01196#pg8>. It shows robustness in combining the steering mechanism with specific target data.

Tom: So, to summarize this paper on "Low-Resource Safety Failures Are Action Failures, Not Representation Failures," the main point is that we don't need to retrain the model when it struggles with low-resource languages; we just need to recalibrate its decision threshold using a few examples of the target language.

Jane: It really shifts our perspective from thinking about teaching a new concept to simply fixing how the existing concept is being used by the model. The real work here is in diagnosing where that calibration breaks down <ref:2606.01196#pg2>.

Lu: And they’ve laid out a pretty clear diagnostic workflow: you test whether harmfulness is represented before deciding how to repair its use, which is a very practical path forward for safety research <ref:2606.01196#pg9>.

Meng: It makes the process much more manageable for engineers. Instead of needing huge datasets for every single language just to define a new safety direction, they suggest reusing what’s already there and just tuning the final switch <ref:2606.01196#pg2>.

Tom: This feels like it moves us away from the idea that we need perfect multilingual alignment from day one and toward a diagnostic-then-fix approach for specific language gaps. It’s a very grounded way to look at it <ref:2606.01196#pg9>.

Title and authors: Jane: It gives us hope that we can actually address these safety disparities without having to overhaul the entire model architecture or spend months on massive retraining cycles for every new language we want to support.

Lu: The implication is that safety alignment becomes much more scalable when you focus on fine-tuning the output mechanism rather than trying to perfectly map all semantic concepts across every single linguistic boundary <ref:2606.01196#pg2>.

Meng: So for those of us building these systems, it means we can probably prioritize targeted calibration efforts on the most problematic low-resource languages instead of spreading our resources thin trying to solve every language problem at once <ref:2606.01196#pg8>.

Tom: Exactly. We have a lot to think about as we move into the next set of papers, but this one really hammers home that calibration is often the weak link in multilingual systems when it comes to safety outcomes <ref:2606.01196#pg0>.

Jane: So, we’ve seen how they diagnose the gap and how they propose a way to fix it by re-tuning the gate. That gives us a solid framework for future work <ref:2606.01196#pg9>.

Lu: I'm curious to see if other models or different alignment techniques can use this same few-shot latent gate mechanism effectively, because the underlying idea of reusing high-resource signals is very powerful for generalization <ref:2606.01196#pg2>.

Meng: From my side, I’m looking at how easily this setup integrates into existing deployment pipelines. If it can be done with just a few examples, that’s a huge win for operational efficiency <ref:2606.01196#pg8>.

Tom: That's the practical angle we need to keep in mind as we look ahead at how this kind of targeted recalibration can be applied across different safety modalities and tasks <ref:2606.01196#pg9>.

Jane: It’s a really interesting piece of work that gives us a concrete intervention to address the gap between high-resource and low-resource AI safety performance.

Lu: We’ll keep an eye on how this few-shot latent gate performs when applied to even more complex, nuanced harmfulness directions in future research <ref:2606.01196#pg2>.

Meng: And I’ll be checking if we can adapt this idea for other types of model steering, not just safety refusal <ref:2606.01196#pg8>.

Tom: That’s all the time we have for this deep dive into "Low-Resource Safety Failures Are Action Failures, Not Representation Failures." We’ll be back next time with some new arXiv papers.

The paper's summary: Tom: So we’re looking at this paper today from arXiv titled "Low-Resource Safety Failures Are Action Failures, Not Representation Failures."

Jane: Basically, they found that when AI safety systems fail in languages that aren't well-represented in training data, it’s not because the model doesn't understand what’s harmful at all.

Tom: That’s right. The core finding is that the underlying knowledge of what to avoid is actually present in those models across different languages.

Lu: They tested the harmfulness direction and found it stays linearly separable, meaning the concept of harm is there in the activations, even when you use a low-resource prompt.

Meng: But they noted that because those low-resource prompts produce weaker signals for that harm direction, the model just doesn't trigger its refusal mechanism reliably.

Tom: So the failure isn't about missing a representation; it’s about the calibration of how the model decides to act on that representation.

Jane: That’s a huge distinction. It means we don't have to worry about trying to teach the model a new concept for every single language just because its data is thin.

Lu: They propose this few-shot latent gate, which essentially reuses that high-resource harmfulness signal and recalibrates the decision threshold with just a few examples of the target language.

Meng: So it’s not about retraining the whole model; it’s about resetting the internal switch using minimal new data to fix that specific calibration issue.

Tom: The results show this gate significantly increases what they call mean refusal selectivity, boosting it from around thirty-three percent up to over fifty-four percent for their strongest baseline.

Jane: That jump in selectivity means the model starts refusing much more often than before the intervention, which is exactly what we want in a safety context.

Lu: They show that this method works well when you combine it with specific low-resource language setups, showing a mean delta increase to around forty-two percent on models like Qwen and Gemma.

Tom: The practical implication here is a diagnostic workflow: first, test if the harmfulness is represented, and then fix how it’s being used by recalibrating that decision threshold with a little supervision.

Jane: It moves the focus from learning new safety rules to simply tuning the existing rules for different contexts.

Meng: For us building systems, this means we can prioritize targeted calibration efforts on the languages where we have real deployment issues instead of spreading resources thin everywhere.

Lu: This approach suggests that safety alignment becomes much more scalable when you focus on fine-tuning the output mechanism rather than trying to perfectly map all semantic concepts across every single linguistic boundary.

Tom: So, instead of aiming for perfect multilingual alignment from day one, the authors suggest a diagnostic-then-fix path for specific language gaps.

Jane: It gives us a concrete way to handle those safety disparities without having to overhaul the entire model architecture or spend months on massive retraining cycles for every new language we want to support.

Lu: I'm curious if other models can use this same few-shot latent gate effectively when dealing with even more complex, nuanced harmfulness directions in future research.

Meng: From my side, I’m checking how easily this setup integrates into our existing deployment pipelines; if it can be done with just a few examples, that’s a big win for operational efficiency.

Tom: This paper really hammers home that calibration is often the weak link in multilingual systems when it comes to safety outcomes.

The paper's improvements: Tom: So we’ve seen how the paper diagnoses that safety failures are about decision calibration rather than missing harmfulness representations, and now we’re looking at what they propose to fix it.

Jane: They introduce this few-shot latent gate, which is basically a smart switch that reuses the high-resource safety signal and recalibrates the decision threshold using just a few examples in the target language.

Lu: This gate routes between steering toward refusal and just looking at the harmfulness direction ablation, which is pretty clever because it keeps both options open.

Meng: So instead of retraining a massive model for every new language, this method lets you fine-tune the output mechanism with minimal data to fix that specific calibration problem.

Tom: It’s about using that existing high-resource signal as a starting point so you don't have to start from scratch when you move to a low-resource setting.

Jane: The results show this gate substantially raises what they call mean refusal selectivity, jumping it from around thirty-three percent up to over fifty-four percent for their strongest baseline.

Lu: And they found that when you use zero-shot routing with this gate, the mean delta goes up to seventy-five point six on Gemma and Qwen models.

Tom: That jump in selectivity is what matters; it means the AI is actually refusing much more often than it was before this adjustment.

Jane: So what’s the real world shift here? It changes how we think about fixing safety gaps in multilingual systems by focusing on tuning the final decision step instead of trying to teach a whole new concept.

Meng: For deployment, this means you don't need huge datasets for every language just to define a new safety vector; you just need the right calibration examples.

Lu: They show this approach is selective, achieving the best delta in seventeen out of eighteen different language pairs tested.

Tom: That selectivity across different languages is pretty impressive, but they did flag a limitation: it still relies on having some initial representation of harm to start with, even if it's just a weak one.

Jane: And that’s the catch; the method doesn't create the harmfulness direction itself if it wasn't present in the high-resource data.

Meng: So we can use this to fix calibration issues where representation is already there, but it won't magically create a concept where none was learned.

Lu: It’s a diagnostic workflow: first you check if the representation is good, and then you recalibrate how it's used with minimal supervision.

Tom: The implication for the world is that we can start fixing these safety disparities more efficiently by focusing on targeted calibration efforts in specific low-resource languages.

Jane: It gives us a practical path forward when we’re facing real deployment hurdles across diverse linguistic environments.

Conclusion: Tom: So we’re wrapping up on "Low-Resource Safety Failures Are Action Failures, Not Representation Failures."

Jane: Basically, they show that when safety fails in languages with less data, it's not a total lack of understanding, but a problem with how the model decides to act on what it knows.

Lu: The main implication is that we can fix these gaps by tuning the decision threshold instead of trying to train a whole new concept for every language.

Meng: It means we don't need massive amounts of data just to define safety rules for every new language; we just need the right calibration examples.

Lalam: I think this is important because it means our models can become much more culturally aware and safe as we deploy them globally.

Tom: Right, so the diagnostic-then-fix workflow is pretty straightforward for engineers to follow.

Jane: Exactly, it’s about resetting the model's decision point using just a few targeted examples of the target language.

Lu: It opens up a whole new way to think about multilingual safety alignment that doesn't require massive retraining efforts across all languages simultaneously.

Meng: I just want to make sure we emphasize that this works because the underlying harmfulness direction is actually present in those models, even if it's weak.

Tom: That’s the core idea—the representation is there, but the action is flawed due to poor calibration.

Jane: It’s a big win for making AI safer and more reliable as we move into more diverse global applications.

Lu: We should definitely keep looking at how this few-shot latent gate can be adapted for other types of model steering in the future.

Meng: I’m checking if this recalibration approach fits well into our existing deployment pipelines, because that's where the real work happens.

Lalam: If we can make AI systems this robust to low-resource language challenges, it really helps build a more inclusive culture for everyone using these tools.

Tom: So, to wrap up on "Low-Resource Safety Failures Are Action Failures, Not Representation Failures," the message is that calibration fixes action problems without needing new representation learning.

Jane: It’s a solid framework for addressing those safety disparities we see across different linguistic settings.

Lu: We'll be watching how this mechanism performs when we push it into more complex, nuanced harmfulness directions next.

Meng: I'm looking forward to seeing if this targeted calibration can scale up efficiently in real-world scenarios.

Rashad Aziz Ikhlasul Akmal Hanif, Fajri Koto

Mohamed bin Zayed University of Artificial Intelligence

cs.CL, cs.AI

Submitted: 2026-05-31

Updated: 2026-10-05

Code: https://github.com/rashadaziz/low-resource-safety

Importance score: 91/100

The gist: The paper investigates why safety alignment learned in high-resource languages fails to transfer effectively to low-resource languages, specifically focusing on whether this failure stems from a lack

Key concepts

Harmfulness Representation
This refers to the internal mathematical direction within a language model's activation space that indicates whether a piece of text is harmful or harmless. The study found this representation exists in low-resource languages, meaning the model 'knows' what is bad, but it doesn't know how to use that knowledge correctly.
Calibration Failure
Calibration failure means the model sets its decision threshold incorrectly. In this context, even though the harmfulness signal is present, low-resource prompts result in lower projections onto this signal. This causes the model to set a much higher bar for triggering a refusal compared to high-resource settings.
Few-Shot Latent Gate
This is the proposed solution. It's a mechanism that reuses the strong harmfulness signal from high-resource data and recalibrates it using only a few examples in the target language. This gate adjusts the decision threshold, effectively teaching the model how to properly convert its existing knowledge into an actionable refusal.
Mean Refusal Selectivity (∆)
This metric measures how selective a model is at refusing harmful content. The study showed that after intervention, this selectivity significantly increased across different models and setups, proving the recalibration successfully improved the model's ability to refuse harmful prompts in low-resource settings.

Terminology

Summary

The paper investigates why safety alignment learned in high-resource languages fails to transfer effectively to low-resource languages, specifically focusing on whether this failure stems from a lack of a harmfulness representation or a failure to act upon an existing one. This research matters because it diagnoses the mechanism behind the observed multilingual safety gap and proposes an intervention—a few-shot latent gate recalibration—that repairs these failures by resetting decision thresholds without requiring new model training.

The gist: What fails to transfer is calibration of the safety decision, not the underlying representation.

How it works

The study diagnoses the failure by testing two hypotheses regarding low-resource failures: whether they reflect missing harmfulness representations or a failure to use an existing one. The diagnosis shows that harmful refusal drops from 87.9% to 43.9% across Qwen2.5-7B, Gemma-2-9B, and Llama-3.1-8B when prompts are translated into low-resource languages, while the harmfulness direction remains linearly separable in low-resource activations nearly as well as high-resource ones. This indicates that the relevant representation is present, but the model fails to convert the representation into refusal.

Diagnosis of Failure

The researchers test whether a direction learned from high-resource activations remains discriminative when applied to low-resource prompts. They find that harmfulness remains linearly separable in lower-resource activations, and a harmfulness direction extracted entirely from high-resource data still causally mediates refusal. The failure is localized to the calibration of the safety decision, as the low-resource prompts produce lower projections onto the harmfulness direction, so the same latent evidence less reliably triggers the model’s implicit refusal threshold.

Proposed Intervention: Few-Shot Latent Gate

To repair this calibration failure without retraining, the authors introduce a few-shot latent gate that reuses the high-resource harmfulness signal and recalibrates its decision threshold using a small number of target-language examples. This gate routes between refusal steering and harmfulness-direction ablation, substantially raising mean refusal selectivity (∆ = harmful − harmless refusal) from 33.6 for the strongest adapted baseline to 54.5.

Evaluation and Results

The intervention is evaluated using a few-shot threshold calibration, where the threshold is chosen to maximize F1 on b harmful and b harmless training examples. In the final steering method, Zero-shot HRL routing is strong, raising mean ∆ from 15.3 to 38.6 on Qwen and 60.1 to 75.6 on Gemma. The method proves selective, achieving the best ∆ in 17 of 18 pairs, with the HRL + 32 LRL setup reaching a mean ∆ from 38.6 to 42.4 for Qwen and Gemma.

Conclusion

The paper concludes that Low-resource safety failures are not failures to represent harmfulness, but rather a failure in calibration, which can be repaired by recalibrating the decision threshold with minimal target-language supervision. The practical path suggested is a diagnostic-then-fix workflow: test whether harmfulness is represented before deciding how to repair its use.

REFERENCES

Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. 2024. Refusal in language models is mediated by a single direction. In Advances in Neural Information Processing Systems 37 (NeurIPS 2024).

Muhammad Falensi Azmi, Muhammad Dehan Al Kautsar, Alfan Farizki Wicaksono, and Fajri Koto. 2025. IndoSafety: Culturally grounded safety for LLMs in Indonesian languages. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 9135–9166, Suzhou, China. Association for Computational Linguistics.

Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez. 2022. Constitutional AI: Harmlessness from AI feedback. Preprint,

Somnath Banerjee, Sayan Layek, Pratyush Chatterjee, Animesh Mukherjee, and Rima Hazra. 2025. Soteria: Language-specific functional parameter steering for multilingual safety alignment. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9347–9364, Suzhou, China. Association for Computational Linguistics.

Yuyan Bu, Xiaohao Liu, Zhaoxing Ren, Yaodong Yang, and Juntao Dai. 2026. Align once, benefit multilingually: Enforcing multilingual consistency for LLM safety alignment. In The Fourteenth International Conference on Learning Representations.

Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2024. Multilingual jailbreak challenges in large language models. In The Twelfth International Conference on Learning Representations.

Clément Dumas, Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2025. Separating tongue from thought: Activation patching reveals language-agnostic concept representations in transformers. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 31822–31841, Vienna, Austria. Association for Computational Linguistics.

Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread.

Leo Gao, Stella Biderman, Sid Black, Laurence Golding et al. 2020. The Pile: An 800GB dataset of diverse text for language modeling.

Yangsibo Huang, Samyak Gupta, Mengzhou Xia et al. 2024. Catastrophic jailbreak of open-source LLMs via exploiting generation. In The Twelfth International Conference on Learning Representations.

Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Durrani, and Husrev Taha Sencar. 2026. There is more to refusal in large language models than a single direction. Preprint,

Priyanshu Kumar et al. 2025. PolyGuard: A multilingual safety moderation tool for 17 languages. In Proceedings of the Second Conference on Language Modeling (COLM 2025).

Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy et al. 2025. Programming refusal with conditional activation steering. In The Thirteenth International Conference on Learning Representations.

Jiaming Liang, Zhaoxin Wang, and Handing Wang. 2026. Multilingual safety alignment via sparse weight editing. Preprint,

Yang Liu, Dan Iter, Yichong Xu et al. 2023. G-eval: NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.

Christina Lu, Jack Gallagher, Jonathan Michala et al. 2026. The assistant axis: Situating and stabilizing the default persona of language models. Preprint,

Mantas Mazeika, Andy Zou, Norman Mu et al. 2023. TDC 2023 (LLM Edition): The trojan detection challenge. In NeurIPS Competition Track.

Long Ouyang, Jeff Wu, Xu Jiang et al. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022).

Wenbo Pan, Zhichao Liu, Qiguang Chen et al. 2025. The hidden dimensions of LLM alignment: A multidimensional analysis of orthogonal safety directions.

Improvements for AI systems

  1. This system can be improved by implementing a few-shot latent gate that reuses the high-resource harmfulness signal and recalibrates its decision threshold using a small number of target-language examples. This intervention substantially rais[es] mean refusal selectivity (∆ = harmful − harmless refusal) from 33.6 for the strongest adapted baseline to 54.5 while preserv[ing] MMLU utility.

  2. The improved system will achieve better selective refusal by routing between refusal steering and harmfulness-direction ablation at the extraction layer, ensuring that Prompts classified as harmful receive refusal steering; prompts classified as harmless have the HRL harmfulness direction removed.

  3. The system can be enhanced to handle low-resource safety failures by using a mechanism where the representation is intact, so the right intervention is to recalibrate how it is used, not to relearn it, specifically by resetting the decision threshold on a handful of target-language examples.

Sources

Related papers