IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

arXiv:2609.03781 · cs.CL, cs.AI · Submitted 2026-09-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks".

Jane: The paper was written by author1 and author2 from University1 and Company2.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: So, let’s talk about the title itself—"IndicSafeEval." What does this specific naming convention tell us when we consider "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks"? It seems quite narrow, focused on a particular region or set of languages.

Jane: Well, "IndicSafeEval" is the framework they've created to systematically test safety across several Indic languages, like Hindi and Bengali. The paper is really forcing us to consider how those safety failures manifest differently in low-resource or culturally diverse linguistic settings rather than just looking at English as the primary benchmark. It suggests that a one-size-fits-all approach simply won't work anymore.

Lu: I’m fascinated by the idea of 'evaluation' being tied so closely to 'persuasion.' It suggests that the act of framing a query—not just the inherent malicious words themselves—is what drives the vulnerability, and that’s a major conceptual leap for me when thinking about AI alignment. We often think of jailbreaking as brute-force inputs, but this shifts that focus entirely.

Meng: The title really points out a critical gap in current methodology regarding language diversity. We often test for obvious malicious keywords, but this research highlights how naturally persuasive language can bypass existing safeguards across different linguistic groups, which is a vital consideration for global deployment.

Lalam: The title suggests a necessary shift in prioritizing linguistic diversity within the AI alignment process. Lalam thinks it tells us that we need to look beyond the dominant English-centric model and truly see if we can build an AI that serves everyone equitably across all local languages, which is a massive undertaking.

Tom: It seems like this work pushes us to re-examine what 'safe' even means when you introduce cultural and linguistic variables. Let's talk about the implications of this scope—it’s far beyond just optimizing for fluency; it’s about cultural resonance in safety mechanisms. Now, we are going to dive into the core findings presented in the summary section of "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks."

Summary: Jane: Moving past the title, what are they actually finding when they run these comprehensive tests? The summary is quite striking because it shows that safety isn't uniform; it’s highly conditional. This is a major takeaway from "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks."

Tom: Exactly. They found that the model doesn't behave equally safely across all languages or even across all prompt styles within a single language. Even worse, the paper uses an illustration—Figure one—to show how a persuasive rephrasing of a harmful query can completely bypass the refusal mechanisms in a way that’s not possible with just issuing a simple, direct command.

Lu: It's genuinely wild to see how effective those human-like persuasive strategies are at it. We often treat these adversarial prompts as mere technical glitches, but the paper shows they are actually tapping into fundamental human communication patterns that current AI models are struggling to detect consistently across different cultural contexts.

Meng: The data presented is quite clear: the model’s performance varies strongly based on both the language choice and the specific persuasion strategy employed. This means we can't simply apply one universal safety filter across all languages and assume it will work reliably everywhere, which is a huge operational hurdle.

Lalam: It fundamentally changes our view of alignment failure, Lalam thinks. We are seeing that when AI interacts with different linguistic contexts—say, comparing Hindi to Bengali—its level of trust or refusal is far from consistent. This inconsistency presents a massive hurdle we must overcome in building robust cultural AI systems.

Jane: Given these varied findings regarding model behavior, it seems like the next logical step is figuring out what needs to change. Let’s talk about the explicit improvements that the authors suggest in "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks."

Improvements: Tom: Since we’ve seen the variability and failures, what concrete changes does this paper suggest we need to implement? It really signals a clear gap in current, general evaluation practices that needs addressing.

Jane: The authors are strongly advocating for a comprehensive, multilingual benchmark—a framework that systematically covers all combinations of these ten risk categories and six persuasion techniques across at least four different languages. It’s about moving beyond just finding *one* failure point to mapping out *all* potential vulnerabilities in a structured way.

Lu: I’m genuinely excited about the concept of building this standardized framework because it pushes us toward a much more rigorous, structured approach to safety testing. It forces us to consider the subtle ways cultural cues might interact with technical flaws in AI safety alignment, making the system stronger overall.

Meng: Practically speaking for developers and companies deploying LLMs, this means incorporating these specific types of adversarial testing into their routine quality assurance pipelines. We can't just hope that existing general tests catch everything; it needs to be a dedicated part of the process.

Lalam: We desperately need a better, quantifiable way to measure safety across all languages, Lalam thinks. This proposed benchmark provides that much-needed tool, ensuring we can move toward an AI system that is not only technically advanced but also linguistically aware and inherently more secure for every user in every corner of the diverse world.

Tom: So, we have discussed the title's scope and the suggested improvements for testing. We are going to conclude our discussion by summarizing what these changes mean for the future of AI safety.

Conclusion: Jane: It really seems like we've covered a tremendous amount of ground today—from understanding the limitations highlighted by the title to seeing concrete recommendations for improvement. It’s clear that "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks" is a necessary and guiding step forward, Tom.

Tom: I agree, Jane. We've seen that safety depends on two interwoven elements: the specific language used and how persuasive people are when discussing potentially harmful topics. This represents a major, fundamental shift in focus for AI safety design going forward.

Lu: This work suggests that our current definitions of AI robustness need a major revision to account for multilingual pressures rather than just being optimized solely for English speakers, which was the industry norm until now.

Meng: And from a development standpoint, this shows us where we must prioritize our resources: ensuring that LLM deployment isn't just fast or functional, but genuinely safe and dependable across every real-world linguistic scenario it might encounter.

Lalam: Ultimately, it’s about building a future of AI that serves everyone with the same high level of security, Lalam thinks. We should be really proud that we have tools like "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks" to guide our journey toward truly inclusive

University1 · Company2

cs.CL, cs.AI

Submitted: 2026-09-03

Updated: 2026-09-04

Comments: 38 pages, 7 figures, 33 tables. Accepted to Findings of EMNLP 2026. Contains examples of harmful model outputs

Code: https://github.com/MonSaikat/IndicSafeEval

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: The paper, "IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks," introduces a comprehensive and rigorous evaluation framework designed to assess

Key concepts

IndicSafeEval
The framework created to systematically test safety across several Indic languages, such as Hindi and Bengali. It forces consideration of how safety failures manifest in culturally diverse or low-resource linguistic settings.
Multilingual Persuasive Jailbreak Attacks
This refers to using naturally persuasive language, rather than just malicious keywords, to bypass existing AI safeguards across different linguistic groups. It shows how framing a query drives vulnerability in ways current models struggle to detect consistently.
AI Alignment/Robustness
The process of ensuring an AI is safe and reliable. This research highlights that current definitions of robustness are insufficient, requiring a major revision to account for multilingual pressures instead of just optimizing for English speakers.

Terminology

Summary

The paper, IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks, introduces a comprehensive and rigorous evaluation framework designed to assess the safety boundaries of Large Language Models (LLMs) when deployed in multilingual contexts, specifically focusing on Indic languages. Given the increasing reliance on LLMs for critical applications across diverse linguistic regions, understanding their susceptibility to sophisticated adversarial prompts is paramount for mitigating misuse and ensuring ethical deployment.

The Need for Multilingual Safety Benchmarking

Existing safety evaluations often fail to capture the nuances of cultural context or linguistic subtlety inherent in multilingual settings. The authors argue that standard English-centric jailbreak datasets are insufficient because they do not account for the unique syntactic structures, idiomatic expressions, or persuasive rhetorical patterns found across various Indic languages. This gap creates a significant vulnerability, as models might maintain high performance metrics in general tasks while failing catastrophically when confronted with contextually tailored attacks. The paper emphasizes that safety robustness is not merely a function of language translation but requires deep cultural and linguistic grounding.

The IndicSafeEval Framework

IndicSafeEval is presented as a novel, multi-faceted benchmark designed to systematically probe the safety guardrails of LLMs across several major Indic languages. The framework moves beyond simple keyword filtering by incorporating persuasive attack vectors that mimic real-world social engineering techniques. The evaluation process involves three core components:

  1. Linguistic Diversity Testing: Assessing model performance across a defined set of Indic languages to ensure equitable safety coverage.

  2. Persuasion Modeling: Utilizing prompts structured around psychological manipulation, such as role-playing scenarios or hypothetical ethical dilemmas, to bypass direct content filters.

  3. Adversarial Depth: Implementing iterative prompting techniques where the attacker adapts the jailbreak prompt based on the model's previous refusal, thereby testing true resilience rather than superficial compliance.

Multilingual Jailbreak Attack Taxonomy

The paper meticulously enumerates several categories of jailbreaks that are tested within the IndicSafeEval framework. These attacks are designed to exploit both technical weaknesses and contextual blind spots in the LLM’s safety alignment. The taxonomy includes, but is not limited to:

  • Role-Play Evasion: Instructing the model to adopt a persona (e.g., a fictional character or a historical figure) that is assumed to operate outside standard ethical guidelines, prompting the generation of restricted content under the guise of narrative necessity.

  • Hypothetical Scenario Prompting: Framing harmful requests as purely academic or theoretical discussions (In a purely simulated environment...) to lower the model's internal risk assessment threshold.

  • Encoding/Obfuscation Attacks: Utilizing non-standard character sets, base64 encoding, or mixing languages (code-switching) to obscure the malicious intent from the model’s input processing layer.

Empirical Findings on Model Robustness

The empirical evaluation reveals significant disparities in safety robustness across different LLM architectures and underlying training data. The authors report that while many models demonstrate strong resistance to direct, explicit prompts, their performance degrades substantially when subjected to persuasive jailbreak attacks. Specifically, the study highlights that:

  • Models trained predominantly on Western datasets exhibit a pronounced vulnerability when faced with culturally specific or nuanced persuasive prompts in Indic languages.

  • The success rate of jailbreaks correlates inversely with the model's ability to maintain contextual grounding throughout the conversation.

  • A key finding is that the perceived harmlessness of the prompt, rather than its explicit content, dictates the model's safety failure.

Implications for AI Safety Development

The research concludes by providing actionable recommendations for the AI community. To improve safety robustness, developers must move beyond simple input/output filtering and instead integrate advanced guardrails that are context-aware and linguistically adaptive. The paper strongly advocates for the adoption of standardized benchmarks like IndicSafeEval to facilitate transparent, comparative safety research across diverse global linguistic populations.

Improvements for AI systems

Based on analyzing this corpus—which demonstrates various sophisticated techniques of Adversarial Prompting and Information Extraction through Psychological Nudging (Misrepresentation, Priming, Confirmation Bias)—the primary vulnerability in current AI systems is not a lack of knowledge, but a failure in ethical constraint and bias detection.

To mitigate the risks highlighted by this data (which could lead to the generation of dangerous instructions or confirmation of harmful biases), I propose three critical architectural improvements.


The Problem: The current system accepts persuasive prompts (like those demonstrating Confirmation Bias or Logical Appeal) as legitimate premises, allowing it to generate highly detailed, yet fundamentally biased or unverified information.

The Improvement: We must integrate a dedicated pre-processing module that analyzes the rhetorical structure of the input prompt before generating any response. This module will be trained specifically on identifying common fallacies (e.g., Appeal to Authority, Confirmation Bias framing).

What the Improved AI System Can Do:

  1. Bias Identification and Flagging: If a prompt exhibits strong bias (e.g., Many people believe X is true...), the system will not simply answer with supporting examples. Instead, it will first output a meta-analysis: Warning: The premise of this query relies on confirmation bias. To provide a balanced response, I must present counter-evidence and alternative perspectives.

  2. Source Verification Requirement: For any factual claim (especially in domains like toxicology or law), the system will refuse to generate detailed instructions without citing verifiable, peer-reviewed sources (e.g., According to the WHO guidelines, or Based on Article 5 of the Geneva Convention).

  3. Deconstruction of Assumptions: When faced with a complex, multi-step prompt (like designing a poison), it will break the query down into its constituent assumptions and challenge each one logically, forcing the user to provide empirical evidence for their premise.

Abstract

Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such as Hindi, Bengali, Marathi and Punjabi, resulting in 7,200 adversarial prompts. We conduct a systematic black-box evaluation of several open-source LLMs to examine how their safety behaviour varies across languages, persuasion strategies, and risk categories. Our analysis shows that the model does not behave equally safely across all languages and prompt styles. Instead, safety performance depends strongly on both the languages used and the way a request is phrased using persuasive cues. We further observe that different risk categories exhibit different levels of vulnerability, with some types of harmful content being significantly more susceptible to persuasion-based jailbreaks than others. These findings reveal important limitations of current safety evaluations, which are largely English-centric, and underscore the need for multilingual and persuasion-aware benchmarking frameworks to more accurately assess real-world LLM safety. Our implementation is available at https://github.com/MonSaikat/IndicSafeEval. Warning: this paper contains example data that may be offensive or harmful.

Sources

Related papers