TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity

arXiv:2412.02371 · cs.CL, cs.CR · Submitted 2024-12-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity".

Tom: The gist The proposed method TSCheater generates high-quality Tibetan adversarial texts by considering the characteristic of Tibetan encoding and the feature that visually similar syllables have similar semantics,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Moving into segment two, let's talk about the title and the authors of this paper, "TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity."

Jane: That title immediately tells you what’s going on—it’s not just a general attack; it’s specifically focusing on using visual similarity as a key component in generating those adversarial texts.

Lu: The authors are from Minzu University of China, and they are clearly working with resources related to Tibetan language research, which gives them the necessary context for this kind of specialized work.

Meng: So, what does this mean practically for the community using these models? Is it just academic curiosity right now?

Tom: It’s more than that; it suggests that when you look at languages like Tibetan, you can develop methods that are tailored to their specific visual properties instead of just applying general NLP techniques.

Jane: They aren't just looking at the text as characters; they are looking at the underlying structure of the script itself, which is what makes this method distinct from other approaches.

Lu: The fact that they mention that this encoding characteristic also appears in other abugidas like Devanagari script really opens up possibilities for generalizing this technique later on.

Meng: Generalizing is good, but how do we know it will work the same way for a completely different script? That’s where the practical concerns come in.

Tom: Well, the paper suggests that by focusing on visual similarity features shared across these abugidas, you create a framework that could be transferable to other Indic scripts down the line.

Jane: It’s about finding commonalities in how these different writing systems structure their syllables and characters that can be exploited for adversarial generation.

Lu: This points toward a research direction where visual encoding analysis becomes an integral part of language model robustness testing, not just an afterthought.

Meng: I guess it means future work needs to look beyond English and focus on languages with complex scripts where visual cues are more pronounced.

The paper's summary: Tom: Now let's talk about what the paper actually summarizes, focusing on the core mechanism of TSCheater.

Jane: The summary boils down to a process where they take an original text and generate a perturbed version, x prime, by adding an imperceptible noise delta.

Lu: But instead of random noise, they use a structured approach based on the granularity of Tibetan script—breaking it down into letters, syllables, and words.

Meng: So the paper is suggesting that the way you attack an AI shouldn't be at the word level alone; it needs to respect how those components are visually related.

Tom: Precisely. They use this structure to build a database of syllable visual similarities called TSVSDB, which is based on cross-correlation matching between syllables.

Jane: This database then lets them identify specific substitution syllables and words that have high scores—those with a CCOEFF NORMED between zero point eight and one point one.

Lu: They figure out the attack effect by calculating Delta P, which is the difference in probability of the model predicting something before and after making these specific substitutions.

Meng: So, they are using this structured database to pinpoint the single best syllable or word change that yields the largest prediction shift.

Tom: And they sort these potential changes using a metric H that combines their substitution score with that Delta P star to decide which ones to use in sequence.

Jane: It’s a structured, multi-step process designed to find high-quality adversarial examples by prioritizing those changes that are both impactful and visually plausible.

Lu: This method addresses the limitation of previous methods that might just be brute-forcing substitutions without considering how the script is actually constructed.

Meng: So, it’s about being smarter about perturbation, not just louder with it, which makes sense for keeping things subtle.

The paper's improvements: Tom: Let's shift over to what the authors claim as improvements over prior work in this area.

Jane: They highlight that TSCheater-s has extremely high visual similarity with the original texts, which is impressive given that it still maintains good semantic similarity.

Lu: That’s significant because usually, when you push for visual subtlety, you lose some of the semantic impact of the change; they managed to keep both high.

Meng: And they claim that TSCheater-s achieves an extremely small perturbation magnitude because the syllable substitution is treated as a partial modification rather than replacing an entire syllable.

Tom: That means the noise isn't just a big swap; it’s a finer tuning of existing visual elements, which is much more subtle.

Jane: While TSCheater-w might not be as good in every single dimension, they state that it performs best in terms of attack effectiveness and human acceptance.

Lu: So the paper suggests there’s a trade-off: you can prioritize maximum impact or maximum subtlety depending on what your goal is.

Meng: It seems like the method offers flexibility, allowing researchers to choose their focus based on whether they need a powerful attack or something that looks almost identical to the original.

Tom: And finally, they point out that this approach yields the first Tibetan adversarial robustness evaluation benchmark called AdvTS, which was generated by existing methods and proofread by humans.

Jane: That benchmark is valuable because it gives people a standardized way to test how robust models are when faced with adversarial examples in low-resource languages.

Lu: It moves the discussion from just creating one specific attack to establishing a standard for measuring robustness across different languages.

Conclusion: Tom: We’re coming to the end of our discussion on this paper "TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity."

Jane: To summarize, this method tackles adversarial text generation in Tibetan by leveraging visual similarity features of the script.

Lu: The core idea is using a structured database based on syllable and word granularity to find substitutions that maximize prediction change while keeping them visually similar.

Meng: The main practical contribution seems to be establishing a benchmark like AdvTS for testing model robustness on these kinds of languages.

Tom: And the implication for broader research is that this technique has potential transferability to other abugidas, not just Tibetan but also Devanagari script.

Jane: So, we've seen how they balance attack power with imperceptibility and have built a framework that is both effective and visually grounded.

Lalam: This work shows how paying attention to the visual structure of a language can unlock new ways to challenge AI systems in ways that are both creative and grounded in linguistics.

Lu: It’s a strong indication that researchers should start incorporating this kind of visual analysis into their robustness studies for languages with less data.

Meng: We just need to keep pushing for practical applications where these kinds of robust models are needed, especially since low-resource languages play such an important role in AI development.

Tom: That’s the takeaway on TSCheater—a simple and sweet method that shows how focusing on encoding characteristics can lead to better adversarial examples for complex language models.

Xi Cao, Quzong Gesang, Yuan Sun, *Nuo Qun*, *Tashi Nyima

Minzu University of China · National Language Resource Monitoring & Research Center Minority Languages Branch, Beijing, China · Tibet University, Lhasa, China · Collaborative Innovation Center for Tibet Informatization by Ministry of Education & Tibet Autonomous Region, Lhasa, China

cs.CL, cs.CR

Submitted: 2024-12-03

Updated: 2026-10-04

Comments: Revised Version; Accepted at ICASSP 2025

Journal ref: ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

DOI: 10.1109/ICASSP49660.2025.10889732

Code: https://github.com/metaphors/TSAttack

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 77/100

The gist: The gist The proposed method TSCheater generates high-quality Tibetan adversarial texts by considering the characteristic of Tibetan encoding and the feature that visually similar syllables have

Key concepts

Tibetan Encoding Granularity
Tibetan script is broken down into units like letters, syllables, and words, which differ from English or Chinese. This structural understanding allows the method to target specific parts of the text for subtle manipulation.
TSVSDB
This is a database created by measuring visual similarity among 18,833 Tibetan syllables. It helps the system identify which syllables look similar to each other, which is crucial for finding effective substitutions that maintain visual realism.
Attack Effectiveness ($ΔP$)
This measures how successfully the adversarial text fools a language model. It is calculated by comparing the probability of predicting a word given the original text versus the perturbed text, aiming to maximize this difference.

Terminology

Summary

The gist The proposed method TSCheater generates high-quality Tibetan adversarial texts by considering the characteristic of Tibetan encoding and the feature that visually similar syllables have similar semantics, which outperforms existing methods in attack effectiveness, perturbation magnitude, semantic similarity, visual similarity, and human acceptance.

Methodology

  1. For a successful textual adversarial attack, let x' = x + δ be the perturbed input text where δ is the ϵ-bounded, imperceptible perturbation (Page 1).

  2. The granularity of Tibetan script can be divided into letters [or Unicodes], syllables, and words, which is different from English (letters [or Unicodes], words) and Chinese (characters [or Unicodes], words) (Page 2).

  3. The normalized cross-correlation matching algorithm is used to measure the visual similarity between 18,833 Tibetan syllables and construct a Tibetan syllable visual similarity database called TSVSDB (Page 2).

  4. The set of substitution syllables Cs consists of syllables with CCOEFF NORMED within (0.8, 1) (Page 3).

  5. The set of substitution words Cw is constructed by taking the Cartesian product operation on the sets corresponding to each word, filtering out elements whose CCOEFF NORMED product of syllables is within (0.8, 1) (Page 3).

  6. The attack effect is determined by calculating ∆P = P(yx) − P(yx') and selecting the substitution syllable s or word w that maximizes this difference, denoted as ∆P∗ (Page 3).

  7. The substitution order is determined using the metric H = Sof tmax(S) · ∆P∗, where S = P(yx) − P(yxˆ), to sort scores in descending order (Page 3).

Evaluation

The method is evaluated across five dimensions: attack effectiveness, perturbation magnitude, semantic similarity, visual similarity, and human acceptance (Page 1).

  • Attack effectiveness is measured by Accuracy Drop Value (ADV), where a larger ADV means a more effective attack (Page 4).

  • Perturbation magnitude is evaluated using Levenshtein Distance (LD) between the original text and the adversarial text, where a smaller LD indicates a more slight perturbation (Page 4).

  • Semantic similarity is assessed by Cosine Similarity (CS), calculated based on word embeddings of Tibetan-BERT, where a larger CS means greater semantic similarity (Page 4).

  • Visual similarity is evaluated using CCOEFF NORMED, where a larger value indicates more visual similarity (Page 4).

  • Human acceptance is rated on a scale of 1 to 5, and texts rated 4 or 5 are considered acceptable by humans (Page 4).

Experimental Setup

The experiments utilize eight victim language models constructed via the paradigm of “PLMs + Fine-tuning” using datasets such as TNCC-title and TU SA (Page 3).

The victim language models were built by fine-tuning PLMs like Tibetan-BERT and CINO on the respective datasets, with hyper-parameters taken from previous work (Page 3).

Baselines compared against TSCheater include TSAttacker, a syllable-level black-box method, and TSTricker, a syllable- and word-level method (Page 4).

Results

TSCheater outperforms existing methods in all five evaluated dimensions (Page 5).

Specifically, TSCheater-s has extremely high visual similarity with the original texts and performs very well in semantic similarity (Page 5).

TSCheater-s achieves an extremely small perturbation magnitude because the perturbation of syllables is no longer a whole substitution but a partial modification (Page 5).

While TSCheater-w does not perform as well as TSCheater-s in three dimensions, it performs best in attack effectiveness and best in human acceptance (Page 5).

The first Tibetan adversarial robustness evaluation benchmark called AdvTS was constructed by existing methods and proofread by humans (Page 5).

Conclusion

TSCheater is a simple and sweet Tibetan adversarial text generation method that considers the characteristic of Tibetan encoding and the feature that visually similar syllables have similar semantics, with transferability to other abugidas like Devanagari script (Page 1).

The adversarial robustness of language models covering low-resource cross-border languages plays a crucial role in AI and national security, so more attention is called for it (Page 5).

Acknowledgments

This work is supported by the National Social Science Foundation of China (22&ZD035), the Key Project of Xizang Natural Science Foundation (XZ202401ZR0040), the National Natural Science Foundation of China (61972436), and the MUC (Minzu University of China) Foundation (GRSCP202316, 2023QNYL22, 2024GJYY43) (Page 5).

References

[1] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. J. Goodfellow, and R. Fergus, “Intriguing properties of neural networks,” in 2nd International Conference on Learning Representations, ICLR 2014 (Page 5).

[2] I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in 3rd International Conference on Learning Representations, ICLR 2015 (Page 5).

[3] R. Jia and P. Liang, “Adversarial examples for evaluating reading comprehension systems,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2021–2031 (Page 5).

[4] S. Goyal, S. Doddapaneni, M. M. Khapra, and B. Ravindran, “A survey of adversarial defenses and robustness in NLP,” ACM Comput. Surv., vol. 55, no. 14s, jul 2023 (Page 5).

[5] S. Hua, S. Jin, and S. Jiang, “The limitations and ethical considerations of ChatGPT,” Data Intelligence, vol. 6, no. 1, pp. 201–239, 02 2024 (Page 5).

[6] Y. Chen, H. Gao, G. Cui, F. Qi, L. Huang, Z. Liu, and M. Sun, “Why should adversarial perturbations be imperceptible? Rethink the research paradigm in adversarial NLP,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 11 222–11 237 (Page 5).

[7] M. Iyyer, J. Wieting, K. Gimpel, and L. Zettlemoyer, “Adversarial example generation with syntactically controlled paraphrase networks,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 1875–1885 (Page 5).

[8] S. Eger, G. G. S¸ahin, A. Ruckl ¨ e, J.-U. Lee, C. Schulz, M. Mes- gar, K. Swarnkar, E. Simpson, and I. Gurevych, “Text processing like humans do: Visually attacking and shielding NLP systems,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 1634–1647 <ref:

Improvements for AI systems

  1. The TSCheater method can be transferred to other abugidas such as Devanagari script, enabling its application to other Indic scripts for adversarial text generation research without requiring a complete redevelopment of the core mechanism.

  2. The system can generate adversarial texts that are superior in multiple dimensions: TSCheater-s have extremely high visual similarity with the original texts and TSCheater performs well in attack effectiveness and best in human acceptance.

  3. The generated adversarial texts exhibit a extremely small perturbation magnitude because the method considers the characteristic that visually similar syllables have similar semantics, leading to a better trade-off between attack power and imperceptibility compared to existing methods.

  4. The research provides the first Tibetan adversarial robustness evaluation benchmark called AdvTS, which is generated by existing methods and proofread by humans, allowing for standardized and rigorous comparative studies of model robustness on low-resource languages.

Sources

Related papers