New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs

summary

Video file (mp4)

The gist

This paper investigates the detection of implicit toxicity expressed through Chinese neologisms—newly lexical items in sense or form—which can serve as new vehicles for toxic expression.

In short

The episode discusses a paper on detecting toxic Chinese neologisms using consensus-based Search-Augmented LLMs. The research addresses how new or evolving slang bypasses moderation systems, showing that current commercial APIs miss over eighty percent of such content. The proposed SeTox framework uses search to provide real-time context, significantly improving detection rates from sixty-two percent to nearly ninety-three percent.

Key concepts

Neologisms
New words or new meanings for old words. The paper focuses on how these terms are used in a toxic way, often sneaking past moderation systems because they look benign on the surface.
Consensus-based Detection
A method to determine toxicity by checking if a term's harmful meaning has actually spread and stabilized in public usage. This prevents flagging benign new words while catching terms that have become widely used as insults.
SeTox Framework
A detection framework that uses search tools to give LLMs real-time context when they encounter unknown or ambiguous terms. This 'search-augmented' approach allows the model to look up current usage, improving accuracy for evolving language.
Search-Augmented LLMs
Large Language Models enhanced with a search tool. This feature allows the model to access fresh, real-time information from the web when analyzing language, enabling it to handle rapidly evolving slang and neologisms effectively.

Terminology used across episodes

This episode discusses

The paper

New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs · Read on arXiv

Shiyao Cui, Qinglin Zhang, Di Wang, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Min He, Han Qiu, Minlie Huang

Tsinghua University · Institute of Computing Technology, Chinese Academy of Sciences · National Computer Network Emergency Response Technical Team Coordination Center of China · Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · JCSS, Tsinghua University (Institute for Network Sciences and Cyberspace) - Science City (Guangzhou) Digital Technology Group Co., Ltd.

DOI: 10.18653/v1/2026.acl-long.1602

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs".

Jane: The paper was written by Shiyao Cui, Qinglin Zhang, Di Wang, Yida Lu, Zhexin Zhang et al. from Tsinghua University and Institute of Computing Technology, Chinese Academy of Sciences and National Computer Network Emergency Response Technical Team Coordination Center of China and Institute of Information Engineering, Chinese Academy of Sciences and School of Cyber Security, University of Chinese Academy of Sciences and JCSS, Tsinghua University (Institute for Network Sciences and Cyberspace) - Science City (Guangzhou) Digital Technology Group Co., Ltd..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone. We've got a fascinating paper on the table today, and it's called "New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs." Jane, I gotta say, the title alone got me hooked.

Jane: Oh, absolutely, Tom. And for our listeners who just tuned in, let me break down what that title actually means. "Neologisms" are just new words or new meanings for old words. And this paper is looking at how these new words can be used to sneak toxic language past moderation systems.

Tom: Right, and it's not just about new words, it's about the "consensus-based" part. That's what really caught my eye. They're not just flagging any weird new term; they're checking whether the toxic meaning has actually caught on with the public.

Jane: Exactly. Think about it like this—if I make up a word right now and use it to insult someone, that's not really a neologism with consensus. But if a term spreads across social media and everyone starts using it in a derogatory way, then it's become a real problem that moderators need to handle.

Tom: And the paper's from a serious group—CoAI at Tsinghua, plus some folks from the National Computer Network Emergency Response Technical Team. That's a big deal, Jane. These are the people who actually deal with content moderation at scale in China.

Jane: For sure. And the fact that they're looking at Chinese neologisms specifically makes sense because Chinese internet culture is incredibly dynamic. New slang and coded terms pop up constantly, often specifically to evade moderation filters.

Tom: So the core problem they're tackling is that these new terms look totally benign on the surface. Like, the paper gives the example of "安卓" which means "Android" as in the operating system, but it's been twisted to mean "low-quality" or "low-grade" as a derogatory label for people.

Jane: And that's the sneaky part. A moderation system sees "Android" and thinks, "oh, that's harmless," but in context, it's being used as an insult. The paper says detection rates for these kinds of terms were below twenty percent for major commercial moderation APIs.

Tom: Below twenty percent! That's terrible. So basically, if you're using these tools, you're missing over eighty percent of the toxic content that's hiding behind these new words.

Jane: Which is exactly why this research is so important. They're not just pointing out the problem; they're building a solution. And that's what I'm really excited to dig into.

Tom: Me too. Let's get into the meat of this paper and see how they're actually tackling this challenge.

Summary: Tom: So, Jane, we've got the title unpacked, and now I want to get into what this paper actually did. The summary here is pretty dense, but the core idea is really elegant.

Jane: It is. The paper builds three things. First, they created a taxonomy to understand where these toxic neologisms come from and what makes them special. Second, they built a big lexicon—nine hundred seventy-four terms, each with detailed metadata. And third, they designed a detection framework called SeTox that uses search to keep up with evolving language.

Tom: Let's talk about that taxonomy first, because I think it's the foundation. They identified four origins: online communities, media and entertainment, public events, and cultural or folk traditions. Makes sense, right? That's where new language gets born.

Jane: Right, and then they looked at the characteristics that make these toxic neologisms so tricky. They're literally benign on the surface, they evolve dynamically, and their toxicity depends on consensus. That last point is crucial—a term shouldn't be flagged as toxic unless its harmful meaning has actually stabilized in public usage.

Tom: And that's where the consensus criteria comes in. They used search engines to check whether a term's toxic meaning has actually caught on. If you search for a term and the top results show it being used in a derogatory way, then it's reached consensus.

Jane: Exactly. And then they categorized the toxicity into five risk categories: derogatory attacks, sensitive topics, immoral behaviors, adult content, and illegal activities. That gives them a structured way to think about what makes these terms harmful.

Tom: And the lexicon itself—nine hundred seventy-four terms with metadata including conventional meaning, current meaning, origin, and an explanation of how the meaning evolved. That's a serious resource.

Jane: It is. And the metadata isn't just for show. They use it to generate training data for their detection model. The example in the paper for "安卓" is perfect—it shows how a term for a mobile operating system became a tool for social stigmatization.

Tom: So they've built this comprehensive resource, but the real question is how they use it to actually detect toxicity in the wild. That's where SeTox comes in, and I think that's the most exciting part of the paper.

Jane: Absolutely. Let's get into the methodology and how they're making this work in practice.

Improvements: Tom: Okay, so now we're getting to the heart of it—the SeTox framework. Jane, this is where the paper really shines. The problem with static models is that language evolves faster than they can be retrained.

Jane: Right, and that's the key insight. Instead of trying to train a model on every possible new term, SeTox gives the model a tool—literally a search tool. When the model encounters a term it doesn't recognize or that seems ambiguous, it can search the web for real-time context.

Tom: And that's the "search-augmented" part of the title. The model doesn't have to know everything upfront. It can look things up on the fly. The paper calls this "test-time scalability"—you don't need to retrain, you just let the model access fresh information.

Jane: And the results are pretty impressive. They trained SeTox on a seven-billion-parameter model, Qwen2 point 5-7B, and it outperformed much larger and more recent models like Qwen3-Max and Gemini-three-flash-preview. That's a huge deal.

Tom: It really is. The detection rate on neologism-containing toxic sentences went from around sixty-two percent with the base model to almost ninety-three percent with SeTox. And for terms the model had never seen before, the improvement was even more dramatic—from about fifty-two percent to ninety-four percent.

Jane: And it's not just about accuracy. The paper also shows that SeTox produces explanations for its decisions. When it flags something as toxic, it can explain why, referencing the neologism's actual meaning as found through search.

Tom: That's important for trust. A moderation system that just says "this is toxic" without explanation is hard to debug and hard to trust. But with SeTox, you can see the reasoning.

Jane: And the ablation study shows how crucial the search component is. When they removed the search tool, the detection rate dropped significantly, especially for unknown terms. The model would sometimes guess the meaning of a neologism and get it wrong, leading to incorrect safety judgments.

Tom: So the search isn't just a nice extra—it's essential. And I love that they showed the framework works across different model families and sizes. It's not tied to one specific model.

Jane: Right, they tested it on models from different families—InternLM, GLM, and the Qwen series—and it consistently improved performance. That suggests SeTox is a general approach, not just a one-off hack.

Tom: This is exactly the kind of practical improvement that could make a real difference in online safety. Let's look at the actual paper content and see what the first page tells us about their motivation.

First Page: Jane: So, Tom, we've talked about the framework and the results, but I want to go back to the first page of the paper because it sets up the problem so well. They open with those examples—"田园女" meaning "country girl" but used as a stigmatizing label for feminism.

Tom: Yeah, that example really sticks with you. On the surface, it sounds harmless, even a bit romantic. But in online discourse, it's become a weapon. And the paper points out that these terms can be either new meanings for existing words or new words entirely.

Jane: And that's the key distinction. "安卓" is an old word with a new toxic meaning. "田园女" is a new coinage that sounds benign but carries toxic baggage. Both are neologisms, but they evolve differently.

Tom: The paper also mentions their pilot study—fifty sentences with these implicit toxic neologisms, and the detection rates from Google's Perspective API and Baidu's Moderation API were below twenty percent. That's the motivation right there.

Jane: It really is. When the best commercial tools are missing eighty percent of the toxic content, you know there's a serious gap. And that gap exists because these terms are designed to fly under the radar.

Tom: And they're designed that way intentionally. People who want to spread hate or harassment learn what gets flagged and find ways around it. It's an arms race between moderators and bad actors.

Jane: Which is why the consensus-based approach is so smart. By checking whether a term's toxic meaning has actually caught on publicly, you avoid both over-flagging benign usage and under-flagging emerging toxicity.

Tom: And the first page really drives home the point that this is a systemic problem. It's not just a few bad words slipping through—it's a whole category of toxic expression that current systems can't handle.

Jane: Right. And that's what makes this paper so important. It's not just identifying the problem; it's offering a concrete, scalable solution. And I think that's a great note to carry into our final thoughts.

Tom: Couldn't agree more. Let's wrap this up.

Conclusion: Tom: Well, Jane, we've covered a lot of ground on "New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs." Let's pull it all together for our listeners.

Jane: Absolutely. The paper tackles a real problem—toxic language hiding behind new words and new meanings. They built a taxonomy to understand these neologisms, curated a lexicon of nine hundred seventy-four terms with rich metadata, and developed SeTox, a search-augmented framework that lets LLMs access real-time web context.

Tom: And the results speak for themselves. A seven-billion-parameter model with SeTox outperforms much larger, more recent models. The detection rate on toxic neologisms jumped from around sixty-two percent to nearly ninety-three percent. That's a massive improvement.

Jane: And it's not just about the numbers. SeTox provides explanations for its decisions, which builds trust and makes the system easier to debug. It's also model-agnostic, working across different model families and scales.

Tom: The implications are huge for online moderation. Platforms that deploy SeTox could catch toxic content that currently slips through, making online spaces safer without needing to constantly retrain models.

Jane: And the future work is exciting too. The paper mentions extending this to other languages and modalities, like memes and videos. The framework is language-agnostic in principle, so it could be adapted beyond Chinese.

Tom: I also appreciate that they're thoughtful about the ethical considerations. This kind of technology could be misused for censorship, so they emphasize controlled deployment in moderation settings.

Jane: Right. It's a powerful tool, and with power comes responsibility. But for the purpose of making online spaces safer, this is a significant step forward.

Tom: Well said, Jane. That's a wrap on "New Terms, New Toxicity." Great paper, great insights, and definitely one to watch for future developments.

Jane: Thanks for joining us, everyone. We'll be back next time with more cutting-edge research from arXiv. Until then, stay curious!

More episodes

← Home