IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamil
cs.CL, cs.AI
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/NLP2CT/IndicDetect
License: http://creativecommons.org/licenses/by/4.0/
The gist: The rapid proliferation of LLMs has further heightened the need to develop dependable AI-generated text detection, especially beyond English.
Terminology
Abstract
The rapid proliferation of LLMs has further heightened the need to develop dependable AI-generated text detection, especially beyond English. Nevertheless, current benchmarks pay little attention to Indic languages and test detectors in idealized settings that do not represent the real world. We present a generalized benchmark for AI-generated text detection in Hindi, Telugu, and Tamil, which we call IndicDetect, designed to assess the robustness of detectors under realistic distribution shifts. IndicDetect comprises highly curated human-written texts matched with LLM-generated counterparts across various domains and generators, and systematically evaluates detectors in the presence of domain shift, generator shift, and adversarial perturbation. Using a single and repeatable evaluation scheme, we evaluate a wide range of statistical and neural detectors. We find substantial robustness failures: supervised neural detectors perform well in-distribution, while training-free methods degrade considerably under unseen generators and adversarial attacks. The severity of these failures varies across languages, with Hindi exhibiting the largest overall degradation under adversarial perturbations. These results highlight that the primary weakness of existing detectors in Indic settings lies in their robustness, not in their peak accuracy. IndicDetect provides standard data splits, an evaluation protocol, and baselines to establish a robust, language-aware foundation for AI-generated text detection in Indic scripts.
Sources
- MEGA: Multilingual Evaluation of Generative AI
- Language Models are Few-Shot Learners
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
- MGTBench: Benchmarking Machine-Generated Text Detection
- AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages
- MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
- Qwen2.5 Technical Report
- M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
- Release Strategies and the Social Impacts of Language Models
- DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text
- DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
- HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus
- On the Generalization of Training-based ChatGPT Detection Methods
- TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation
- Qwen3 Technical Report
- Evaluating AIGC Detectors on Code Content
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering