VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 17 pages, 5 figures
Code: https://github.com/trieuntu/VietAIDetector
Project page: https://py-pdf.github.io/fpdf2/7https://doi.org/10.57967/hf/6234
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: In recent years, distinguishing between AI-generated text and human-written text has remained a challenge.
Terminology
Abstract
In recent years, distinguishing between AI-generated text and human-written text has remained a challenge. In this paper, we introduce VietAIDetector, an open-source tool designed specifically for detecting Vietnamese AI-generated text. It allows users to interact through a Gradio web interface with inputs ranging from raw Vietnamese text to common text file formats, including scanned documents and exceptionally long texts that exceed the context size of the employed Large Language Models (LLMs). The core component of the tool employs a Zero-Shot approach to detect AI-generated text without requiring domain-specific training data, building upon the previous VietBinoculars and Binoculars research. The tool is built upon a Vietnamese-specific language model and has been evaluated on out-of-domain datasets, demonstrating superior performance compared to existing methods primarily developed for English. Additionally, users can select optimal detection thresholds based on F1 score, accuracy, or TPR@0.05FPR requirements. The results are presented through the web interface, allowing users to easily review and verify suspicious texts or download them as a PDF report. The tool is publicly available at https://github.com/trieuntu/VietAIDetector
Sources
- VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
- PhoGPT: Generative Pre-training for Vietnamese
- Release Strategies and the Social Impacts of Language Models
- GPTZero: Robust Detection of LLM-Generated Texts
- PhoBERT: Pre-trained language models for Vietnamese
- Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering