JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
Ran Li, Huiguo He, Jiahuan Cao, Junle Liu, Hiuyi Cheng, Lianwen Jin
South China University of Technology
cs.CV, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 19 pages, 13 figures. Accepted to the Dataset Track of ACM Multimedia 2026 for oral presentation
Code: https://github.com/Ran00w/JieZi
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 95/100
The gist: The paper introduces Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process of ancient Chinese characters.
Terminology
Summary
The paper introduces Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process of ancient Chinese characters. ACCE is organized into four progressive levels: basic character identification (L1), glyph-form analysis (L2), meaning exegesis (L3), and diachronic evolution analysis (L4). These levels are grounded in established principles of traditional philology.
To support this task, the authors construct two complementary resources:
-
JieZi-Dataset: The first large-scale, expert-audited VQA training dataset for ACCE, comprising over 500K QA pairs and 130K glyph images spanning six script stages (Oracle Bone, Bronze, Warring States, Seal, Clerical, and Regular). It is constructed via a pipeline that reduces factual errors by constraining generation with expert-designed templates and source-text references from the authoritative etymological dictionary Hanzi Yuanliu Dazidian. Human verification is applied at each key stage to ensure scholarly accuracy. The dataset covers all ten ACCE subtasks, and each character is accompanied by detailed, multi-dimensional textual descriptions (median per-entry metadata length is 781 tokens, mean is 1,002 tokens).
-
JieZi-Bench: An evaluation benchmark aligned with the exegesis process, constructed and verified by human experts to ensure evaluation reliability. It contains 1,024 glyph images and approximately 8K QA pairs, with reference answers curated from four authoritative lexicographic works (Kangxi Dictionary, Shuowen Jiezi, Shuowen Jiezi Zhu, and Revised Mandarin Chinese Dictionary) held separate from the training data to prevent leakage. The benchmark includes splits with unseen characters and unseen glyphs to test generalization.
Experiments on multimodal large language models show that current models perform well on basic identification but struggle with glyph analysis, semantic reasoning, and diachronic understanding. For example, general MLLMs score moderately on categorical tasks (SCRC 40–77%) but struggle with fine-grained analysis (CHAR 10–37%), and performance deteriorates further on L4 where even the best non-fine-tuned model scores below 40% on EVOI. Models trained on more Chinese data perform notably better (e.g., Doubao-Seed-2.0-pro and Kimi-K2.5 outperform GPT-5.4), suggesting that domain-relevant data coverage is a primary bottleneck.
Fine-tuning on JieZi-Dataset substantially improves performance across all four levels. Even a lightweight 2B model outperforms Gemini-3.1-Pro and GPT-5.4 on structural parsing tasks, and the 9B model achieves state-of-the-art results across the board. The improvement increases with model capacity, with larger gains on deeper reasoning tasks. Generalization analysis shows that exegesis remains robust despite recognition failures on unseen characters and glyphs, confirming that training on JieZi-Dataset induces transferable paleographic knowledge rather than overfitting to seen character identities. Older scripts (e.g., Bronze) remain the most challenging.
The paper also includes ablation studies showing that each data construction stage contributes distinct gains: structured metadata improves classification-oriented tasks, VQA reformatting benefits generative subtasks, and open-source augmentation strengthens visual robustness. Data scaling analysis shows monotonic improvement from 25% to 100% of the dataset, confirming the value of the full 500K scale. A paraphrase robustness test indicates that performance gains are not mainly driven by surface-level template pattern matching but by learned glyph-grounded exegetical content.
The authors conclude that their work contributes the first resource covering the complete exegesis workflow across multiple script types, establishing a standardized foundation and a reproducible baseline for computational paleography. Code and dataset are available at https://github.com/Ran00w/JieZi.
Improvements for AI systems
Based on this paper, here are specific improvements I can make to AI systems, and what the improved systems can do:
1. Hierarchical Visual-Linguistic Reasoning for Script Evolution
-
Improvement: Add a multi-stage reasoning module that processes glyph images through progressive abstraction levels (shape → component → meaning → historical context), mirroring the paper’s L1–L4 structure.
-
What it can do: The AI can now answer questions like “Why does this Bronze script character differ from its Oracle Bone form?” by decomposing visual features into component-level analysis and then linking to semantic shifts, rather than treating the image as a flat classification problem.
2. Template-Constrained Generation with Source-Grounded Factuality
-
Improvement: Implement a two-pass generation pipeline: first, retrieve relevant etymological references (e.g., from Hanzi Yuanliu Dazidian), then constrain the decoder with expert-designed templates that force the model to cite or paraphrase those sources, reducing hallucination.
-
What it can do: The AI can produce scholarly-grade explanations for character meanings with verifiable citations, and it will refuse to invent etymologies when no source exists—critical for digital humanities tools and educational platforms.
3. Cross-Script Transfer Learning via Shared Glyph Components
-
Improvement: Pre-train a vision encoder on the six script stages (Oracle Bone → Regular) with a contrastive loss that aligns identical characters across scripts, then fine-tune on ACCE tasks.
-
What it can do: The AI can recognize a character in an unfamiliar script (e.g., Warring States) even if it only saw Seal script during training, enabling robust OCR and annotation for ancient manuscripts where script variants are common.
4. Diachronic Evolution Prediction Module
-
Improvement: Add a temporal reasoning head that takes a character’s glyph at one stage and predicts its likely form at a later stage, trained on the L4 EVOI subtask.
-
What it can do: The AI can generate plausible intermediate forms for extinct or undocumented script transitions, aiding paleographers in reconstructing missing links in character evolution chains.
5. Uncertainty-Aware Answering for Low-Confidence Glyph Analysis
-
Improvement: Integrate a confidence estimator that flags answers when the model’s visual features are ambiguous (e.g., damaged glyphs or rare variants), based on the paper’s finding that older scripts (Bronze) remain challenging.
-
What it can do: The AI will explicitly say “I am 40% confident this is the component ‘言’ due to corrosion” instead of giving a false definitive answer, making it safer for archival digitization where errors propagate.
6. Multi-Source Answer Aggregation from Lexicographic Works
-
Improvement: Train a reranker that combines reference answers from Kangxi Dictionary, Shuowen Jiezi, and others, weighting by task type (e.g., Shuowen for etymology, Kangxi for usage).
-
What it can do: The AI can produce a consensus answer with a confidence score that reflects agreement across authoritative sources, and it can explain discrepancies (e.g., “Shuowen says X, but Kangxi adds Y due to later usage”).
7. Data-Efficient Fine-Tuning for Low-Resource Scripts
-
Improvement: Use the paper’s finding that open-source augmentation strengthens visual robustness—apply synthetic glyph degradation (noise, partial occlusion) during fine-tuning to mimic real archaeological conditions.
-
What it can do: The AI becomes resilient to blurry, fragmented, or low-resolution images from field photographs, improving real-world performance on excavated artifacts without needing more labeled data.
8. Generalization-Aware Training with Unseen Character Splits
-
Improvement: Adopt the benchmark’s unseen-character and unseen-glyph splits as a training objective—force the model to solve tasks using only component-level reasoning, not memorized character identities.
-
What it can do: The AI can handle novel characters (e.g., newly discovered variants) by decomposing them into known radicals and applying learned exegetical rules, rather than failing on out-of-distribution inputs.
9. Paraphrase-Robust Answer Generation
-
Improvement: Add a paraphrase adversarial training step where the model must produce the same exegetical content regardless of question phrasing (e.g., “What does this mean?” vs. “Explain the semantic shift here”).
-
What it can do: The AI will not rely on surface-level template matching; it will extract the underlying glyph-grounded meaning, making it robust to diverse user queries in educational or research interfaces.
10. Lightweight Deployment for On-Device Paleography
-
Improvement: Distill the 9B model into a 2B student using the paper’s finding that even 2B models beat larger generalists on structural parsing—optimize for mobile or edge use in museum guides or field archaeology.
-
What it can do: A portable AI assistant can identify and explain ancient characters in real time from a camera feed, with offline capability, bringing expert-level exegesis to non-specialists in remote locations.
Abstract
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis. To address this limitation, we introduce Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process. ACCE is organized into four progressive levels: basic character identification, glyph-form analysis, meaning exegesis, and diachronic evolution analysis. To support this task, we construct two complementary resources. JieZi-Dataset is the first large-scale, expert-audited VQA training dataset for ACCE, comprising over 500K QA pairs. It is constructed via a pipeline that reduces factual errors by constraining generation with expert-designed templates and source-text references. Human verification is further applied at each key stage to ensure scholarly accuracy. JieZi-Bench is an evaluation benchmark aligned with the exegesis process, constructed and verified by human experts to ensure evaluation reliability. It consists of four levels with reference answers curated from authoritative lexicographic works held separate from the training data. Experiments on multimodal large language models show that current models perform well on basic identification but struggle with glyph analysis, semantic reasoning, and diachronic understanding. Fine-tuning on JieZi-Dataset substantially improves performance across all four levels. Code and dataset are available at https://github.com/Ran00w/JieZi.
Sources
- Qwen3-VL Technical Report
- C$^{3}$Bench: A Comprehensive Classical Chinese Understanding Benchmark for Large Language Models
- OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- An open dataset for the evolution of oracle bone characters: EVOBC
- OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
- Oracle Bone Inscriptions Multi-modal Dataset
- A comprehensive survey of oracle character recognition: challenges, benchmarks, and beyond
- Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
- Kimi K2.5: Visual Agentic Intelligence
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Oracle-MNIST: a Dataset of Oracle Characters for Benchmarking Machine Learning Algorithms
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models