SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding
cs.CL, cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4o System Card
- LLaVA-OneVision: Easy Visual Task Transfer
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- DocVQA: A Dataset for VQA on Document Images
- Towards VQA Models That Can Read
- C$^{3}$Bench: A Comprehensive Classical Chinese Understanding Benchmark for Large Language Models
- Seed1.5-VL Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- AnchiBERT: A Pre-Trained Model for Ancient ChineseLanguage Understanding and Generation
- OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- WenyanGPT: A Large Language Model for Classical Chinese Tasks
- Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
- F\`ux\`i: A Benchmark for Evaluating Language Models on Ancient Chinese Text Understanding and Generation
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering