All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
cs.CV
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/YesianRohn/ScriptMoE
Terminology
Sources
- Qwen2.5-VL Technical Report
- Benchmarking Chinese Text Recognition: Datasets, Baselines, and an Empirical Study
- PaddleOCR 3.0 Technical Report
- Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
- GLM-OCR Technical Report
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- KOSMOS-2.5: A Multimodal Literate Model
- HunyuanOCR Technical Report
- Qwen3.5-Omni Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
- DeepSeek-OCR: Contexts Optical Compression
- DeepSeek-OCR 2: Visual Causal Flow
- PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models