One Image, No Tokens: A Controlled Study of Glyph-Based Chinese Language Modeling
cs.CV, cs.AI
Submitted: 2026-07-04
Updated: 2026-09-26
Terminology
Sources
- LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
- C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
- Language Modelling with Pixels
- PIXAR: Auto-Regressive Language Modeling in Pixel Space
- CLIPPO: Image-and-Language Understanding from Pixels Only
- Glyce: Glyph-vectors for Chinese Character Representations
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models