From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

arXiv:2610.11818 · cs.CV, cs.AI · Submitted 2026-10-08 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-10-08

Updated: 2026-10-08

Code: https://github.com/uddipan77/Analysis-of-Lightweight-Vision-Language-Models-for-Document-OCR-and-Str

Terminology

Sources

Related papers