SketchVLM: Vision language models can annotate images to explain thoughts and guide users

arXiv:2604.22875 · cs.CV, cs.AI · Submitted 2026-04-23 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-04-23

Updated: 2026-08-31

Code: https://github.com/allenai/molmo

Project page: https://sketchvlm.github.io

Terminology

Sources

Related papers