Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models

arXiv:2610.02876 · cs.CV · Submitted 2026-10-02 · Read on arXiv

cs.CV

Submitted: 2026-10-02

Updated: 2026-10-02

Related papers