Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

arXiv:2609.16255 · cs.LG, cs.AI, cs.CV · Submitted 2026-09-14 · Read on arXiv

cs.LG, cs.AI, cs.CV

Submitted: 2026-09-14

Updated: 2026-09-14

Comments: 14 pages, 2 figures, 5 tables. Published in MultiMedia Modeling (MMM 2026), LNCS 16412

Journal ref: MultiMedia Modeling (MMM 2026), Lecture Notes in Computer Science, vol. 16412, pp. 567-580, Springer, Singapore, 2026

DOI: 10.1007/978-981-95-6950-2_40

Code: https://github.com/OpenGVLab/InternVL

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers