Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning
cs.CV, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/Qwen-Applications/TD-LTTS
Terminology
Sources
- Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Qwen2.5-VL Technical Report
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
- Training Large Language Models to Reason in a Continuous Latent Space
- Spotlight on Token Perception for Multimodal Reinforcement Learning
- Vision-aligned Latent Reasoning for Multi-modal Large Language Model
- Efficient Test-Time Scaling for Small Vision-Language Models
- Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
- Xiaomi MiMo-VL-Miloco Technical Report
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- Visual-RFT: Visual Reinforcement Fine-Tuning
- Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Seeing with You: Perception-Reasoning Coevolution for Multimodal Reasoning
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Generating Sequences by Learning to Self-Correct
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models