CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
cs.CV, cs.AI, cs.CL
Submitted: 2025-06-21
Updated: 2026-05-25
Journal ref: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 5134-5143
Code: https://github.com/Teacher-Tom/CLiViS
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory
- DeepSeek-V3 Technical Report
- A Survey on Vision-Language-Action Models for Embodied AI
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Video-R1: Reinforcing Video Reasoning in MLLMs
- LLaMA: Open and Efficient Foundation Language Models
- OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Qwen2.5 Technical Report
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- MemGPT: Towards LLMs as Operating Systems
- Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models