PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/tusu-code/20260121-icml2026-2
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Distilling the Knowledge in a Neural Network
- The 2017 DAVIS Challenge on Video Object Segmentation
- Improved Neural Radiance Fields Using Pseudo-depth and Fusion
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- FitNets: Hints for Thin Deep Nets
- Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
- PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
- VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
- InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
- GLaMM: Pixel Grounding Large Multimodal Model
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
- F-LMM: Grounding Frozen Large Multimodal Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection