UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception
cs.CV, cs.AI
Submitted: 2026-09-29
Updated: 2026-09-29
Project page: https://4dvlab.github.io/UniAfford
Terminology
Sources
- Qwen2.5-VL Technical Report
- SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
- Segment Anything
- Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
- UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model
- Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP
- Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- Decoupled Weight Decay Regularization
- Learning Affordance Grounding from Exocentric Images
- AffordanceLLM: Grounding Affordance from Vision Language Models
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- Side Adapter Network for Open-Vocabulary Semantic Segmentation
- GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
- SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
- LLMBind: A Unified Modality-Task Integration Framework
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models