GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
cs.CV, cs.AI
Submitted: 2026-07-04
Updated: 2026-09-17
Comments: Accepted version. Published in IEEE Transactions on Geoscience and Remote Sensing, DOI: 10.1109/TGRS.2026.3734378. 22 pages
Journal ref: IEEE Transactions on Geoscience and Remote Sensing, 2026
DOI: 10.1109/TGRS.2026.3734378
Code: https://github.com/Avalon-S/GeoSelect
Project page: https://avalon-s.github.io/GeoSelect
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
- Referring Remote Sensing Image Segmentation via Bidirectional Alignment Guided Joint Prediction
- SAM 3: Segment Anything with Concepts
- InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
- Enabling Training-Free Text-Based Remote Sensing Segmentation
- AUG: A New Dataset and An Efficient Model for Aerial Image Urban Scene Graph Generation
- RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
- ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
- Qwen3 Technical Report
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Falcon Mamba: The First Competitive Attention-free 7B Language Model
- The Llama 3 Herd of Models
- Generalized Referring Expression Segmentation on Aerial Photos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models