RefAM: Attention Magnets for Zero-Shot Referral Segmentation
cs.CV
Submitted: 2025-09-26
Updated: 2026-09-20
Code: https://github.com/black-forest-labs/flux
Project page: https://refam-diffusion.github.io
Terminology
Sources
- Quantifying Attention Flow in Transformers
- Why do LLMs attend to the first token?
- Layer-wise Relevance Propagation for Neural Networks with Local Renormalization Layers
- Transformer Interpretability Beyond Attention Visualization
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Interpreting CLIP's Image Representation via Text-Based Decomposition
- ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
- Vision Transformers Don't Need Trained Registers
- Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
- Ref-Diff: Zero-shot Referring Image Segmentation with Generative Models
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
- No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
- What the DAAM: Interpreting Stable Diffusion Using Cross Attention
- Learning Visual Grounding from Generative Vision and Language Model
- Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
- Learning Referring Video Object Segmentation from Weak Annotation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models