Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding
cs.CV
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- ETC: Encoding Long and Structured Inputs in Transformers
- Longformer: The Long-Document Transformer
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Generating Long Sequences with Sparse Transformers
- Rethinking Attention with Performers
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- Reformer: The Efficient Transformer
- U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation
- SPFormer: Enhancing Vision Transformer with Superpixel Representation
- SAM 2: Segment Anything in Images and Videos
- Three-level Hierarchical Transformer Networks for Long-sequence and Multiple Clinical Documents Classification
- QuadTree Attention for Vision Transformers
- Ultra-Long Sequence Distributed Transformer
- ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling
- DeepSeek-OCR: Contexts Optical Compression
- Topology-Preserving Segmentation Network: A Deep Learning Segmentation Framework for Connected Component
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models