Separators in Enhancing Autoregressive Pretraining for Vision Mamba
cs.CV, cs.AI
Submitted: 2026-03-04
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting considerable attention for its efficient handling of long-sequence tasks.
Terminology
Abstract
The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting considerable attention for its efficient handling of long-sequence tasks. Its inherent causal structure makes it particularly well suited for autoregressive pretraining. However, existing autoregressive pretraining methods in vision are largely limited to short-sequence settings and may not fully use Mamba's capacity to model longer contexts. To investigate this setting, we introduce SeparaTors for AutoRegressive pretraining (STAR), a new autoregressive pretraining method for Vision Mamba that explicitly marks the boundaries between different images. STAR increases the patch-token sequence length from 144 to 640 by packing four images and four separator clusters. This is approximately 4.4 times the ARM patch-token sequence length. The increase is achieved without changing the resolution of any individual image: we use 192 times192 inputs for autoregressive pretraining and 224 times224 inputs for downstream classification fine-tuning. With this long-sequence pretraining scheme, STAR-B achieves 83.5% EMA top-1 accuracy on ImageNet-1K after 1,600 epochs of pretraining. The learned representation also transfers beyond in-distribution classification: compared with ARM, STAR-B improves COCO box AP from 46.11 to 46.84 and mask AP from 40.74 to 41.45, while raising the mean top-1 accuracy across five ImageNet robustness benchmarks from 55.1% to 56.8%. Under the evaluated four-image setting, these results indicate that separator-based long-sequence pretraining improves recognition robustness and dense visual prediction relative to ARM.
Sources
- MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
- VMamba: Visual State Space Model
- GPT-4 Technical Report
- EfficientVMamba: Atrous Selective Scan for Light Weight Visual Mamba
- Autoregressive Pretraining with Mamba in Vision
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Mamba-R: Vision Mamba ALSO Needs Registers
- Mamba-UNet: UNet-Like Pure Visual Mamba for Medical Image Segmentation
- MambaOut: Do We Really Need Mamba for Vision?
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models