Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
cs.CV, cs.LG
Submitted: 2025-07-24
Updated: 2026-08-25
Code: https://github.com/cominder/Iwin-Transformer
Terminology
Sources
- TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
- Reformer: The Efficient Transformer
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Linformer: Self-Attention with Linear Complexity
- Longformer: The Long-Document Transformer
- Generating Long Sequences with Sparse Transformers
- LocalViT: Analyzing Locality in Vision Transformers
- Token Merging: Your ViT But Faster
- Gaussian Error Linear Units (GELUs)
- Adam: A Method for Stochastic Optimization
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Big Transfer (BiT): General Visual Representation Learning
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Aggregated Residual Transformations for Deep Neural Networks
- Video Swin Transformer
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
- Autoregressive Image Generation without Vector Quantization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models