ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
cs.CV, cs.CL
Submitted: 2025-05-27
Updated: 2026-09-08
Comments: Camera-ready version for EMNLP 2026 Findings
Code: https://github.com/zooblastlbz/ID-Align
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- GPT-4 Technical Report
- NVLM: Open Frontier-Class Multimodal LLMs
- A Survey of Multimodal Large Language Model from A Data-centric Perspective
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- Round and Round We Go! What makes Rotary Positional Encodings useful?
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- InternLM2 Technical Report
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
- Base of RoPE Bounds Context Length
- Pointer Sentinel Mixture Models
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
- DeepSeek-V3 Technical Report
- Self-Attention with Relative Position Representations
- Encoding word order in complex embeddings
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models