Subspace Alignment for Vision-Language Model Test-time Adaptation
cs.CV, cs.AI
Submitted: 2026-01-13
Updated: 2026-08-27
Comments: 23 pages, 11 figures
Code: https://github.com/zhichenz98/SubTTA_EMNLP26
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
- Mint: A Simple Test-Time Adaptation of Vision-Language Models against Common Corruptions
- Matcha: Mitigating Graph Structure Shifts with Test-Time Adaptation
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
- Panda: Test-Time Adaptation with Negative Data Augmentation
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
- Visual question answering: from early developments to recent advances -- a survey
- Language in the Flow of Time: Time-Series-Paired Texts Weaved into a Unified Temporal Narrative
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- MADE: Graph Backdoor Defense with Masked Unlearning
- ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification
- MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
- CATS: Mitigating Correlation Shift for Multivariate Time Series Classification
- A Collaborative Ensemble Framework for CTR Prediction
- Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting
- TUCKET: A Tensor Time Series Data Structure for Efficient and Accurate Factor Analysis over Time Ranges
- LLaMA: Open and Efficient Foundation Language Models
- Tent: Fully Test-time Adaptation by Entropy Minimization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models