DynaTokens: Controlling Token Dynamics for Continual Video-Language Understanding
cs.CV, cs.LG
Submitted: 2026-03-02
Updated: 2026-10-04
Comments: Accepted to the EMNLP 2026 Main Conference
License: http://creativecommons.org/licenses/by/4.0/
The gist: Continual VideoQA with multimodal LLMs remains challenging because sequential adaptation induces task interference, while storing task-specific prompts becomes impractical as task sequences grow.
Terminology
Abstract
Continual VideoQA with multimodal LLMs remains challenging because sequential adaptation induces task interference, while storing task-specific prompts becomes impractical as task sequences grow. We introduce DynaTokens, a transformer-based token generator that dynamically produces fine-tuning tokens on demand, enabling task-adaptive prompt updates through shared generation weights. To mitigate forgetting, we introduce meta-learning-inspired regularisers that look ahead to avoid task-specific sharp update directions while anchoring the evolving generator to prior-task behaviours. We theoretically connect this objective to sharpness-aware optimisation, showing how it favours flatter cross-task minima and improves retention. DynaTokens combines gradient-free routing based on robust pretrained token and visual embeddings with lightweight auxiliary multimodal supervision, reducing router drift during continual adaptation. Across standard continual VideoQA benchmarks, DynaTokens achieves higher average accuracy and substantially lower forgetting than strong baselines. It also improves zero-shot generalisation and remains effective in longer domain-incremental sequences with extended task shifts. Finally, we introduce a challenging ImageQA->VideoQA protocol and show that DynaTokens enables robust cross-modal continual transfer.
Sources
- Gemini: A Family of Highly Capable Multimodal Models
- On First-Order Meta-Learning Algorithms
- Representation Learning with Contrastive Predictive Coding
- DeepSeek-V3 Technical Report
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Progressive Neural Networks
- Instant Personalized Large Language Model Adaptation via Hypernetwork
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models