CE 4 L: Continual Ego, Exo, and Ego-Exo Learning
cs.CV, cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 23 pages. Accepted by ICML 2026
Code: https://github.com/AnAppleCore/CE4L
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
- LoRA: Low-Rank Adaptation of Large Language Models
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Continual Reinforcement Learning by Planning with Online World Models
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
- OpenVLA: An Open-Source Vision-Language-Action Model
- The Power of Scale for Parameter-Efficient Prompt Tuning
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- R3M: A Universal Visual Representation for Robot Manipulation
- CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
- MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning
- Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- Domain Generalizable Continual Learning
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
- InternVideo: General Video Foundation Models via Generative and Discriminative Learning
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models