Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos
cs.CV, cs.LG
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/Leoxu3/PointTube-JEPA
Terminology
Sources
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- 3D-JEPA: A Joint Embedding Predictive Architecture for 3D Self-Supervised Representation Learning
- STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- Cross4D-JEPA: Dense Cross-modal Correspondence Distillation for 4D Point Cloud Representation Learning
- Diffusion Masked Pretraining for Dynamic Point Cloud
- Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models