SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
cs.CV, cs.LG
Submitted: 2026-09-30
Updated: 2026-09-30
Project page: https://semantoken.github.io
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
- Improving Flexible Image Tokenizers for Autoregressive Image Generation
- Cosmos World Foundation Model Platform for Physical AI
- VidTok: A Versatile and Open-Source Video Tokenizer
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- VideoGPT: Video Generation using VQ-VAE and Transformers
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models