LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations
cs.LG, cs.AI, cs.CV
Submitted: 2026-09-23
Updated: 2026-09-24
Terminology
Sources
- Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Perception Encoder: The best visual embeddings are not at the output of the network
- Is the Modality Gap a Bug or a Feature? A Robustness Perspective
- Unsupervised Alignment of Embeddings with Wasserstein Procrustes
- When Models Manipulate Manifolds: The Geometry of a Counting Task
- Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
- TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- MTEB: Massive Text Embedding Benchmark
- Expanding Language-Image Pretrained Models for General Video Recognition
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- Scaling Language-Centric Omnimodal Representation Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks