How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
cs.MM, cs.CV, cs.LG
Submitted: 2026-08-27
Updated: 2026-08-27
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Qwen2.5 Technical Report
Related papers
- ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
- A Rate-Distortion-Classification Approach for Lossy Image Compression