Omni-Interactive Universal Embedder
cs.AI, cs.CV
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/openclaw/openclaw
Terminology
Sources
- YouTube-8M: A Large-Scale Video Classification Benchmark
- MAEB: Massive Audio Embedding Benchmark
- Perception Encoder: The best visual embeddings are not at the output of the network
- SAM 3: Segment Anything with Concepts
- e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
- Qwen2-Audio Technical Report
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
- Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- Efficient and High-Fidelity Omni Modality Retrieval
- E5-V: Universal Embeddings with Multimodal Large Language Models
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Learning Personalized Agents from Human Feedback
- ReMatch: Boosting Representation through Matching for Multimodal Retrieval
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- SAM Audio: Segment Anything in Audio
- Representation Learning with Contrastive Predictive Coding
- MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection