UniData: Universal Multimodal Instruction Generation Pipeline
cs.AI, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- Program Synthesis with Large Language Models
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- The Llama 3 Herd of Models
- Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
- SEED-Bench-2: Benchmarking Multimodal Large Language Models
- VideoChat: Chat-Centric Video Understanding
- MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- GPT-4o System Card
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
- mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection