COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
cs.AI
Submitted: 2026-06-27
Updated: 2026-08-28
Terminology
Sources
- Qwen3-VL Technical Report
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Emerging Properties in Unified Multimodal Pretraining
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
- AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
- Step1X-Edit: A Practical Framework for General Image Editing
- JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
- Qwen2.5-VL Technical Report
- InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
- Emu3: Next-Token Prediction is All You Need
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
- UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection