Gestalt: Large Multimodal Interplay Model
cs.CV
Submitted: 2026-09-30
Updated: 2026-10-08
Code: https://github.com/GeWu-Lab/Gestalt
Project page: https://gewu-lab.github.io/Gestalt
Terminology
Sources
- Qwen2.5-VL Technical Report
- Kimi K3: Open Frontier Intelligence
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- FineVision: Open Data Is All You Need
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- Emerging Properties in Unified Multimodal Pretraining
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
- MIBench: Evaluating LMMs on Multimodal Interaction
- Core Knowledge Deficits in Multi-Modal Language Models
- Gated Multimodal Units for Information Fusion
- Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Emu3: Next-Token Prediction is All You Need
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models