The Telephone Game: Evaluating Semantic Drift in Unified Models
cs.CV, cs.CL
Submitted: 2025-09-04
Updated: 2026-08-27
Code: https://github.com/mollahsabbir/telephone-game-semantic-drift
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Unified models (UMs) combine visual understanding (I2T) and generation (T2I) in a single framework.
Terminology
Abstract
Unified models (UMs) combine visual understanding (I2T) and generation (T2I) in a single framework. We focus on T2I and I2T, where cross-consistency---what a model understands, it should be able to generate---is a promise of unification and a necessity when composing both capabilities. Yet, existing benchmarks evaluate them in isolation: FID/GenEval for T2I; MME/MMBench for I2T. We show this gap is consequential: models scoring competitively on these benchmarks can fail severely when understanding and generation are composed, losing entities, attributes, spatial relations, and counts, resulting in semantic drift. To quantify drift, we introduce the Semantic Drift Protocol (SDP), inspired by the Telephone Game: starting from a caption or image, we alternate I2T and T2I over multiple generations and measure semantic preservation. We propose Mean Cumulative Drift (MCD), an embedding-based measure of content retention across three representation spaces, and Multi-Generation GenEval (MGG), extending GenEval's object-level compliance scoring across generations. To stress-test models beyond COCO-style data, we create a benchmark of 400 image-text pairs sampled from NoCaps and DOCCI, emphasizing novel objects and fine-grained descriptions. Applying SDP to seven models reveals that drift varies dramatically and is not predicted by single-pass scores: BAGEL retains high semantic fidelity over multiple generations, while VILA-U and Janus variants collapse within five generations, despite comparable isolated metrics. We identify six recurring failure modes and find degradation is typically catastrophic rather than gradual: once a critical error occurs, subsequent generations compound it. SDP exposes failure modes that single-pass benchmarks miss, enabling a more faithful assessment of unified model reliability. Code and benchmark: https://github.com/mollahsabbir/telephone-game-semantic-drift
Sources
- VQA: Visual Question Answering
- Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
- MaskGIT: Masked Generative Image Transformer
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- Meshed-Memory Transformer for Image Captioning
- Emerging Properties in Unified Multimodal Pretraining
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
- Evaluating Object Hallucination in Large Vision-Language Models
- Visual Instruction Tuning
- MMBench: Is Your Multi-modal Model an All-around Player?
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Scaling Open-Vocabulary Object Detection
- BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
- DOCCI: Descriptions of Connected and Contrasting Images
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Learning Transferable Visual Models From Natural Language Supervision
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models