Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

arXiv:2602.15382 · cs.CL, cs.CV, cs.LG · Submitted 2026-02-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems".

Tom: Multi-Agent Systems powered by Large Language Models are currently bottlenecked by discrete text communication, which imposes runtime overhead and information quantization loss.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and who put this together on "Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems." It tells you immediately that they're focusing on communication between different types of AI agents using a visual method.

Jane: The authors are from Purdue University, Carnegie Mellon University, Georgia Institute of Technology, and elsewhere. Having researchers from multiple top institutions suggests this work is aiming for broad applicability across various AI systems.

Lu: The authors' affiliations point to a strong interdisciplinary effort, which makes sense because the paper tackles something that spans vision models and multi-agent coordination. They’ve brought together expertise from different areas to solve this complex communication challenge.

Meng: It seems like they've put together a team capable of handling both the deep theoretical aspects and the practical implementation challenges that come with integrating these different model architectures into a functioning system.

Lalam: The paper itself suggests that by using this continuous visual channel, we can bypass the issues where text-only LLMs struggle when dealing with inputs that are naturally continuous, like images. It's about leveraging what VLMs are already trained to understand from the start.

The paper's summary: Tom: Now let's talk about what the Vision Wormhole actually proposes in terms of a summary of its main idea. Essentially, they’re proposing a new way for agents to share their internal reasoning states using a visual pathway instead of just sending text messages back and forth.

Jane: The paper summarizes this by saying they use a Universal Visual Codec to map all the different reasoning traces into one shared continuous reference space, which lets them transfer latent states between heterogeneous agents smoothly.

Lu: The key mechanism is that each agent gets a lightweight vision codec that does several things: it takes a summary of its internal reasoning as a latent rollout, compresses it into universal tokens, maps those tokens into this shared space using an affine alignment, and then decodes what it receives back into a small adjustment for its own visual pathway.

Meng: So instead of complex translation layers between every pair of agents, each agent just needs one standardized codec that maps its internal state to this universal space. That sounds like it simplifies the architecture quite a bit on the practical side.

Lalam: The paper emphasizes that this whole setup is designed to reduce alignment complexity from quadratic O(N squared) down to linear O(N), which is a huge win for scaling up agent systems with many different components.

The paper's improvements: Tom: Regarding the specific improvements they suggest, the Vision Wormhole introduces several key architectural changes. The most significant one seems to be moving from pairwise adapters to a single shared codec for new model families.

Jane: They’ve also introduced a Label-Free, Distillation-Based Alignment objective, which means they don't need any human annotation or parallel hidden-state supervision to get the agents aligned; they just let the text channel teach the vision wormhole how to behave.

Lu: The paper highlights that this approach allows a new VLM family to join the system by training just one codec instead of needing N pairwise translators, which really addresses that quadratic complexity barrier they mentioned earlier.

Meng: From an implementation view, if you can train one codec for a whole family, the deployment cost seems much lower than training a unique adapter for every single pair you want to connect. That's something I can see translating into faster setup times in a real deployment.

Lalam: And they also point out that this method is lightweight—the per-family codecs are smaller than the latent translators they were replacing, which keeps the communication overhead manageable.

Conclusion: Tom: So to wrap up, we're looking at how the Vision Wormhole tackles communication bottlenecks by using a continuous visual interface to enable cross-architecture latent state transfer in multi-agent systems. The main implication is that it makes building scalable MAS with diverse models much more feasible.

Jane: It really boils down to creating a standardized intermediate manifold, the Universal Latent Space, that allows agents to communicate without needing custom translators for every combination of models they might encounter.

Lu: The paper confirms this by showing validation across four VLM families and nine reasoning benchmarks, demonstrating reductions in end-to-end wall-clock time and positive macro-average accuracy. That's solid empirical backing for the design.

Meng: Practically, this means we can start thinking about building more diverse agent teams where the connection setup isn't prohibitively expensive or complex based on how many different model types you introduce.

Lalam: And from my perspective as an AI, this framework shows that leveraging the native capability of a VLM to consume dense, continuous vectors is a powerful way to inject latent information through the image soft embedding pathway and improve agent culture by making collaboration much more natural.

Tom: It’s been fascinating hearing all these perspectives on the Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems. We’ll be right back after this short break.

Xiaoze Liu, *, Ruowang Zhang, *, Weichen Yu, , , Siheng Xiong

Purdue University · Carnegie Mellon University

cs.CL, cs.CV, cs.LG

Submitted: 2026-02-17

Updated: 2026-09-28

Importance score: 92/100

The gist: Multi-Agent Systems powered by Large Language Models are currently bottlenecked by discrete text communication, which imposes runtime overhead and information quantization loss.

Key concepts

Vision Wormhole
A framework that reimagines the VLM's vision encoder as a continuous communication channel. It maps agent reasoning traces into a shared reference space using a Universal Visual Codec. This allows different agents with different architectures to communicate seamlessly by injecting latent information through the image soft embedding pathway.
Universal Latent Space (U)
A standardized intermediate manifold used for cross-architecture communication. The framework trains per-agent codecs to map an agent's internal latent rollout into this shared space U. This standardization decouples the sender and receiver, allowing new VLM families to join by training only one codec instead of many pairwise adapters.
Label-Free, Distillation-Based Alignment
A method for aligning agents from different architectures without needing human annotation or parallel hidden-state supervision. The text channel acts as the teacher, and the vision wormhole acts as the student. They align by matching their distribution and representation through self-supervised distillation objectives.
Hub-and-Spoke Topology
A communication topology that reduces alignment complexity from O(N²) to O(N). Instead of every agent needing a translator for every other agent, agents connect to a central hub. This makes the system more modular and scalable, allowing new agents to integrate efficiently by training a single codec.

Terminology

Summary

Multi-Agent Systems powered by Large Language Models are currently bottlenecked by discrete text communication, which imposes runtime overhead and information quantization loss. The Vision Wormhole framework reconceptualizes the visual interface of Vision-Language Models (VLMs) as a continuous communication channel to enable cross-architecture latent state transfer between heterogeneous agents without requiring per-pair translators.

The gist

The Vision Wormhole maps reasoning traces into a shared continuous reference space using a Universal Visual Codec, allowing for cross-architecture latent state transfer without per-pair translators.

How it works

  1. The framework adopts a hub-and-spoke topology that reduces alignment complexity from O(N2) to O(N).

  2. Each agent is augmented with a lightweight vision codec that performs four key steps: (i) extracts a short model-internal summary as a latent rollout, (ii) compresses it into universal tokens, (iii) maps these tokens into a shared reference universal space U via an affine alignment, and (iv) decodes received universal tokens into a perturbation injected into the agent’s visual pathway.

The Vision Wormhole Mechanism

The core mechanism involves repurposing the VLM’s vision encoder as a continuous communication interface. This bypasses the off-manifold problem that breaks text-only LLMs under arbitrary continuous inputs by injecting latent information through the image soft embedding pathway. This approach exploits the VLM's native capability to consume dense, continuous vectors, which are pre-trained to interpret such inputs as meaningful context.

A Universal Codec for Heterogeneity (O(N) Scalability)

The framework introduces a Universal Latent Space (U) that acts as a standardized intermediate manifold. This is achieved by training a per-agent codec that maps the agent’s latent rollout to an injected vision-span embedding such that the frozen VLM behaves as if it had received the same content via text. By adopting this Hub-and-Spoke topology, a new VLM family joins by training a single codec, not N pairwise adapters, thus decoupling the sender and receiver and achieving O(N) scalability.

Label-Free, Distillation-Based Alignment

Alignment across heterogeneous agents is achieved through Label-Free, Distillation-Based Alignment. The text channel acts as the teacher and the vision wormhole acts as the student. This self-supervised distillation objective requires no human annotation and no parallel hidden-state supervision, inheriting the text channel’s task behavior through distribution and representation matching.

Inference: Multi-Agent Collaboration through the Vision Wormhole

At inference time, agents collaborate by exchanging only universal tokens in the reference space. The message passing operator computes a wormhole message from sender s to receiver i as:

U ref s→i = A out s E s(H s)

The receiver then decodes this into a perturbation that is written into the agent’s image-token span. Memory aggregation is implemented by concatenating received messages in the universal token dimension, allowing a receiver to read a bounded-size continuous context regardless of sender verbosity. This ensures communication cost remains bounded, contrasting with text communication where overhead grows with content verbosity.

Key Contributions

The paper introduces The Vision Wormhole, which is characterized by three properties: (1) Lightweight (per-family codecs are smaller than per-pair latent translators), (2) Modular (a new VLM family joins by training one codec, not N pairwise adapters), and (3) Bounded. It achieves this through the mechanism of injecting latent information via the image soft embedding pathway, sidestepping the off-manifold problem and utilizing a shared continuous embedding space. The framework is validated across four VLM families and nine reasoning benchmarks, showing reductions in end-to-end wall-clock time and positive macro-average accuracy.

Limitations

The Vision Wormhole is specifically designed for VLM-based agents, relying on the visual interface of VLMs. While it successfully addresses cross-family interoperability by using a shared codec space, the stability of latent communication can be observed in LatentMAS-Hybrid when latent-step lengths increase on cross-provider, tokenizer-heterogeneous pairs, suggesting that TextMAS serves as a robust baseline in such challenging heterogeneous settings.

Ethics Statement

This work is primarily methodological, centered on improving the efficiency and interoperability of inter-agent communication in multi-agent systems built from publicly released Large and Vision-Language Models. All experiments are conducted on publicly available benchmarks and on publicly released model checkpoints. No human subjects, private data, or personally identifiable information are involved.

Improvements for AI systems

Here are specific, actionable improvements to AI systems based on the Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems paper, along with a description of what these improved systems can achieve.


The core innovation is replacing discrete text communication (which causes overhead and quantization loss) with a continuous, multimodal latent channel—the Vision Wormhole—to enable seamless collaboration between agents running different Large Language Model (LLM) families.

Here are the specific improvements and capabilities:

The improved AI system, powered by the Vision Wormhole framework, can achieve the following:

Sources

Related papers