Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning
Yuhang Cao
Nanjing University
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 15 pages, 5 figures, 1 table
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: Continuous Interaction Diffusion (CID) is a diffusion-native model–runtime architecture that integrates tool interaction into iterative denoising, addressing the mismatch between turn-based tool
Terminology
Summary
Continuous Interaction Diffusion (CID) is a diffusion-native model–runtime architecture that integrates tool interaction into iterative denoising, addressing the mismatch between turn-based tool protocols and continuously revisable diffusion generation. The paper formalizes the architecture, runtime, and training objectives, and defines an evaluation protocol for task quality and end-to-end efficiency, but makes no empirical performance claims.
The paper states: "We introduce Continuous Interaction Diffusion (CID), a diffusion-native model–runtime architecture that integrates tool interaction into iterative denoising. CID separates a model-read-only fact channel, a thought channel represented by a Typed Cognitive Tensor, and a display channel. Information needs can emerge before a textual or JSON call is fully serialized, allowing perceptual bindings to launch external reads while denoising continues. Returned results are projected into the evolving thought state and can revise earlier cognition and display regions. Persistent bindings reuse static results without repeated external execution and refresh changing sources when needed. CID is designed to expose evidence earlier, overlap tool latency with model computation, reduce duplicate external work, and preserve useful computation after new evidence arrives."
The architecture maintains the system state as Ss = (Fs, Ts, Ys, Bs, Js), where Fs is the fact channel, Ts the thought channel, Ys the display channel, Bs the active perceptual bindings, and Js the in-flight external jobs. The fact channel contains information controlled outside the generative state that the model can read but cannot overwrite; the thought channel is a structured cognitive field represented as a Typed Cognitive Tensor (TCT); and the display channel is the evolving token canvas that converges to the user-visible response.
The Typed Cognitive Tensor represents cognitive cells as cs,i = (hs,i, rs,i, as,i, qs,i, us,i, τs,i, ls,i), where hs,i is continuous semantic content, rs,i is a soft role distribution (e.g., hypothesis, information need, percept, plan, constraint, conclusion), as,i are sparse symbolic anchors (tool identifiers, entities, paths, numeric values, schema fields, output spans), qs,i are links to fact items, bindings, or other cognitive cells, us,i is epistemic or binding uncertainty, τs,i is a local diffusion level controlling editability, and ls,i is lifecycle state (active, waiting, stable, retired).
Tool use is represented as a role within the thought channel, not as a separate text region. A model-side intent adapter reads the cognitive field and registered source descriptors, exposing candidate needs ik = (nk, πk, αk, ωk, ηk, χk), where nk is a continuous need embedding, πk a distribution over registered sources or tools, αk partially bound arguments, ωk uncertainty, ηk freshness or persistence demand, and χk links to affected cognitive or display regions. This separates three stages: need emergence → source selection → argument binding.
Persistent perceptual bindings bj = (ij, mj, αj, ρj, zj, χj) connect an information need to an external source, its current arguments, refresh policy, cached state, and affected cognitive regions. The runtime distinguishes external refresh (re-executing or re-reading a source) from cognitive refresh (recomputing how an already available external value should influence the current cognitive state without repeating external I/O). For static sources, the runtime can reuse cached results while repeatedly re-projecting their meaning into the evolving thought state; for changing sources, the binding can request refreshed values or consume a stream.
The percept encoder converts each tool result into a context-dependent projection Pjs = Eψ(vj, pj, Ts, Ys), whose interpretation depends on the current thought and display states. The thought denoiser integrates these projections with gated cross-attention or residual updates: Tes = Ts + Σj∈Bs gj,s A(Ts, Pjs), where gj,s is predicted from need persistence, relevance, and source freshness. New evidence may change the editability of existing cells through local reopening, adjusting per-cell diffusion levels so that conflicting cognitive cells become more diffused and revisable while supported cells stabilize.
The runtime is asynchronous: model denoising may continue while read-only work is outstanding, but only until a learned current-information equilibrium signal fires. The runtime then pauses if required work remains and resumes on external progress; otherwise it forms a terminal candidate. External work is keyed by source and canonicalized arguments, and a completion applies only when an active binding still has the same work key, making obsolete work cancellable or discardable. Equivalent jobs are deduplicated without suppressing repeated perception.
The training objective is L = LT + λY LY + λI Lintent + λB Lbind + λP Lassim + λR Lrefresh + λG Lground + λC Lconv, where the first two terms train thought and display diffusion, and the others train intent exposure, binding, event-conditioned revision, refresh behavior, symbolic grounding, and the equilibrium signal. Training randomizes event arrival, source freshness, and cache state; before a needed event arrives, the model should maintain uncertainty and expose a binding need rather than fabricate a value; after arrival, the model should assimilate the percept, revise affected cells, preserve unrelated cognition, and update the display.
The evaluation protocol defines five research questions: RQ1 (do useful information needs emerge before explicit calls), RQ2 (can persistent bindings hide external latency), RQ3 (does repeated perception improve assimilation), RQ4 (is the Typed Cognitive Tensor necessary), and RQ5 (is diffusion essential). Task families include static reference and copying, dynamic state tracking, delayed retrieval, streaming evidence, and competing sources. Baselines include autoregressive ReAct with blocking tools, an event-driven asynchronous autoregressive agent, a dLLM using conventional explicit tool-call regions, P-ReAct as instantiated by DLLM-Searcher, and CID variants with one-time percept injection, persistent static re-projection, full CID with static and dynamic refresh, and oracle bindings. Metrics cover task quality, interaction latency, assimilation, binding behavior, and revision quality.
The paper's contributions are: (1) formulating the mismatch between turn-based tool protocols and continuously revisable diffusion generation, defining continuous tool interaction as a model–runtime problem; (2) introducing a three-channel architecture with an externally controlled fact channel, a typed continuous thought channel, and a discrete display channel; (3) defining persistent perceptual bindings that distinguish repeated cognitive refresh from repeated external execution, supporting both static and changing information sources; and (4) specifying an asynchronous runtime, local diffusion clocks, model-side intent interface, training objectives, and a concrete evaluation protocol for read-only tools.
The paper discusses limitations including architectural complexity, supervision for cognition, latent-state stability, fact-channel policy, compute and memory cost, privacy and unobservable reasoning, read-only scope, human cognition analogy, and multimodal extension. The conclusion states: "We introduced Continuous Interaction Diffusion, a model–runtime co-designed architecture that replaces discrete tool-call rounds with persistent asynchronous perception. CID separates externally controlled facts, continuous typed cognition, and user-visible text; represents information needs inside a Typed Cognitive Tensor; and maintains perceptual bindings that can re-project static results or refresh changing sources throughout denoising. CID makes a falsifiable claim: diffusion should let new external information revise existing cognition and display state while useful computation continues."
Improvements for AI systems
Based on this paper, here are the specific improvements I can make to AI systems:
1. Implement asynchronous tool interaction during generation. Instead of blocking generation on tool calls, I can expose information needs mid-generation, launch external reads in parallel with continued denoising, and assimilate results as they arrive—reducing end-to-end latency by overlapping I/O with computation.
2. Add a persistent perceptual binding layer. I can cache static tool results and re-project them into the evolving thought state multiple times without re-executing external calls, while selectively refreshing dynamic sources based on freshness policies—eliminating duplicate work and enabling evidence revision.
3. Introduce a typed cognitive tensor for structured reasoning. I can maintain a continuous thought field where each cell has a role (hypothesis, need, percept, plan), uncertainty, editability level, and links to facts or bindings—allowing targeted revision of conflicting cells while preserving stable ones, rather than regenerating the whole response.
4. Enable local diffusion-level control. I can dynamically adjust how editable each cognitive cell is: when new evidence conflicts, I increase diffusion (editability) for affected cells and decrease it for supported ones—enabling surgical updates to reasoning without destabilizing unrelated content.
5. Separate fact, thought, and display channels. I can keep externally controlled facts read-only, maintain a private continuous thought space, and only converge to user-visible text at the end—preventing the model from overwriting ground truth and allowing reasoning to evolve without premature commitment to output tokens.
6. Implement intent-driven tool selection before full argument binding. I can detect that a need exists (e.g., need current stock price
) before serializing the full API call, launch a source selection and partial binding early, then complete arguments as context clarifies—reducing the critical path for tool use.
7. Add a learned equilibrium signal for pausing. I can train the model to recognize when it has enough current information to proceed, avoiding unnecessary waits for in-flight jobs while still pausing when critical evidence is missing—balancing responsiveness with correctness.
8. Support event-conditioned revision training. I can train on randomized event arrival and cache states so the model learns to maintain uncertainty before evidence arrives, then assimilate percepts and revise only affected cells—improving robustness to non-deterministic tool latency.
9. Deduplicate equivalent external jobs without suppressing perception. I can key jobs by source and canonicalized arguments, discard obsolete work when bindings change, but still allow repeated cognitive projection of the same cached result—reducing redundant I/O while preserving assimilation quality.
10. Enable streaming evidence assimilation. For changing sources, I can consume streams and continuously update the thought state, allowing the display to converge progressively rather than in a single shot—useful for live data or iterative reasoning tasks.
What the improved AI system can do: It can answer questions that require live or evolving information (e.g., What's the current temperature and how does it affect my travel plan?
) by starting to reason, launching a weather API call mid-generation, revising its plan when the result arrives, and producing a final answer with lower latency and fewer redundant calls. It can handle multi-step research tasks where each fact check overlaps with continued reasoning, and it can correct earlier conclusions when new evidence contradicts them—all while preserving unrelated parts of its response. It can also operate with static references (e.g., a PDF) by caching and re-interpreting them as the reasoning evolves, without re-reading the file.
Abstract
Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. For autoregressive models, tool use naturally fits the generation process: the model emits a tool call, waits for the result, and then continues generating. Diffusion language models (dLLMs), however, reason by repeatedly refining many parts of their output in parallel, making this stop-and-resume interaction pattern unnecessarily restrictive. It can force tool decisions before the model's reasoning has stabilized, delay useful observations until a discrete call finishes, and introduce redundant refinement and tool execution, potentially hurting both task accuracy and inference efficiency. We introduce Continuous Interaction Diffusion (CID), a diffusion-native model--runtime architecture that integrates tool interaction into iterative denoising. CID separates a model-read-only fact channel, a thought channel represented by a Typed Cognitive Tensor, and a display channel. Information needs can emerge before a textual or JSON call is fully serialized, allowing perceptual bindings to launch external reads while denoising continues. Returned results are projected into the evolving thought state and can revise earlier cognition and display regions. Persistent bindings reuse static results without repeated external execution and refresh changing sources when needed. CID is designed to expose evidence earlier, overlap tool latency with model computation, reduce duplicate external work, and preserve useful computation after new evidence arrives. We formalize the architecture, runtime, and training objectives, and define an evaluation protocol for task quality and end-to-end efficiency. This first paper focuses on read-only tools and makes no empirical performance claims.
Sources
- Diffusion-LM Improves Controllable Text Generation
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
- Structured Denoising Diffusion Models in Discrete State-Spaces
- Large Language Diffusion Models
- Asynchronous Tool Usage for Real-Time Agents
- Simple and Effective Masked Diffusion Language Models
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Continuous Latent Diffusion Language Model
- Tool Calling is Linearly Readable and Steerable in Language Models
- Training Large Language Models to Reason in a Continuous Latent Space
- ReAct: Synergizing Reasoning and Acting in Language Models
- Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
- Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models
- Dream 7B: Diffusion Large Language Models
- Planned Diffusion
- DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection