MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models

arXiv:2609.39182 · cs.CV, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models".

Tom: World Models are emerging as a frontier in computer vision, but their robustness remains largely unexplored,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: We’ve covered a lot about "MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models," starting with what latent hallucination is and moving into how MEND tackles it with its detection, localization, and correction loop.

Jane: And we looked at the implications of this work across different environments like WALL and POINTMAZE, focusing on the authors' findings regarding AUROC detection scores up to zero point eight zero and per-token AUPRC reaching zero point eight seven for localization accuracy.

Lu: The main point is that MEND introduces a single conditional score network that acts as a unified mechanism for detecting, localizing, and correcting latent hallucination at inference time without needing ground-truth error labels.

Meng: From an engineering viewpoint, the practical contribution is the ability to deploy a reliable method for auditing world models by adding an intrinsic correction loop that tries to move predictions toward a valid state, even with the limitation that it only recovers errors tangent to the data manifold.

Lalam: Ultimately, this work points toward building AI systems where internal checks are not just about accuracy but also about self-correction at the latent level, which could fundamentally shift how we approach model robustness and trustworthiness.

Tom: Exactly. This paper gives us a solid technique to actively mitigate compounding errors in autoregressive predictions by using a score field derived from real transitions for inference-time repair.

Jane: It’s important to remember that while it’s not perfect, it successfully demonstrated how errors remain spatially sparse, allowing for targeted intervention rather than broad corrections.

Lu: And the future work seems focused on integrating this concept of score-based correction into larger planning architectures to ensure long-term validity in complex scenarios.

Meng: I think the key takeaway for us is that we now have a blueprint for making world models more reliable by adding this specific detection and correction module.

Lalam: We’re excited about how this kind of intrinsic mechanism can influence the development of future generative AI, moving toward systems that are inherently more self-aware in their predictions.

Conclusion: Tom: So, we've been deep in the weeds of MEND, and now we’re getting to wrap things up with what this paper actually means for us on a broader scale.

Jane: I think focusing on the title itself really helps—"MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models." It sounds like it’s tackling a really messy problem in world models without needing all that tedious ground-truth labeling.

Lu: Exactly! The authors are presenting a single conditional score network that does three things at once: it detects the error, pinpoints where the error is, and even suggests how to fix it during inference. That level of integrated functionality is really something special in how we think about model robustness.

Meng: From my side as an engineer, I’m curious about the practical implication of not needing those expensive labels. If we can do this without massive datasets specifically annotated for failures, that opens up a huge path for deploying world models in real-world scenarios where labeling is just impossible.

Lalam: I see it through the lens of culture here. If we can build AI systems that are inherently capable of auditing their own predictions and correcting mistakes internally, it fosters a different kind of trust in these complex generative systems. It moves us closer to more reliable reasoning structures within the AI ecosystem.

Tom: That’s a huge point, Lalam—building intrinsic self-correction. The paper shows that this isn't just fixing a single wrong prediction; it’s about correcting the underlying structure of how the model generates those states.

Jane: And when you look at the results, even though it only corrects errors that are tangent to the data manifold, we still see tangible reductions in latent error during rollouts. It shows that targeted intervention has a real effect on overall model performance.

Lu: That spatial sparsity of the error is what makes this approach so elegant from a theory standpoint—it suggests that even when the whole state is wrong, the deviation can be localized to just a few tokens. I’m already thinking about how this could apply to long-horizon planning tasks where errors compound quickly.

Meng: That localization aspect is crucial for me because it makes it actionable. Being able to isolate the problem patch rather than just getting a generic error signal means we can focus our compute on the fix, which is essential for making this viable in production environments.

Lalam: For culture, I think this work signals a shift from simply training bigger models to training smarter systems that can manage their own internal uncertainty and self-repair. That’s where the long-term value lies for how we design and trust these advanced AI applications.

Tom: So, to wrap up this segment, MEND gives us a way to check our world models on the fly without needing perfect labels, which points toward a future where AI systems are less brittle when they encounter novel situations.

Jane: It’s clear that the work by the authors has provided a solid framework for auditing and repairing latent states in world models.

Lu: This technique opens up fascinating avenues for exploring how we can systematically enforce consistency across complex, multi-step generative processes.

Meng: I’m eager to see how this detection loop scales when we move from simple environments like WALL to the more complex dynamics of real-world physics simulations.

Lalam: We’re going to keep digging into these concepts because this work genuinely contributes a new mechanism for ensuring AI systems are more reliable and transparent in their generative processes.

Tom: Next up, we're going to look at some of the specific experimental results on WALL and POINTMAZE to see exactly how robust this detection is under different conditions.

Ali Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar

University of Melbourne · Monash University

cs.CV, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: Accepted at DICTA 2026 (International Conference on Digital Image Computing: Techniques and Applications). Camera-ready version

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: World Models are emerging as a frontier in computer vision, but their robustness remains largely unexplored, leading to the identification of latent hallucination—where predicted next latent states

Key concepts

Latent Hallucination
This occurs when a model predicts a next state that is mathematically impossible given the true rules of the environment. It means the predicted state is either unreachable or results from an invalid action, causing errors to compound rapidly in sequential predictions.
Conditional Score Field
This is the output of MEND, which acts as a single network providing three functions: detection (measuring distance from valid states), localization (showing how much each part of the image needs to move), and correction (defining the direction to move toward a valid state).
Denoising Score Matching
This is the training method used. The network learns by trying to reverse a noise process applied to valid next states. It teaches the model to accurately predict the conditional score of real transitions, allowing it to learn the structure of valid world dynamics without needing any error labels.
Tweedie’s Formula
This mathematical formula is used during correction. It uses the learned score field and its variance to calculate a new state estimate that is pulled toward the set of valid successor states, effectively defining the precise direction needed to fix the hallucination.

Terminology

Summary

World Models are emerging as a frontier in computer vision, but their robustness remains largely unexplored, leading to the identification of latent hallucination—where predicted next latent states decode to scenes that never occur—which compounds autoregressively and silently. This paper introduces MEND (Masked Empirical-Bayes Neural Denoising), a single conditional score network that serves as a unified mechanism for detecting, localizing, and correcting this latent hallucination at inference time without requiring ground-truth error labels.

The gist

MEND is a single conditional score network trained by denoising score matching on real transitions, whose score field detects hallucination magnitude, localizes it to specific image patches, and defines an inference-time correction direction.

What is Latent Hallucination?

Latent hallucination occurs when a model produces a next latent state that lies outside the set of valid successor states reachable under the true environmental dynamics, i.e., it is Unreachable state (zˆt+1 ∈/ R(zt)) or an Incorrect action outcome (zˆt+1 ∈ R(zt)∖ R(zt, at)). Because rollouts are autoregressive, a single hallucinated step reduces the probability of the entire rollout and causes the error to compound with depth. The error is characterized as being spatially sparse: a small subset of tokens carries a disproportionate share of the total error.

How MEND Works

MEND instantiates a general detect–localise–correct loop using a single differentiable object: a conditional score field, denoted as the network output, sθ(˜z zt, σ). This score field serves three complementary roles simultaneously:

  1. Detection: The scalar D =∥sθ(ˆzt+1 zt)∥ squared measures how far the prediction is from the set of reachable states.

  2. Localisation: The per-token field∥sθ(ˆzt+1 zt)n∥ is a direct estimate of how far each token must move, localizing errors to specific image patches.

  3. Correction: By Tweedie’s formula, the posterior mean of the clean state is zˆt+1 + σ squared sθ, meaning the score itself defines the correction direction.

The score network is trained using denoising score matching (Eq. 4), which learns the conditional score ∇ log qσ(˜z zt) of a Gaussian smoothing of valid next states, without requiring any hallucination or accuracy labels. This training ensures that the network learns the exact conditional score and that the resulting field is world-model-agnostic.

The Detection, Localisation, and Correction Loop

The loop operates iteratively:

  1. Evaluate the detector: Calculate D˜ = (D − µacc)/σacc to normalize against known correct predictions.

  2. Localize: Identify the most suspicious tokens using the rank of the displacement vector 'd' (mraw ← rank(d)).

  3. Masked Correction: Apply a correction step by moving only the suspicious tokens, while freezing confident ones, following Tweedie’s formula to move toward the valid manifold. This process repeats until a predefined budget is reached or validity is achieved.

Experimental Results and Analysis

MEND was evaluated on two environments: WALL (Markov predictor) and POINTMAZE (three-frame history). Detection performance was measured by AUROC, while localisation accuracy was per-token AUPRC against the ground-truth error map.

(Table I results are summarized below)

Detection is robust, achieving an AUROC of up to 0.80 without using actions on WALL and exceeding a single-Gaussian density baseline in POINTMAZE. Localisation is highly effective, with per-token AUPRC reaching up to 0.87 for POINTMAZE hallucinations, demonstrating that the error remains concentrated in a small subset of tokens even when the whole state is far off the manifold.

Correction reliably reduces latent error; for example, on WALL, MEND achieved a reduction of-6.4% in latent error over 1 step compared to no correction. However, correction is bounded by the part of the error that lies tangent to it [the data manifold], meaning only the normal part of the displacement is recoverable.

Ablation and Future Directions

Ablation studies confirm that using only one factor from the Bayes decomposition (D1: reachability) provides strong detection and localisation, while including D2 (action-consistency) offers little benefit in deterministic settings because "the inverse dynamics residual is largely independent of whether a prediction is valid or hallucinated.

Improvements for AI systems

Here are the specific improvements to AI systems based on the MEND framework described in the paper:

  1. Improved Robustness Against Latent Hallucination in World Models: The system will no longer suffer from compounding errors where a single hallucinated state corrupts subsequent predictions.

  2. Inference-Time Correction of Errors: The system can actively correct predicted latent states during inference by iteratively nudging suspicious tokens toward the known valid data manifold, leading to reliable reductions in single-step latent error and improved overall prediction accuracy.

  3. Label-Free Detection and Localization: The system utilizes a single conditional score field to simultaneously detect if a prediction is hallucinated (global score), localize exactly which image patches are responsible for the error (per-token localization map), and determine the direction for correction.

  4. Data-Efficient Error Mitigation: Unlike methods requiring expensive data collection and retraining on failure cases, MEND operates entirely at inference time on a frozen world model, making it practical for deployment in environments where fine-tuning is not feasible or desirable.

  5. Spatially Sparse Error Handling: The system leverages the finding that hallucination is spatially sparse (a small subset of tokens carries the error), allowing for highly targeted and efficient correction procedures rather than uniform state updates.

This improved AI system can perform:

  1. Accurate, reliable planning and simulation in complex environments (like navigation or robotics) by ensuring that the predicted future states are physically plausible and consistent with the environment's true dynamics, even when the model imagines a path it never actually takes.

  2. Enhanced safety guarantees for autonomous agents by providing an immediate check on the validity of predicted actions or states before execution, preventing errors stemming from latent space distortions.

  3. More efficient use of existing world models in real-time applications by providing a mechanism to clean potentially faulty outputs without needing external ground-truth error labels.

Abstract

World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.

Sources

Related papers