MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models".
Tom: World Models are emerging as a frontier in computer vision, but their robustness remains largely unexplored,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: We’ve covered a lot about "MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models," starting with what latent hallucination is and moving into how MEND tackles it with its detection, localization, and correction loop.
Jane: And we looked at the implications of this work across different environments like WALL and POINTMAZE, focusing on the authors' findings regarding AUROC detection scores up to zero point eight zero and per-token AUPRC reaching zero point eight seven for localization accuracy.
Lu: The main point is that MEND introduces a single conditional score network that acts as a unified mechanism for detecting, localizing, and correcting latent hallucination at inference time without needing ground-truth error labels.
Meng: From an engineering viewpoint, the practical contribution is the ability to deploy a reliable method for auditing world models by adding an intrinsic correction loop that tries to move predictions toward a valid state, even with the limitation that it only recovers errors tangent to the data manifold.
Lalam: Ultimately, this work points toward building AI systems where internal checks are not just about accuracy but also about self-correction at the latent level, which could fundamentally shift how we approach model robustness and trustworthiness.
Tom: Exactly. This paper gives us a solid technique to actively mitigate compounding errors in autoregressive predictions by using a score field derived from real transitions for inference-time repair.
Jane: It’s important to remember that while it’s not perfect, it successfully demonstrated how errors remain spatially sparse, allowing for targeted intervention rather than broad corrections.
Lu: And the future work seems focused on integrating this concept of score-based correction into larger planning architectures to ensure long-term validity in complex scenarios.
Meng: I think the key takeaway for us is that we now have a blueprint for making world models more reliable by adding this specific detection and correction module.
Lalam: We’re excited about how this kind of intrinsic mechanism can influence the development of future generative AI, moving toward systems that are inherently more self-aware in their predictions.
Conclusion: Tom: So, we've been deep in the weeds of MEND, and now we’re getting to wrap things up with what this paper actually means for us on a broader scale.
Jane: I think focusing on the title itself really helps—"MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models." It sounds like it’s tackling a really messy problem in world models without needing all that tedious ground-truth labeling.
Lu: Exactly! The authors are presenting a single conditional score network that does three things at once: it detects the error, pinpoints where the error is, and even suggests how to fix it during inference. That level of integrated functionality is really something special in how we think about model robustness.
Meng: From my side as an engineer, I’m curious about the practical implication of not needing those expensive labels. If we can do this without massive datasets specifically annotated for failures, that opens up a huge path for deploying world models in real-world scenarios where labeling is just impossible.
Lalam: I see it through the lens of culture here. If we can build AI systems that are inherently capable of auditing their own predictions and correcting mistakes internally, it fosters a different kind of trust in these complex generative systems. It moves us closer to more reliable reasoning structures within the AI ecosystem.
Tom: That’s a huge point, Lalam—building intrinsic self-correction. The paper shows that this isn't just fixing a single wrong prediction; it’s about correcting the underlying structure of how the model generates those states.
Jane: And when you look at the results, even though it only corrects errors that are tangent to the data manifold, we still see tangible reductions in latent error during rollouts. It shows that targeted intervention has a real effect on overall model performance.
Lu: That spatial sparsity of the error is what makes this approach so elegant from a theory standpoint—it suggests that even when the whole state is wrong, the deviation can be localized to just a few tokens. I’m already thinking about how this could apply to long-horizon planning tasks where errors compound quickly.
Meng: That localization aspect is crucial for me because it makes it actionable. Being able to isolate the problem patch rather than just getting a generic error signal means we can focus our compute on the fix, which is essential for making this viable in production environments.
Lalam: For culture, I think this work signals a shift from simply training bigger models to training smarter systems that can manage their own internal uncertainty and self-repair. That’s where the long-term value lies for how we design and trust these advanced AI applications.
Tom: So, to wrap up this segment, MEND gives us a way to check our world models on the fly without needing perfect labels, which points toward a future where AI systems are less brittle when they encounter novel situations.
Jane: It’s clear that the work by the authors has provided a solid framework for auditing and repairing latent states in world models.
Lu: This technique opens up fascinating avenues for exploring how we can systematically enforce consistency across complex, multi-step generative processes.
Meng: I’m eager to see how this detection loop scales when we move from simple environments like WALL to the more complex dynamics of real-world physics simulations.
Lalam: We’re going to keep digging into these concepts because this work genuinely contributes a new mechanism for ensuring AI systems are more reliable and transparent in their generative processes.
Tom: Next up, we're going to look at some of the specific experimental results on WALL and POINTMAZE to see exactly how robust this detection is under different conditions.
Ali Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar
University of Melbourne · Monash University
cs.CV, cs.AI
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: Accepted at DICTA 2026 (International Conference on Digital Image Computing: Techniques and Applications). Camera-ready version
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 77/100
The gist: World Models are emerging as a frontier in computer vision, but their robustness remains largely unexplored, leading to the identification of latent hallucination—where predicted next latent states
Key concepts
- Latent Hallucination
- This occurs when a model predicts a next state that is mathematically impossible given the true rules of the environment. It means the predicted state is either unreachable or results from an invalid action, causing errors to compound rapidly in sequential predictions.
- Conditional Score Field
- This is the output of MEND, which acts as a single network providing three functions: detection (measuring distance from valid states), localization (showing how much each part of the image needs to move), and correction (defining the direction to move toward a valid state).
- Denoising Score Matching
- This is the training method used. The network learns by trying to reverse a noise process applied to valid next states. It teaches the model to accurately predict the conditional score of real transitions, allowing it to learn the structure of valid world dynamics without needing any error labels.
- Tweedie’s Formula
- This mathematical formula is used during correction. It uses the learned score field and its variance to calculate a new state estimate that is pulled toward the set of valid successor states, effectively defining the precise direction needed to fix the hallucination.
Terminology
Summary
World Models are emerging as a frontier in computer vision, but their robustness remains largely unexplored, leading to the identification of latent hallucination—where predicted next latent states decode to scenes that never occur—which compounds autoregressively and silently. This paper introduces MEND (Masked Empirical-Bayes Neural Denoising), a single conditional score network that serves as a unified mechanism for detecting, localizing, and correcting this latent hallucination at inference time without requiring ground-truth error labels.
The gist
MEND is a single conditional score network trained by denoising score matching on real transitions, whose score field detects hallucination magnitude, localizes it to specific image patches, and defines an inference-time correction direction.
What is Latent Hallucination?
Latent hallucination occurs when a model produces a next latent state that lies outside the set of valid successor states reachable under the true environmental dynamics, i.e., it is Unreachable state (zˆt+1 ∈/ R(zt))
or an Incorrect action outcome (zˆt+1 ∈ R(zt)∖ R(zt, at))
. Because rollouts are autoregressive, a single hallucinated step reduces the probability of the entire rollout and causes the error to compound with depth. The error is characterized as being spatially sparse: a small subset of tokens carries a disproportionate share of the total error.
How MEND Works
MEND instantiates a general detect–localise–correct loop using a single differentiable object: a conditional score field, denoted as the network output, sθ(˜z zt, σ). This score field serves three complementary roles simultaneously:
-
Detection: The scalar D =∥sθ(ˆzt+1 zt)∥ squared measures how far the prediction is from the set of reachable states.
-
Localisation: The per-token field∥sθ(ˆzt+1 zt)n∥ is a
direct estimate of how far each token must move,
localizing errors to specific image patches. -
Correction: By Tweedie’s formula, the posterior mean of the clean state is zˆt+1 + σ squared sθ, meaning the score itself defines the correction direction.
The score network is trained using denoising score matching (Eq. 4), which learns the conditional score ∇ log qσ(˜z zt) of a Gaussian smoothing of valid next states, without requiring any hallucination or accuracy labels. This training ensures that the network learns the exact conditional score
and that the resulting field is world-model-agnostic.
The Detection, Localisation, and Correction Loop
The loop operates iteratively:
-
Evaluate the detector: Calculate D˜ = (D − µacc)/σacc to normalize against known correct predictions.
-
Localize: Identify the most suspicious tokens using the rank of the displacement vector 'd' (mraw ← rank(d)).
-
Masked Correction: Apply a correction step by moving only the suspicious tokens, while freezing confident ones, following Tweedie’s formula to move toward the valid manifold. This process repeats until a predefined budget is reached or validity is achieved.
Experimental Results and Analysis
MEND was evaluated on two environments: WALL (Markov predictor) and POINTMAZE (three-frame history). Detection performance was measured by AUROC, while localisation accuracy was per-token AUPRC against the ground-truth error map.
(Table I results are summarized below)
Detection is robust, achieving an AUROC of up to 0.80 without using actions on WALL and exceeding a single-Gaussian density baseline in POINTMAZE. Localisation is highly effective, with per-token AUPRC reaching up to 0.87 for POINTMAZE hallucinations, demonstrating that the error remains concentrated in a small subset of tokens
even when the whole state is far off the manifold.
Correction reliably reduces latent error; for example, on WALL, MEND achieved a reduction of-6.4% in latent error over 1 step compared to no correction. However, correction is bounded by the part of the error that lies tangent to it [the data manifold],
meaning only the normal part
of the displacement is recoverable.
Ablation and Future Directions
Ablation studies confirm that using only one factor from the Bayes decomposition (D1: reachability) provides strong detection and localisation, while including D2 (action-consistency) offers little benefit in deterministic settings because "the inverse dynamics residual is largely independent of whether a prediction is valid or hallucinated.
Improvements for AI systems
Here are the specific improvements to AI systems based on the MEND framework described in the paper:
-
Improved Robustness Against Latent Hallucination in World Models: The system will no longer suffer from compounding errors where a single hallucinated state corrupts subsequent predictions.
-
Inference-Time Correction of Errors: The system can actively correct predicted latent states during inference by iteratively nudging suspicious tokens toward the known valid data manifold, leading to reliable reductions in single-step latent error and improved overall prediction accuracy.
-
Label-Free Detection and Localization: The system utilizes a single conditional score field to simultaneously detect if a prediction is hallucinated (global score), localize exactly which image patches are responsible for the error (per-token localization map), and determine the direction for correction.
-
Data-Efficient Error Mitigation: Unlike methods requiring expensive data collection and retraining on failure cases, MEND operates entirely at inference time on a frozen world model, making it practical for deployment in environments where fine-tuning is not feasible or desirable.
-
Spatially Sparse Error Handling: The system leverages the finding that hallucination is spatially sparse (a small subset of tokens carries the error), allowing for highly targeted and efficient correction procedures rather than uniform state updates.
This improved AI system can perform:
-
Accurate, reliable planning and simulation in complex environments (like navigation or robotics) by ensuring that the predicted future states are physically plausible and consistent with the environment's true dynamics, even when the model
imagines
a path it never actually takes. -
Enhanced safety guarantees for autonomous agents by providing an immediate check on the validity of predicted actions or states before execution, preventing errors stemming from latent space distortions.
-
More efficient use of existing world models in real-time applications by providing a mechanism to
clean
potentially faulty outputs without needing external ground-truth error labels.
Abstract
World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.
Sources
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- Mastering Diverse Domains through World Models
- DINOv2: Learning Robust Visual Features without Supervision
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- Do Deep Generative Models Know What They Don't Know?
- Multiscale Score Matching for Out-of-Distribution Detection
- Score-Based Generative Modeling through Stochastic Differential Equations
- Diffusion Posterior Sampling for General Noisy Inverse Problems
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models