LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

arXiv:2608.11691 · cs.LG, cs.CL · Submitted 2026-08-12 · Read on arXiv

Xinhao Zhong, Yuxia Qiao, Junhao Li, Hao Fang, Yi Sun, Bin Chen

Harbin Institute of Technology, Shenzhen · Tsinghua University · Pengcheng Laboratory

cs.LG, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection Abstract Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with

Terminology

Summary

LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

Abstract

Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in its reasoning trace. This leakage is substantially more pronounced in natively RL-trained MLRMs than in their non-reasoning base models, revealing a privacy risk that existing unlearning methods are not designed to address. RL-induced exploration leaves sensitive content with a distinctive token-level entropy signature that is largely absent from base models. Based on this observation, LEMUR is proposed as a fully training-free, inference-time unlearning framework for natively RL-trained multimodal models. LEMUR uses entropy dynamics as a control signal to identify when sensitive reasoning begins and when sanitization should stop. During this interval, it redirects the reasoning trajectory through entropy-modulated visual-anchor latent injection, replacing committed tokens with sanitized, probability-weighted embeddings re-grounded in the input image. Across diverse MLRMs, LEMUR consistently outperforms existing unlearning methods in suppressing both reasoning-trace and answer leakage, while better preserving non-sensitive utility and output fluency.

Introduction

Reinforcement-learning (RL) post-training has reshaped how multimodal models reason. Modern multimodal large reasoning models (MLRMs) such as R1-Onevision and Vision-R1 are trained to explore, emitting a long chain of thought (CoT) inside an explicit ⟨think⟩...⟨/think⟩ region before answering, and this exploratory reasoning drives much of their gain on visual question answering. However, the stronger exploratory ability that drives this reasoning also brings greater privacy risks: such models tend to leak more sensitive information during the reasoning process than their non-reasoning counterparts. This calls for machine unlearning, the removal of a designated subject's information on request without retraining from scratch.

Reasoning models expose a failure mode that answer-centric views miss: a model can omit a private fact from its final answer yet still recite it inside the reasoning trace. The result is a tension in which answer-level cleaning leaves the trace leaking, while perturbing the trace hard enough to stop the leak degrades the model's reasoning ability. Existing methods are mismatched with RL-trained MLRMs in two fundamental ways. First, they lack a reliable mechanism for monitoring the diverse and exploratory reasoning trajectories induced by RL training, so leakage can persist when the model recalls additional sensitive attributes of the same subject or restates the same private fact using semantically equivalent or synonymous expressions. Second, the fine-tuning and activation-steering interventions on which these methods rely can be overly disruptive for such models, substantially degrading the reasoning capabilities that distinguish them.

Studying at the level of individual tokens how a memorized attribute surfaces during decoding, the authors find that its recital in RL-trained MLRMs produces a characteristic two-stage entropy signature largely absent from their non-reasoning base models: the leak begins at a high-entropy decision point where the model hesitates among several candidate values, and once it commits to the attribute the per-token entropy collapses to near-zero and remains low until the attribute span ends, recovering only at its boundary. This exposed structure is attributed to RL exploration, under which the model actively deliberates over and commits to memorized content within the trace rather than emitting it only in the answer.

Guided by this signature, LEMUR is a training-free unlearning framework that operates entirely at decoding time. Given a forget set, LEMUR constructs a forget-relevance space that captures the protected content associated with the target subject, while using token-level entropy as a complementary internal signal of impending recall. When the model emits a forget-relevant token or exhibits an anomalous deviation in its per-token entropy trajectory, LEMUR switches from standard autoregressive decoding to a latent decoding regime. Instead of re-injecting the committed one-hot token, it redirects the latent trajectory toward a sanitized visual anchor, with the injection strength dynamically modulated by the current entropy. LEMUR determines the exit point using an adaptive entropy threshold estimated from the statistics of the current trajectory. To prevent repeated oscillations between discrete and latent decoding, a refractory cooldown window is introduced and the number of allowed transitions is capped.

The contributions are summarized as follows:

  • RL-trained MLRMs suffer markedly more severe reasoning-trace leakage than their base MLLMs, and this leakage carries a distinctive token-level entropy signature.

  • Building on the unique entropy signature of leakage, LEMUR is a training-free framework that unlearns by steering the decoding process, switching into a latent decoding state and injecting an entropy-controlled visual anchor to redirect the reasoning trace.

  • Across a range of reasoning MLRMs, LEMUR achieves state-of-the-art unlearning, outperforming both training-based and training-free baselines in leakage, utility, and fluency.

Related Work

Reasoning Multimodal Large Language Models: Large reasoning models scale test-time computation to emit explicit chains of thought, learning through reinforcement learning with verifiable rewards to deliberate before committing to an answer. This paradigm has moved to the multimodal setting, where MLRMs such as R1-Onevision and Vision-R1 couple visual perception with linguistic reasoning and markedly improve visual question answering. Unlike instruct multimodal large language models (MLLMs), whose rationales are prompted or distilled, their reasoning is acquired natively through RL. The explicit trace, however, is a double-edged sword: while it improves reasoning, it can also surface memorized private content even when the final answer is clean, a behavior best diagnosed at the level of individual decoding steps.

Machine Unlearning: Machine unlearning removes the influence of designated data from a trained model without retraining it from scratch. It was first developed for text-only LLMs, through gradient ascent and its stabilized variants, preference optimization, and Bayesian or continual formulations, as well as training-free inference-time interventions that act on prompts or output logits. These ideas were then carried to MLLMs through single-image, modality-aware, and other multimodal-specific schemes that act on the final answer. Reasoning models raise a sharper challenge, since answer-only forgetting leaves the fact recoverable from the trace whereas aggressive trace perturbation collapses reasoning into degenerate repetition. Closest to this work, Li et al. formalize this reasoning-preserving setting and propose R-MUSE, a training-free activation-steering method, yet like the other approaches above it targets instruct MLLMs and acts on hidden activations rather than on the decoding process itself. LEMUR is therefore the first unlearning framework designed for natively RL-trained MLRMs and the first to intervene at decoding time, using the entropy dynamics of the reasoning trace to redirect generation through entropy-controlled latent injection of a visual anchor.

Method

Problem Definition: Machine unlearning (MU) for multimodal large reasoning models (MLRMs) aims to remove targeted forgetting knowledge while minimizing the degradation of general capabilities. Let Mθ denote the original MLRM with parameters θ. The forget set Df = (Ij, Tj) contains the subjects to be forgotten, and Dn = (Ik, Tk) is the remaining normal data outside the forget set, on which the model's general capability must be preserved; each I is an image and each T = (q, a) is a question–answer pair for visual understanding.

Unlike a non-reasoning MLLM that maps a query directly to an answer, an MLRM forms its input by concatenating the N vision tokens xv produced by the vision encoder with the M text tokens xt, i.e. x = xv ⊕ xt, and decodes autoregressively. At step t it predicts the next-token distribution pt = Mθ(· x, y<t) ∈ ΔV−1, yt ∼ pt, where y<t = (y1,..., yt−1) are the previously generated tokens and V is the vocabulary. Decoding proceeds in two stages: the model first emits a reasoning trajectory r1:m = (r1,..., rm) enclosed in ⟨think⟩...⟨/think⟩, and then the final answer a1:n = (a1,..., an), so the complete response is y = (r1:m, a1:n). The explicit trace r1:m is precisely the additional channel through which a forgotten attribute can resurface, even when it is absent from a1:n.

Unlearning objective: Let Acc(Mθ′(I, q), a) ∈ [0, 1] measure whether the model's response to (I, q) is consistent with the gold answer a, and write the expected correctness of Mθ′ over a set D as A(θ′; D) = E(I,q,a)∼D[Acc(Mθ′(I, q), a)]. Unlearning seeks an unlearned model Mθ̂ that drives this correctness to its minimum on the forget set while leaving it unchanged on the remaining normal data: θ̂ ∈ arg minθ′ A(θ′; Df) s.t. A(θ′; Dn) ≈ A(θ; Dn). For reasoning models, this requirement is subject-level rather than tied to the queried answer a alone: across the full response y = (r1:m, a1:n), neither the reasoning trace nor the final answer should reveal any of the subject's private attributes, including those that answer other possible questions about the same subject. Unlike conventional methods that fine-tune θ, LEMUR keeps θ̂ = θ and achieves unlearning purely at inference by intervening in the decoding distribution.

Entropy-augmented Sensitivity Switching: A reasoning model produces its trace autoregressively from the distribution pt, and its step-wise uncertainty is read through the token-level entropy Ht(v) = −Σv∈V pt(v) log pt(v), which is large when several candidates compete and small when one token dominates. Recalling a memorized sensitive attribute leaves a characteristic two-stage trace in this signal: in an initial deliberation stage, the model briefly weighs the candidate values of the attribute, so Ht(v) rises sharply while no single value yet dominates; once it commits, the memorized span is recited almost deterministically and Ht(v) collapses to near-zero, recovering only at the span boundary. This rise-then-collapse pattern delimits exactly the segment over which an unlearning intervention must remain active.

Following prior unlearning work, the sensitive content is first identified through an explicit lexical cue. For the queried subject s a forbidden token set Φs ⊂ V is maintained covering its protected attributes, and the probability mass PtΦ = Σv∈Φs pt(v) is tracked, flagging a step when its forbidden mass exceeds a threshold (PtΦ ≥ ρ). This threshold is reached only after the model has concentrated enough probability on the forbidden tokens to recite the value at low entropy, so it fails to react during the initial deliberation stage, when the model first begins to be drawn toward the protected value and the probability is still spread across the competing variants and the synonyms, so that no individual token reaches ρ. The entropy Ht(v) is therefore used as an additional cue that assists the lexical test rather than replacing it. This deliberation stage is marked by high Ht(v), since the protected value and its variants compete without a dominant winner, so the mass bar is lowered to ρlo < ρ whenever the step is genuinely uncertain (Ht(v) ≥ τ), letting a diffuse aggregate of synonymous candidates suffice to trigger:

gt = (PtΦ ≥ ρ) ∨ (Ht(v) ≥ τ ∧ PtΦ ≥ ρlo)

The two terms mirror the two stages of the phenomenon: the lexical term fires on the low-entropy committed span, whereas the entropy-augmented term recovers the high-entropy deliberation stage and its diffuse synonym mass without firing on contentless high-entropy tokens such as discourse connectives, which carry no forbidden mass (PtΦ < ρlo). Decoding accordingly runs in one of two modes mt ∈ D, S, ordinary discrete generation D and a sensitive mode S, and switches from D into S as soon as gt fires, so that the intervention spans the sensitive segment from its uncertain beginning to its deterministic completion.

Entropy-aware Visual Anchor Injection: In sensitive mode S (triggered by gt), LEMUR no longer feeds the sampled token back to the model. Instead it feeds back a continuous embedding, built from two parts: (i) a constrained latent feedback that removes the forbidden mass yet stays differentiable, and (ii) an entropy-controlled injection of a visual anchor and a safe-answer anchor that steers the output away from the memorized span.

Constrained Latent Feedback: In mode S the step distribution is first restricted by removing the forbidden tokens and renormalizing over the survivors, p̃t(v) = pt(v) ⊮[v ∉ Φs] / Σu∉Φs pt(u), which guarantees that no sensitive value can be emitted while preserving the relative ordering of all admissible candidates. Rather than committing to a single discrete token, the restricted distribution is summarized as its expected embedding, êt = Σv∈V p̃t(v) Ē[v], and êt is fed back as the next input embedding. Because this retains the full competition among admissible continuations rather than collapsing it onto one token, the model carries forward a soft, gradient-preserving state in which the forgotten attribute has already been suppressed, leaving room for the anchors to take effect before the trajectory recommits.

Visual Redirection anchors: Suppressing the forbidden mass alone leaves the latent state under-determined and prone to drift back toward the memorized value once it re-enters discrete decoding. Two fixed anchors are therefore injected that supply an explicit, non-sensitive target for the redirection. The visual anchor evis is the averaged embedding of the pretrained visual special tokens Vvis (e.g.,,,), which regrounds the reasoning in the visual modality, while the safe-answer anchor esafe averages the embeddings of a few fixed refusal and uncertainty templates, such as I'm not sure and I cannot identify this person from the image, which pulls the continuation toward a benign answer. They are combined into a single composite anchor a = β evis + (1 − β) esafe, where β ∈ [0, 1] balances grounding the response in the image against deflecting it toward an explicit safe phrasing.

Entropy-controlled Injection: The composite anchor is injected into the latent feedback by convex interpolation, et = (1 − γt) êt + γt a, where the injection strength γt is set by the step entropy Ht(v). The strength is entropy-dependent because the two stages of a memorized span call for different amounts of steering. When entropy is high, the attribute value is still undecided, so a strong anchor can cheaply steer the outcome; when entropy is low, the model is already reciting fluent text that masking has made safe, so a strong anchor would only distort it. The strength is scaled in proportion to Ht(v), using γ as the strength at the reference entropy τ: γt = min(γmax, (Ht(v)/τ)γ), capped at γmax so the feedback stays on the embedding manifold. High-entropy steps thus get stronger steering and low-entropy steps weaker steering, with the strength equal to γ when Ht(v) = τ. Applying these operations at every step of the sensitive segment pushes each distribution toward image-consistent, non-sensitive content, so that when decoding returns to mode D the model produces a benign answer instead of the forgotten attribute, with no update to θ.

Dynamic Entropy-controlled Phase Duration: The switch decides only when a latent phase begins, and its effectiveness depends equally on how long that phase is held: the intervention must stay active across the entire memorized span yet release as soon as the model resumes ordinary generation, because releasing too early reopens the low-entropy committed recital to a forbidden completion while holding the latent channel past the span suppresses benign tokens and degrades fluency. The entropy trajectory supplies the boundary signal directly, since entropy collapses inside the span and recovers at its end, so a phase should persist while Ht(v) stays low and exit once it climbs back. The level to which entropy recovers is subject-dependent, and a subject whose discrete-mode generation is itself low-entropy never crosses a fixed global threshold and leaves the phase to run unchecked, so the release threshold is made adaptive rather than constant.

The model's baseline uncertainty is tracked with an exponential moving average of the entropy over the discrete-mode steps, updated only while mt = D: H̄t = (1 − η) H̄t−1 + η Ht(v), so that entropy can be judged relative to what the model exhibits on ordinary text for the same subject. A phase opened at step t0 then fixes its duration through the exit indicator zt = (¬gt ∧ Ht(v) ≥ κ H̄t) ∨ (t−t0 ≥ Wmax), which terminates the latent encoding once the forbidden mass has cleared (¬gt) and entropy has recovered above the adaptive threshold κ H̄t, and caps the total length at Wmax as a hard safeguard against runaway phases. Decoding reverts to mode D at the first step where zt holds, so the dynamic threshold κ H̄t is precisely what governs the length of each latent-encoded segment, lengthening the phase on subjects that deliberate at high entropy and shortening it on subjects that recite with little uncertainty, in both cases matching the intervention window to the extent of the memorized span without a manually tuned constant.

To preserve utility, back-to-back latent phases can still degrade the fluency of the generated text: if the gate re-fires the instant a phase ends, the latent intervention chains into degenerate repetition and disrupts the surrounding discrete generation. A short cooldown is therefore imposed: after each exit, at least C discrete steps must elapse before a new phase may open, so the gate is suppressed whenever fewer than C steps have passed since the last release. This leaves the latent intervention free to act on genuinely distinct sensitive spans while keeping it from latching onto the fluent text that immediately follows a suppressed one.

Experiments

Experimental Setup: All experiments are conducted on a dataset reconstructed on top of MLLMU-Bench. Its corpus of fictitious subjects, each paired with a portrait image and curated private QA pairs, is reused and partitioned into forget, retain, and celebrity splits. The original question–answer pairs contain no reasoning trace, so a strong multimodal teacher (Qwen3.5-35B-A3B) is used to distill a ⟨think⟩/⟨answer⟩ chain for every pair. The teacher sees the image and the subject's other attributes and writes a first-person reasoning trace that leads to the answer. Evaluation is performed at different forget ratios over three task types: classification, fill-in-blank, and generation. The models to be unlearned are natively RL-trained MLRMs: R1-Onevision-7B and Vision-R1-7B as primary backbones. As LEMUR is training-free and leaves the weights untouched, it is applied to the vanilla checkpoint, and compared against training-based GA, NPO, MMUnlearner, and R2 MU and the training-free state-of-the-art R-MUSE.

Metrics: Task Accuracy is the mean of the classification and fill-in-blank accuracies on a split, which should be low on forget yet high on retain and celebrity. On the generation task, Target Recall (TR) measures the fraction of the subject's queried attributes that appear in the output. Subject-level Reasoning Leakage (SRL) measures whether the reasoning trace reveals any of the queried subject's curated attributes, excluding values already given in the prompt, reported as the average over the three tasks. Gemini-2.5-Pro is used as an automatic judge of the Reasoning Retention Ability (RRA) by evaluating the fluency and naturalness of the text generated across all tasks.

Main Results: Table 1 reports all five metrics across the forget, retain, and celebrity splits, and LEMUR achieves the strongest forgetting on the forget split by pushing classification accuracy, fill-in-blank accuracy, and generation target recall below every baseline. This advantage becomes most meaningful at the reasoning level, where the baselines behave very differently. Answer-oriented methods such as MMUnlearner suppress the final answer while leaving subject-level reasoning leakage almost untouched because they never intervene on the reasoning trace, and even the reasoning-aware baselines reduce leakage only partially. LEMUR drives reasoning leakage far below all of them, showing that it erases the target concept from the intermediate reasoning as well as from the final answer, achieving genuine reasoning-process forgetting rather than answer-only suppression.

This aggressive forgetting does not come at the usual cost to utility. On the retain and celebrity splits LEMUR keeps its classification, fill-in-blank, and generation scores at the vanilla level, whereas the training-based baselines lose visible ground as their parameter updates spill over from the forget set onto retained knowledge. The Reasoning Retention Ability (RRA) makes this contrast even clearer, since the gradient-based baselines depress RRA even on non-forget data while LEMUR keeps it close to the vanilla level on every split including forget, so the model continues to generate fluent and well-formed reasoning even where the queried facts have been removed instead of collapsing into the repetitive or degenerate text that stronger interventions tend to produce. These trends are reproduced consistently across both RMLLM backbones and all forget ratios.

Component Ablation: The inference-time components are ablated on the Onevision-R1-7B 5% forget setting. Vanilla is the original model with no unlearning, and Base is the most basic unlearning intervention (lexical forbidden-token masking); then LEMUR's components are added cumulatively on top of Base: entropy-augmented sensitivity switching (ESS), visual anchor injection (VAI) at fixed strength and its entropy-aware variant (EVAI), and dynamic entropy-controlled phase duration (DEPD). Base relies solely on the lexical detection to flag and rewrite the sensitive tokens, and this most basic form of intervention already reduces forget-set accuracy, recall, and leakage over Vanilla, but the reduction is weak and its crude masking simultaneously erodes retain utility together with the model's generation and reasoning quality. Complementing this lexical test with the entropy cue in ESS recovers the uncertain deliberation stage that the mass threshold alone would miss, and the more precise detection of the memorized span markedly strengthens the forgetting effect over Base. Injecting the visual anchor in VAI supplies the suppressed latent state with an explicit image-grounded target, which redirects the trajectory toward safe, visually consistent content and improves forgetting further from the visual side. Because a fixed injection strength tends to over-steer once the span is already committed, EVAI makes the strength entropy-adaptive so that the anchor acts strongly only where the span is still uncertain, and this both sharpens the forgetting and begins to recover the retain utility that the more aggressive injection had depressed. Finally, DEPD applies the dynamic entropy-controlled constraint to keep the intervention aligned with the extent of the memorized span rather than a fixed window, and by timing the latent phase to the span it substantially raises the model's utility while sacrificing almost none of the forgetting ability.

Transferable Ability: LEMUR's complete inference-time pipeline is transferred to Qwen2.5-VL with an identical cumulative component analysis on the forget split. The method remains broadly effective under this new architecture, albeit with notable shifts in component-wise contributions. Unlike the RL-trained backbone, Qwen2.5-VL does not produce the pronounced entropy surges typically associated with transitions into memorized content regions. As a result, when entropy cues are coupled with the lexical test in the ESS, they function primarily as an additional gating condition rather than as a genuinely informative detector, yielding only modest improvements over the Base configuration. The situation changes substantially once the visual pathway is engaged. Both the injection of the visual anchor and its entropy-adaptive strength modulation contribute significantly to improving forgetting metrics. In this setting, the visual anchor provides the dominant corrective signal, while the entropy cue continues to serve as a supplementary outcome consistent with the multimodal nature of MLLMs, wherein visual evidence plays a central role in guiding generation. Given the attenuated entropy signal, the adaptive exit mechanism tends to prolong the phase spans, which in turn helps recover reasoning retention ability. When benchmarked against R-MUSE, a baseline method specifically designed for Qwen2.5-VL, the transferred LEMUR pipeline still achieves superior overall forgetting performance, demonstrating that the effectiveness of LEMUR is not contingent upon an RL-trained foundation, underscoring its generality and transferability across different MLLM backbones.

Additional experiments: Generalization beyond privacy data is verified on a general visual-reasoning corpus (VQAv2) using the identical pipeline, showing that the entropy signature LEMUR exploits is not confined to privacy-oriented data and LEMUR remains effective in this general-domain setting. Robustness to a higher forget ratio of 15% on both RMLLM backbones is evaluated, with the overall picture matching the 5% and 10% settings closely. Transfer to a third RL-trained MLRM, OpenVLThinker-7B, is verified across all three forget ratios (5%, 10%, and 15%), with the qualitative pattern unchanged: LEMUR delivers the strongest forgetting on the forget split while preserving retain and celebrity utility and keeping the Reasoning Retention Ability at the vanilla level.

Qualitative Analysis: Per-instance outputs of the unlearned R1-Onevision-7B model on the 5% forget split are inspected across all three MLLMU-Bench tasks. A consistent qualitative pattern emerges: recognition is preserved, private recall is corrupted. The unlearned model still sees the subject correctly—its ⟨think⟩ traces open with faithful visual descriptions—so LEMUR does not degrade generic perception. What breaks is the recall step: the moment the chain reaches a private attribute, it substitutes a plausible but incorrect value drawn from the model's prior rather than the memorized ground truth. The three tasks fail in mutually consistent ways: on classification the model confidently selects a wrong option while narrating a fabricated justification; on fill-in-the-blank the blank is completed with the same hallucinated attribute; on open-ended generation the free-form answer commits to the wrong attribute in prose. Failure modes are otherwise benign: a minority of forget-split generations exhibit mild degeneration (token repetition or truncated spans) once the recall pathway is suppressed, but these do not leak the protected attribute and are confined to the forget subjects; retain- and celebrity-split outputs remain fluent and accurate.

Conclusion

A privacy risk specific to natively MLRMs was identified: a memorized sensitive fact can still be reproduced in the reasoning trace even after it has been removed from the final answer. This behavior was traced to a distinctive token-level entropy signature induced by RL training and largely absent from non-reasoning base models. Building on this observation, LEMUR was proposed as a training-free, inference-time unlearning framework that turns entropy dynamics into a control signal for decoding-time intervention. LEMUR monitors forget-relevant content and anomalous entropy dynamics, transitions from discrete autoregressive decoding to a latent decoding regime, and redirects the reasoning trajectory through entropy-modulated injection of a sanitized visual anchor. Across a range of RL-trained MLRMs, LEMUR substantially reduces both answer-level and reasoning-trace leakage while preserving non-sensitive utility and output fluency, consistently outperforming existing training-based and training-free baselines. More broadly, the results suggest that the decoding dynamics of reasoning models provide a promising foundation for training-free privacy control. In future work, this entropy-driven decoding-time perspective will be extended to broader safety objectives.

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Entropy-aware privacy guard for reasoning traces: Implement a real-time token-level entropy monitor that detects when a model begins deliberating over memorized sensitive content (high-entropy hesitation followed by low-entropy commitment). This allows AI systems to automatically trigger sanitization before private facts are recited in chain-of-thought, not just in final answers.

  2. Training-free, inference-time unlearning: Replace costly fine-tuning or gradient-based unlearning with a decoding-time intervention that requires no weight updates. The improved system can remove specific subjects' information on demand while preserving general capabilities, making it practical for dynamic privacy requests in deployed models.

  3. Visual-anchor redirection for multimodal reasoning: When sensitive content is detected, the system injects entropy-modulated visual anchor embeddings (derived from image special tokens) to re-ground the reasoning trajectory in the input image. This enables the model to continue producing fluent, image-consistent reasoning without leaking the protected attribute, effectively re-routing thoughts toward safe, visually supported conclusions.

  4. Adaptive intervention duration via entropy dynamics: Use an exponential moving average of entropy during normal decoding to set a subject-specific threshold for when to exit the sanitized latent phase. This prevents both premature release (which reopens leakage) and over-long suppression (which degrades fluency), automatically adapting to each user's or subject's baseline uncertainty.

  5. Two-stage leakage detection combining lexical and entropy cues: Combine forbidden-token probability mass with entropy thresholds to catch both the committed recital stage (low entropy, high token mass) and the deliberation stage (high entropy, diffuse synonym mass). This makes the system robust to paraphrasing and synonymous expressions of sensitive facts, which purely lexical filters miss.

  6. Refractory cooldown to prevent degenerate repetition: After each sanitization phase, enforce a minimum number of discrete decoding steps before allowing another intervention. This prevents the system from oscillating between latent and discrete modes, which would otherwise produce repetitive or incoherent output.

  7. Subject-level unlearning beyond answer-level: Ensure that no private attribute of a forget subject appears anywhere in the full response—including reasoning traces, fill-in-blank completions, and open-ended generation—not just the queried answer. This makes the system compliant with broader right to be forgotten requirements that cover all possible questions about a subject.

  8. Transferable privacy control across model architectures: The entropy-visual-anchor pipeline works on both RL-trained reasoning models (where entropy signals are strong) and instruction-tuned models (where visual anchors dominate). This allows a single unlearning framework to be deployed across heterogeneous model families without retraining or architecture-specific tuning.

What the improved AI system can do:

  • Accept a user request to forget a specific person or subject and immediately suppress all private attributes from both reasoning and answers, without retraining.

  • Maintain fluent, natural reasoning and high accuracy on non-sensitive tasks even after aggressive forgetting.

  • Detect and neutralize leakage attempts that use synonyms, paraphrases, or indirect references to protected facts.

  • Operate in real-time during inference, making it suitable for interactive applications where privacy requests arrive dynamically.

  • Preserve visual perception and general reasoning quality, as demonstrated by unchanged performance on retain and celebrity splits.

  • Automatically calibrate intervention strength and duration per subject, avoiding both under-suppression and over-suppression.

Abstract

Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in its reasoning trace. This leakage is substantially more pronounced in natively RL-trained MLRMs than in their non-reasoning base models, revealing a privacy risk that existing unlearning methods are not designed to address. We show that RL-induced exploration leaves sensitive content with a distinctive token-level entropy signature that is largely absent from base models. Based on this observation, we propose LEMUR, a fully training-free, inference-time unlearning framework for natively RL-trained multimodal models. LEMUR uses entropy dynamics as a control signal to identify when sensitive reasoning begins and when sanitization should stop. During this interval, it redirects the reasoning trajectory through entropy-modulated visual-anchor latent injection, replacing committed tokens with sanitized, probability-weighted embeddings re-grounded in the input image. Across diverse MLRMs, LEMUR consistently outperforms existing unlearning met hods in suppressing both reasoning-trace and answer leakage, while better preserving non-sensitive utility and output fluency. These results demonstrate that RL-induced entropy dynamics provide a distinctive signal for privacy leakage and that exploiting this signal enables effective training-free unlearning for reasoning-capable multimodal models.

Sources

Related papers