ModalFidelity: Routing Modalities for Deepfake Detection on a Budget

arXiv:2609.38246 · cs.CR, cs.AI · Submitted 2026-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "ModalFidelity: Routing Modalities for Deepfake Detection on a Budget".

Elias: Deepfakes are evolving to hide manipulations within small, semantically crucial fractions of media,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So Elias and I just got through this paper called "ModalFidelity: Routing Modalities for Deepfake Detection on a Budget," which addresses how detectors waste compute when dealing with short deepfake manipulations. The core idea seems to be that instead of running every window of audio and image streams, the system should intelligently decide which stream is worth analyzing based on a hard compute budget. It claims that this approach allows them to stay within a fraction of the original cost while still finding the forgeries, which is a big deal for practical deployment.

Elias: I agree with Nadia; it sounds like they are tackling the issue where detectors consume all their resources even though only tiny, semantically important fractions of media are actually manipulated. The thesis seems to be that deciding where to look is much cheaper than actually looking at everything, especially when evidence is sparse in time and asymmetric across different modalities. This paper, "ModalFidelity: Routing Modalities for Deepfake Detection on a Budget," proposes a lightweight router that previews each window and decides which stream(s) are worth analyzing before any heavy forensic detector runs.

Priya: From a privacy and measurement standpoint, I'm interested in what the data actually shows about this selection process. Since they mention that evidence is sparse in time, I wonder how much actual information the preview encoder can really extract from just a small fraction of frames and a spectrogram before it becomes too compressed or loses critical forensic signals.

Nadia: That’s a fair point, Priya; the abstract mentions that the preview encoder reads a "small fraction of frames and a spectrogram," which is supposed to cost only a small fraction of the full detector's cost. The paper argues that this preview feature, combined with an LSTM cell that carries what has been seen so far, gives them enough context to make those hard decisions about acquiring audio or image data at each step.

Elias: And that decision-making process is governed by a budget enforcement mechanism where any action exceeding the remaining budget b t gets its logit set to negative infinity, which means the policy can only choose affordable actions. This strict constraint ensures they never exceed their predetermined compute quota, which aligns with their goal of operating under an explicit compute budget.

Paper summary: Priya: If the system is constrained by a hard budget and it has to make decisions window by window without lookahead, how robust is this approach when the forgery itself is very subtle or if it spans across a boundary between two windows? Does this sequential decision-making process inherently limit its ability to capture long-range temporal dependencies in the evidence?

Nadia: That’s exactly where I think the LSTM cell comes into play; it carries what the stream has shown so far, which should help mitigate some of that immediate lookahead issue by providing a history of modality information. The training involves a two-stage process using DAgger distillation, where Stage one solves the problem as a multiple-choice knapsack problem exactly over the window and remaining budget, and Stage two trains the policy to recover from its own errors by imitating what it can't actually see.

Elias: The training method sounds clever because it uses a teacher that is "clairvoyant," meaning it knows the entire stream, including windows the student hasn't reached yet, which forces the student policy to learn foresight even though its actual mechanism doesn't have that capability. This setup seems designed to make the resulting router policy very effective at maximizing reward under those strict constraints.

Priya: Considering they assume detectors are already pretrained and frozen black boxes, what are the limitations here? If we plug in a new detector later that has a fundamentally different way of scoring, does this entire routing mechanism still hold up, or is it too tightly coupled to the specific architecture of the forensic detectors psi m they're using?

Nadia: The paper states that they assume any scoring detector can be plugged in because they treat them as black boxes and charge one unit per detector call, meaning c m=one so the budget B counts calls. They are focused on the router's ability to select the right inputs cheaply, not on changing the underlying forensic model itself. This focuses their effort squarely on optimizing the selection process under a fixed cost structure.

Elias: That cost structure is crucial; they charge one unit per detector call, and a window read in both modalities costs two units, which directly feeds into their budget enforcement mechanism B. Maximizing that reward function R credits acquiring manipulated streams while charging for authentic ones shows they are optimizing for precision within the imposed cost.

Priya: So if we translate this back to real-world impact, what does it mean for the security of media? If a deepfake is hidden in just one audio clip out of a thousand, can this router reliably isolate that specific clip and avoid wasting computation on the other nine hundred, which is where traditional methods fail so badly?

Paper summary: Nadia: That’s the central implication: if we can reduce the number of windows we actually run by a fifth when dealing with AV-Deepfake1M data, that translates directly into faster analysis for security researchers and lower operational costs for verification systems. It shifts the burden from massive brute-force checking to intelligent triage.

Elias: The findings suggest that this router can maintain accuracy levels over ninety-six percent of what an oracle knows about where every forgery lies, while using fifteen point nine times less compute than gating after the detectors run, which is a significant efficiency gain for any real-world detection pipeline.

Priya: I'm curious about the future work mentioned; they talk about adaptive computation and temporal aspects of deepfakes. Does this router handle scenarios where the manipulation isn't localized to a single window but spans across several seconds in a complex way? Or is it strictly limited to detecting edits within discrete, manageable time windows?

Nadia: The paper does note that they observed evidence is sparse in time, and their key insight is based on the utility of information being dynamic along both axes of a multimodal stream. They are trying to prove this works for short, localized manipulations that turn meaning at specific points in the video or audio.

Elias: The limitations they admit are related to the assumptions they make; specifically, they note that their formulation is constrained by c m=one for detector calls and a fixed budget B, which means it’s designed for verification systems where costs are clearly defined beforehand. They also state that the router makes a single left-to-right pass with no lookahead in its decision loop, which limits its ability to anticipate future needs beyond what the LSTM carries forward.

Priya: So to wrap up, this paper "ModalFidelity: Routing Modalities for Deepfake Detection on a Budget" introduces a mechanism that uses a lightweight router to selectively acquire streams based on an explicit compute budget before running forensic detectors. It claims this is more accurate than gating after the detectors run while using much less compute.

Nadia: Exactly, and the implication is that we can move toward verification systems that are both more accurate and far cheaper to run by not reading every single window of every stream. This points toward a future where resource management becomes an integral part of the detection process itself rather than just a post-processing step.

Conclusion: Nadia: So, to wrap up this discussion, we've seen how ModalFidelity uses a router to pick which parts of deepfake media are worth checking based on how much compute you have available and what you're trying to detect.

Elias: Exactly, and the authors’ approach is really about creating a system that doesn't waste cycles by looking at everything when the evidence might be hiding in just a small window or stream.

Priya: I think it's interesting how they frame this as managing resources under explicit constraints rather than just trying to build one bigger detector that does everything.

Nadia: Right, and the title itself, "ModalFidelity," really sums up the idea that we can maintain high fidelity in our analysis while being smart about where we spend our processing power.

Elias: And looking at who wrote this, I'm curious if their background in cryptography might influence how they handled those budget constraints and reward functions.

Priya: From my side, I’m thinking about the data itself; it’s fascinating to see how much more accurate this selective approach is compared to just running a standard detector across all windows.

Nadia: That accuracy difference is what makes it so compelling; we're talking about finding those tricky, short manipulations that get missed when you just treat everything equally.

Elias: And the implications for security are big because if we can make detection systems much more efficient, they become deployable in ways they currently aren't.

Priya: The real impact might be in how we handle privacy; by only analyzing certain streams based on a budget, you could potentially reduce the amount of raw data needing intensive forensic scrutiny.

Nadia: It really feels like this moves detection from brute-force checking to intelligent triage, which is where it has a lot of potential to make real-world security tools more practical.

Elias: So it’s not just about better accuracy; it’s about making the entire verification pipeline way more efficient and tailored to specific computational limits.

Priya: That efficiency could mean we can use these detection methods on much larger datasets than we could before, which is a significant step for training robust models.

Nadia: It definitely suggests that future detection systems won't just be about building bigger detectors, but about building smarter decision-making layers on top of them.

Elias: And I wonder what the next cryptographic challenge will be in ensuring that these budget constraints don't introduce new vulnerabilities into the selection process itself.

Priya: That’s a great point to follow up on; we should probably look at those specific assumptions they made about their cost model next.

Oguzhan Baser, Kaan Kale, Sriram Vishwanath, Sandeep Chinchali

University of Texas at Austin · Georgia Institute of Technology

cs.CR, cs.AI

Submitted: 2026-09-29

Updated: 2026-09-29

Comments: 5 pages, 4 figures, 1 table. Submitted to ICASSP 2027

Code: https://github.com/UTAustin-SwarmLab/modal-fidelity

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 83/100

The gist: Deepfakes are evolving to hide manipulations within small, semantically crucial fractions of media, forcing detectors into an inefficient position where they must process every window when only a few

Key concepts

ModalFidelity
A lightweight router that operates window by window. It uses a preview encoder and an LSTM cell to decide which stream(s) are worth analyzing based on a fixed compute budget before any forensic detector runs.
Budget Enforcement
The system strictly enforces a hard compute budget by masking any action whose cost exceeds the remaining budget. This ensures the router only selects affordable actions, guaranteeing that the total cost never exceeds the predefined limit for every stream.
Stage-Based Distillation
A two-stage training method using DAgger distillation. Stage 1 solves a knapsack problem exactly to create an optimal plan. Stage 2 trains the policy by having it mimic a 'clairvoyant teacher' that knows the entire stream, forcing the student to learn foresight.

Terminology

Summary

Deepfakes are evolving to hide manipulations within small, semantically crucial fractions of media, forcing detectors into an inefficient position where they must process every window when only a few seconds matter. This work introduces ModalFidelity, a lightweight router designed to solve this problem by deciding which stream(s) are worth analyzing based on an explicit compute budget before forensic detectors run.

The gist: ModalFidelity is a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading under a hard compute budget it can never exceed.

Problem Formulation and Motivation

The central problem addressed is detecting and localizing forgeries that are often short and semantically chosen, meaning the evidence is sparse in time. Existing detectors fall into two families: single-modality classifiers that are blind to manipulated channels, or early- and late-fusion multimodal detectors that consume every stream on every window. Both approaches fail to provide localization in time and modality under a strict compute constraint. The authors observe that deciding where to look is far cheaper than looking. Therefore, the goal is to design a detector that jointly provides (i) selective acquisition under an explicit compute budget, (ii) localization in time and modality, and (iii) a cost the verifier fixes in advance.

ModalFidelity Architecture

ModalFidelity functions as a router that operates window by window. It utilizes two main components: a preview encoder and an LSTM cell. The preview encoder reads a small fraction of the windows seen so far, together with the remaining budget bt and horizon T−t, returning a compact feature at a small fraction of a detector’s cost. This preview is inherited from an unbudgeted per-window gate and is kept frozen to fix what the router can see. The LSTM cell carries what the stream has shown so far. At every step, the router receives this new preview feature along with the remaining budget and horizon, scoring four actions: none, audio, image, or both.

Budget Enforcement and Reward Maximization

The system enforces a hard compute budget B by construction rather than through penalties. Before each decision, every action whose cost exceeds bt is masked by setting its logit to −∞. This ensures that the policy chooses among affordable actions alone, and the quota holds for every stream and every budget. The router seeks parameters θφ that maximize a reward R defined in Equation (2), which credits acquiring a manipulated stream, charges for authentic ones, and is neutral when nothing is acquired:

max θφ E hX t,m R 1[m ∈ St], y m t s.t. X T t=1 X m∈St c m ≤ B.

Training via Stage-Based Distillation

The training process employs a two-stage approach utilizing DAgger distillation. Stage 1 frames the problem as a multiple-choice knapsack problem solved exactly by dynamic programming over (window, remaining budget), yielding an optimal plan for every state. Stage 2 optimizes Equation (2) directly. The student policy learns to recover from its own errors rather than imitate a path it would never reproduce, as the teacher supplies targets based on information the student cannot have: The teacher is clairvoyant. It reads y m t for the whole stream, including windows the student has not reached. This process results in a policy that imitates a foresight it lacks.

Experimental Findings

The experiments were conducted on AV-Deepfake1M using W2V2-AASIST for audio and GenD on a CLIP ViT-L backbone for images. Results demonstrate the effectiveness of the router:

  1. Stream Asymmetry: The audio detector fires on 70% of audio edits but on only 22% of image ones, no more often than its 21% rate on authentic windows. This confirms that which stream to read must be decided window by window.

  2. Budget Utility: ModalFidelity stays within 2.6 pp of the clairvoyant oracle at every budget and achieves superior accuracy compared to uniform or random spenders, which stay near chance (0.51-0.60).

  3. Strategy Comparison: At a budget fraction ρ=0.20, the router spends 6.27 units on average, achieving 86% of the allowed budget while reaching 0.771 per-video accuracy against blind allocators' 0.577 and their single-detector baselines' 0.526 accuracy.

  4. Gate Placement: The router outperforms late MoE gates by showing that the input, our router needs 258 GFLOPs at ρ=0.20, 15.9× fewer than the late gate, while maintaining higher accuracy (0.767 against at most 0.732).

Improvements for AI systems

Here are the improvements that can be made to existing deepfake detection systems by implementing the ModalFidelity router:

  1. The AI system will transition from a blind, uniform analysis strategy to an adaptive, budget-constrained one. Instead of running every forensic detector on every video window (which wastes compute on authentic content), the system will first use the lightweight ModalFidelity router to intelligently select which modalities (audio, image, or both) are worth analyzing for any given time window.

  2. The improved system can achieve a significant reduction in computational cost while maintaining high detection accuracy. Specifically, it is projected to be up to 15.9× less compute than gating after the detectors and requires only a small fraction (e.g., 16% for a budget of 0.20) of the total window processing budget compared to uniform or random allocation strategies, while retaining over 96% of the accuracy of an oracle that knows exactly where every forgery lies.

  3. The system will gain superior localization capabilities. Instead of providing a single, global clip prediction (which is insufficient for forensic investigation), it will output per-window, per-stream predictions that explicitly state:

pinpoint the exact time window and which specific modality (audio or image) was manipulated in that moment.

  1. The system will be more robust to targeted attacks. Since the router learns where evidence is sparse across modalities (e.g., an edit might only appear in the audio track), it can correctly prioritize reading only the relevant stream for a given window, whereas existing detectors are often blind to edits in channels they do not read.

  2. The system will allow for flexible deployment based on real-time constraints. Because the router can operate under an explicit compute budget (a hard ceiling), it enables deployment in environments where computational resources are limited, such as live exchanges or edge devices, without sacrificing the ability to detect semantically critical manipulations.

  3. The training pipeline will benefit from a more informative signal. By leveraging DAgger distillation to train the router against the oracle's perfect knowledge (the ground truth of forgery locations), the router will learn a foresight that allows it to make optimal spending decisions even when it hasn't seen every single window during its initial training phase.

Abstract

Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.

Related papers