Moirae: A Multimodal Agent Collaborative Framework for Dynamic Android Malware Detection

arXiv:2608.27994 · cs.CR, cs.SE · Submitted 2026-08-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Moirae: A Multimodal Agent Collaborative Framework for Dynamic Android Malware Detection".

Jane: The paper was written by Xueying Zeng, Youquan Xian, Yanze Li, Bowen Hu, Ziqi Shan et al. from Beihang University and Beijing University of Posts and Telecommunications and Guangxi Normal University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We are looking at a fascinating new paper called "Moirae: A Multimodal Agent Collaborative Framework for Dynamic Android Malware Detection" today. It sounds like something out of a sci-fi novel, but it’s actually a very serious approach to mobile security.

Jane: It really is, Tom, and the name Moirae is quite poetic since it refers to the Greek goddesses who weave the threads of fate. The authors, led by Xueying Zeng from Beihang University, are essentially trying to weave different types of digital evidence together to catch bad actors.

Tom: I love that imagery, Jane, because they aren't just looking at one thing and calling it a day. They are building a team of AI agents that work together to see the full picture.

Lu: That collaborative aspect is what gets me incredibly excited about this research! Usually, we see AI models acting like single-purpose tools, but this framework uses multiple specialized agents that can actually reason with one another. It’s moving us toward a more "social" form of intelligence where different perspectives are used to solve a complex problem.

Meng: I wonder how much coordination is actually required for that to work in a real production environment. If you have all these agents talking back and forth, does the system become too heavy or slow to actually use on a phone?

Jane: That's a fair question, Meng, and the paper actually addresses how they manage that complexity through specialized roles. Instead of one giant model trying to do everything, they use smaller, focused agents for things like looking at screenshots or reading system logs.

Lalam: This approach feels much more human-centric in its logic. By mimicking how a person might observe a suspicious app—looking at the screen, seeing what buttons are pressed, and noticing weird background activity—we are creating a digital security layer that understands intent rather than just matching patterns.

Tom: It really shifts the goalposts from "does this code look bad" to "is this app actually behaving badly." We should probably talk about exactly how they gather all those different pieces of evidence.

Summary: Jane: To understand how Moirae works, we have to look at how it gathers its clues while an app is actually running in an emulator. It doesn't just sit there looking at a file; it watches the app live.

Tom: Right, and they aren't just watching one stream of data, Jane. They are capturing what they call three complementary dimensions: the visual stuff you see on the screen, the way a user interacts with UI elements like buttons or menus, and those deep system-level API calls happening in the background.

Jane: Exactly, and it’s like being a detective who is simultaneously watching a suspect's face, listening to their words, and checking their fingerprints. They use something called the ReAct paradigm, which lets these agents "reason" and "act" by looking at the evidence and deciding what to investigate next.

Lu: The way they align those different threads of data is pure genius! Imagine a visual cue like a pop-up appearing on the screen, and the system instantly connects that exact moment to a specific, hidden command being sent to the phone's operating system. It creates this beautiful, causal chain of evidence that is incredibly hard for malware to hide.

Meng: I’m curious about how they handle the massive amount of noise in those system logs. If you're recording every single method call an app makes, wouldn't that overwhelm the agents with useless information?

Jane: They actually built a specific module for that, Meng, which acts like a filter to keep only the most important system calls. This way, they can turn those massive, messy logs into something compact and meaningful that the agents can actually process without getting lost.

Lalam: It’s such a holistic way to build digital trust. By connecting what is visible to the user with what is happening in the hidden layers of the device, we are finally bridging the gap between human perception and machine execution.

Tom: It sounds like they've solved a lot of the "blind spot" problems we see in older detectors, so let's look at how well this actually works when things get tough.

Improvements: Tom: The results from the experiments are honestly staggering, especially when you consider that they tested this on apps they had never seen before. They hit an accuracy of ninety point zero six percent without any special fine-tuning for those specific datasets.

Jane: That's the part that really matters, Tom, because it proves they can handle what researchers call "concept drift." In plain English, that means when malware creators change their tactics to stay ahead of security software, Moirae is still able to recognize the underlying malicious intent.

Tom: I saw in their data that while older models were struggling and seeing their performance tank as the years went by, Moirae stayed incredibly stable. It was still hitting around ninety percent accuracy even on samples from two thousand twenty-one despite being tested against much older training data.

Lu: This is a massive leap toward truly autonomous security! We are moving away from models that just memorize old signatures and toward models that actually understand the "why" behind a behavior. It's like teaching a guard to recognize the intent to steal rather than just memorizing what every thief's face looks like.

Meng: I noticed they also talked about token efficiency, which is something I think we all need to care about. They managed to compress those long, messy behavioral traces by anywhere from ten to one hundred times before handing them to the judge agent.

Jane: That’s a huge practical win, Meng, because it means they aren't wasting massive amounts of computational power on redundant data. It makes the whole reasoning process much faster and more scalable for real-world use.

Lalam: And that efficiency is what allows this technology to eventually become a seamless part of our digital lives. If security is smart but lightweight, it can protect us constantly without us ever feeling the weight of it, which builds a much deeper sense of safety in our global digital culture.

Tom: It really seems like they've found a way to make AI reasoning actually work for security at scale.

Conclusion: Jane: We have covered a lot of ground today, from the Greek-inspired name to the complex multi-agent system that makes Moirae so effective. It’s clear that combining visual, interaction, and system data is a game-changer for Android security.

Tom: It really is, Jane. By focusing on intent rather than just code patterns, they've built something that can actually grow alongside the threats it's fighting.

Lu: I just keep thinking about how this sets a new standard for all multi-agent research. If we can coordinate agents this well to solve security problems, imagine what they could do for scientific discovery or complex system management!

Meng: From my side, the engineering discipline they showed in managing token costs and data noise is what makes me most optimistic. This isn't just a cool academic experiment; it’s a blueprint for how to actually deploy LLM-based reasoning in the real world.

Lalam: My final thought is about the peace of mind this brings. As our lives become more intertwined with mobile devices, having a protector that truly understands the nuance of human interaction and machine action will be essential for maintaining our trust in technology.

Jane: Well, on that high note, we'll wrap up our discussion on "Moirae: A Multimodal Agent Collaborative Framework for Dynamic Android Malware Detection."

Tom: Thanks for joining us! We'll see you next time when we break down another incredible paper. Goodbye everyone!

Beihang University · Beijing University of Posts and Telecommunications · Guangxi Normal University

cs.CR, cs.SE

Submitted: 2026-08-28

Updated: 2026-09-16

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: Moirae is a multimodal agent collaborative framework designed for dynamic Android malware detection to address the challenge of "concept drift." It matters because traditional machine learning

Key concepts

Multi-agent Collaborative Framework
Instead of using one massive model, Moirae employs multiple specialized AI agents for specific tasks, such as analyzing screenshots or reading system logs. These agents communicate and reason with one another to build a complete picture of an app's behavior and intent.
ReAct Paradigm
This approach allows AI agents to both "reason" and "act" based on the evidence they observe. It enables the system to create a causal chain of evidence, connecting visible events on a screen directly to hidden background commands happening within the operating system.
Concept Drift
Concept drift occurs when malware creators change their tactics to evade security software. Moirae overcomes this by focusing on the underlying intent behind an app's behavior rather than just matching old code patterns, allowing it to remain effective against newer, unseen threats.

Terminology

Summary

Moirae is a multimodal agent collaborative framework designed for dynamic Android malware detection to address the challenge of concept drift. It matters because traditional machine learning detectors are vulnerable to concept drift due to their reliance on implementation-specific features, whereas Moirae leverages Large Language Models (LLMs) to move beyond implementation-specific feature fitting and instead infer high-level malicious intent.

The core motivation

The Android ecosystem faces persistent and rapidly evolving malware threats, making traditional detectors vulnerable because they rely on features whose distributions change over time. While LLMs offer strong reasoning, current detectors are often code-centric or depend on a single-dimensional source of evidence, leaving them susceptible to obfuscation. Moirae addresses this by capturing runtime evidence across three complementary dimensions: visual presentation, UI interaction transitions, and system-level API operations. By linking what users see with how they interact and what the application performs, the framework can infer high-level malicious intent from evolving low-level implementations.

The multi-agent architecture

Moirae utilizes a three-phase process to reconstruct cross-level behavioral chains. During the dynamic preprocessing phase, the system drives an application within an emulator to synchronously record screenshots, XML state trees, and ART method invocation traces. A Cross-Modal Causal Evidence Fusion module then aligns these datasets by defining a precise time window to accurately bind underlying API call sequences to the superficial UI interaction events.

The core reasoning is performed by specialized agents using the ReAct (Reasoning and Acting) paradigm:

  • The Vision Agent (A vis) performs global scanning and in-depth scrutiny of the GUI to identify visual inducements such as deceptive pop-up ads or extortion interfaces.

  • The UI/Event Agent (A ue) conducts temporal analysis on XML view trees to find semantic misalignments between visible UI intentions and actual behaviors.

  • The API Agent (A api) executes micro-level dynamic forensics to condense physical behavior into a system-level abuse feature representation.

  • The Judge Agent (A judge) aggregates these findings to output a final verdict, a threat family classification, and a highly logically interpretable structured cross-validation evidence chain.

Experimental validation and robustness

Experiments on temporally and distributionally unseen datasets show that Moirae achieves an accuracy of 90.06% without fine-tuning, outperforming state-of-the-art baselines. While traditional supervised detectors experience substantial performance degradation when faced with new malware variants, Moirae demonstrates strong zero-shot generalization against Android malware concept drift.

In temporal robustness tests using AndroZoo samples from 2017 to 2021, baseline models showed significant decay; for example, CL-Malware's accuracy dropped from 92.46% to 61.50%. In contrast, Moirae's accuracy remained stable between approximately 87%–92%. Any remaining performance loss in recall is attributed to the inherent difficulty of triggering and observing complete malicious behaviors during dynamic analysis.

Token efficiency and scalability

To manage the computational demands of LLMs, Moirae employs a hierarchical evidence extraction method that transforms lengthy behavioral traces into compact semantic representations. This mechanism achieves a compression ratio of approximately 10× to 100×, which prevents the agent from exceeding its context window and introducing substantial inference latency and computational cost. This allows for efficient reasoning where the Judge Agent consumes only a small fraction of the total tokens after receiving these compressed representations.

Improvements for AI systems

1. Implementation of Cross-Modal Causal Evidence Fusion Layers

  • The Improvement: Integrate a temporal alignment mechanism that binds high-level semantic events (UI interactions/XML state changes) to low-level system execution traces (API method invocations) using a defined time window (t).

  • What the improved AI can do: The system can move beyond simple correlation to true causal reasoning. It will be able to detect semantic misalignments—for example, identifying when a visually benign Close button trigger-links to a hidden background API call for SMS transmission, effectively defeating obfuscation techniques that hide malicious intent behind legitimate-looking interfaces.

2. Hierarchical Multi-Agent Evidence Abstraction (Token Compression Pipeline)

  • The Improvement: Replace single-prompt long-context analysis with a multi-stage pipeline where specialized extractor agents transform high-volume, redundant telemetry (raw system logs, pixel matrices, and XML trees) into dense, compact semantic representations (10 times to 100 times compression).

  • What the improved AI can do: The system can perform deep forensic analysis on extremely long execution sessions that would otherwise exceed LLM context windows or cause lost-in-the-middle hallucinations. It allows a central Judge Agent to reason over massive datasets with high efficiency and low computational latency by processing only distilled behavioral intent rather than raw, noisy data.

3. Heterogeneous Multi-Agent Adjudication with Configurable Decision Biases

  • The Improvement: Deploy an ensemble of specialized agents (Visual, Interaction, and API experts) paired with a final Judge Agent that utilizes a personality-based decision framework (e.g., integrating high-precision models like Gemini for verification and high-recall models like Qwen for discovery).

  • What the improved AI can do: The system can be dynamically tuned for specific deployment environments. In consumer-facing applications, it can be configured to prioritize Precision (minimizing false positives to prevent user frustration); in high-security enterprise or banking environments, it can be pivoted to prioritize Recall (ensuring no potential threat is missed), providing a flexible security posture.

4. Transition from Implementation-Specific Feature Fitting to Intent-Based Reasoning

  • The Improvement: Shift the training and reasoning objective from statistical feature matching (API names, permission strings, or code signatures) to high-level behavioral intent modeling (detecting patterns of deception, extortion, or unauthorized data exfiltration).

  • What the improved AI can do: The system will achieve high Temporal Robustness against concept drift. Even as malware authors evolve their code via reflection, dynamic loading, or new API combinations to evade traditional detectors, the AI will remain effective because it recognizes that the underlying malicious objective (e.g., a deceptive prompt followed by a sensitive data grab) remains constant regardless of the implementation.

Abstract

The Android ecosystem faces persistent and rapidly evolving malware threats. Existing machine learning detectors are vulnerable to concept drift because they rely on implementation-specific features whose distributions change over time. Large language models (LLMs) offer strong semantic understanding and zero-shot reasoning, but current LLM-based detectors typically depend on code-centric or single-dimensional evidence, making them susceptible to obfuscation and limiting comprehensive behavior analysis. We present, a multimodal agent collaborative framework for dynamic Android malware detection. dynamically collects multimodal runtime evidence and employs ReAct-based specialized agents to analyze complementary behavioral views. The detection process begins by identifying visual deception cues, modeling UI state transitions, and integrating runtime API behaviors to fuse multi-dimensional evidence across user-visible interfaces and hidden backend operations. Experiments on temporally and distributionally unseen datasets show that achieves an accuracy of 90.06% without fine-tuning, outperforming state-of-the-art baselines and demonstrating strong zero-shot generalization against Android malware concept drift.

Sources

Related papers