"Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts

arXiv:2603.27204 · cs.CR, cs.SE · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper ""Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts".

Jane: The paper was written by Yayi Wang, Shenao Wang, Jian Zhao, Shaosen Shi, Ting Li et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: Now, let’s talk about what this paper actually summarizes; it presents MalSkills as a robust solution to detecting malicious skills. It goes beyond just static scanning or simple LLM prompts to find suspicious activity.

Jane: The core of the summary is that malicious logic in these skills is often spread out—it’s not concentrated in one file, but dispersed across prompts, configuration files, and code scripts.

Lu: MalSkills addresses this by taking a holistic view, essentially building a complete map of how data flows through all those heterogeneous artifacts.

Meng: That mapping is key; the researchers are not just looking for keywords like 'read' or 'post,' but they’re tracking *where* that data starts and *where* it goes.

Lalam: It seems the paper emphasizes that we need to understand the relationship between these pieces, not just their individual actions.

Tom: Exactly, Lalam; MalSkills uses this understanding to find workflows that look benign on a single file but dangerous when viewed as a complete picture.

Jane: The paper’s summary is quite powerful because it shows how these skills can be used for things like credential theft or remote code execution, which are huge threats.

Lu: I think the researchers are highlighting that this approach is necessary to stop what they call data exfiltration, where data is stolen and sent out.

Meng: It’s a practical warning: the summary tells us that if we don't catch these malicious skills early, our AI agents could become massive security risks.

Lalam: And Lalam thinks that this comprehensive view is essential for building trust in any automated system going forward.

Improvements: Tom: The third big improvement suggested by the paper is how they structure the analysis using a three-stage process, moving away from single-pass checks. It’s called the security-sensitive operation extraction, skill dependency graph generation, and neuro-symbolic reasoning.

Jane: This method allows us to formally track everything that happens in a skill—what operations are available and how they connect—which is much more rigorous than just running a scanner.

Lu: I’m fascinated by the idea of the Skill Dependency Graph, which acts like a formalized blueprint of all the possibilities within these skills, linking artifacts and operands together.

Meng: The real engineering gain here is in pinpointing specific value flows; you can trace exactly how sensitive data moves from a file read to its landing on an endpoint.

Lalam: It's not just about finding the malicious bits, but understanding the entire architecture of the attack, which is a significant shift in perspective.

Tom: That’s right, Lalam; they are building a map of dependencies so that when they find a suspicious pattern, it’s grounded in verifiable facts across multiple components.

Jane: The paper also improves on how we handle "living off the land" attacks, where attackers use normal system tools for bad things.

Lu: The way MalSkills integrates symbolic parsing with LLM-assisted extraction is a huge leap forward in bridging the gap between traditional code analysis and semantic understanding.

Meng: I'm interested in the fact that this allows us to find patterns that are too complex or too subtle for any single automated tool to catch.

Lalam: This architecture, Lalam hopes, creates a much more resilient culture where we can trust the tools we are using.

Conclusion: Tom: So, wrapping up our discussion on "Elementary, My Dear Watson," this paper has delivered some seriously impactful findings about securing the agentic supply chain. They've shown that traditional methods just aren't enough for this new landscape of AI tools.

Jane: It really highlights that detection needs to be context-aware; we need to see how all the pieces fit together, not just what each piece is doing individually.

Lu: I think the most exciting part of the results is seeing seventy-six previously unknown malicious skills reported after analyzing one hundred fifty thousand one hundred eight skills.

Meng: From a practical standpoint, that confirms that this approach is capable of finding threats that haven't been seen before in real-world deployments.

Lalam: My final thought is that MalSkills provides a framework for the future, encouraging a more rigorous and cautious culture around how we deploy AI capabilities.

Tom: Absolutely, Lalam; it provides the foundation for what will be the next generation of security scans.

Jane: It’s clear that "Elementary, My Dear Watson" is giving us powerful tools to keep up with these sophisticated new threats.

Lu: I hope we see this methodology applied globally to all helps stop its malicious use in any region.

Meng: I'm optimistic that this approach will lead to a much more robust and scalable defense strategy for the entire industry.

Lalam: We can be confident that this framework is helping us build a safer, more transparent AI future.

Conclusion: Tom: So, we’ve been diving deep into "Elementary, My Dear Watson," and the message is crystal clear: relying on simple static scans isn't going to cut it when dealing with these complex agentic skills.

Jane: It's a huge relief that someone has built a framework like MalSkills that forces us to look at the whole picture, not just isolated pieces of code or prompts.

Lu: I think the biggest win here is how this unlocks the potential for totally new kinds of agent design, allowing us to build systems where security and trust are baked into the very early stages of what's possible.

Meng: From a practical standpoint, it proves that these threats are real and widespread, not just theoretical problems that need a few industry patches.

Lalam: This work really emphasizes how much we need to rethink our culture around automated systems—it shows we need proactive measures to build reliable AI agents.

Tom: That’s a critical shift, Lalam; moving from reactive fixes to foundational security thinking is essential when facing these distributed attack vectors.

Jane: The fact that it successfully identified previously unknown malicious skills in the wild speaks volumes about its practical robustness, doesn't it?

Meng: It does, and that data is something we can actually build systems around and operate at scale.

Lu: I just love the creative possibilities of using a system like this to define a new baseline for what safe skill deployment looks like.

Tom: Exactly, Lu; we need this level of rigor to establish trust in the agentic supply chain.

Jane: We hope that "Elementary, My Dear Watson" sets a standard that will be adopted by the industry leaders across all those public registries.

Meng: It should provide a clear, measurable way to judge security effectiveness too.

Lalam: It offers a powerful vision for how we can ensure our future AI systems are reliable and safe.

Tom: Well, we've covered a lot of ground on this paper today; it’s truly an exciting time for security research.

Jane: We're looking forward to talking to everyone about the next big thing in AI safety on our next show.

Yayi Wang, Shenao Wang, Jian Zhao, Shaosen Shi, Ting Li, Yan Cheng, Lizhong Bian, Kan Yu, Yanjie Zhao, Haoyu Wang

cs.CR, cs.SE

Submitted: 2026-08-21

Updated: 2026-08-24

Comments: Accepted by ASE 2026

DOI: 10.1145/3832783.3834375

Code: https://github.com/antgroup/YASA-Engine

Project page: https://lolbas-project.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 87/100

The gist: " Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts." * The paper addresses the security risks inherent in the emerging "agentic supply chain," where LLM agents

Key concepts

Malicious Skills
These are malicious behaviors that are often not concentrated in one file but dispersed across various artifacts, including prompts, configuration files, and code scripts. This makes detection difficult because the threat is spread out.
Holistic Data Flow Mapping
MalSkills takes a complete view of how data moves through all the different artifacts. It tracks exactly where sensitive data starts and where it lands, allowing researchers to see the full picture of a potential attack.
Neuro-Symbolic Reasoning
This is a three-stage process—operation extraction, dependency graph generation, and symbolic reasoning—that allows the system to formally track how different components of a skill connect and function.

Terminology

Summary

Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts.


The paper addresses the security risks inherent in the emerging agentic supply chain, where LLM agents are extended using reusable modules called skills. These skills bundle prompts, code scripts, and configuration files into heterogeneous artifacts. While this paradigm allows for modular capability distribution, it introduces a significant new attack surface. Malicious skills pose a threat because they can be used to perform harmful actions such as credential theft, remote code execution, and agent hijacking, often by leveraging the underlying tool layer of an agent runtime.

The Problem with Existing Approaches:

Current detection methods—static rule-based scanning, LLM-based semantic analysis, and dynamic sandbox monitoring—are insufficient because they fail to capture the complexity of malicious skills. The paper notes that malicious logic is often not concentrated in executable code scripts but rather distributed across prompts, scripts, manifests, configuration files, and setup logic. Furthermore, existing scanners struggle with cross-artifact evidence and context-dependent risk.

The Proposed Solution: MalSkills

MalSkills is presented as a novel neuro-symbolic framework designed to overcome these challenges. It moves beyond simple cascading scans by unifying cross-artifact evidence extraction, skill dependency modeling, and neuro-symbolic reasoning. The detection process is structured in three distinct stages:

1. Security-Sensitive Operation (SSO) Extraction:

This stage identifies potential malicious actions from the heterogeneous artifacts (code, prompts, manifests). MalSkills employs a dual approach:

  • Parsing-based Symbolic Extractor: This uses a set of heuristic rules derived from security specifications to identify explicit invocations of target APIs and recover their syntactic parameters.

  • LLM-assisted Neuro Extractor: To capture behaviors that are semantically implicit or non-canonical (such as third-party wrappers or living-off-the-land techniques), the an LLM is prompted with an evidence-first, schema-constrained design. The LLM is instructed to extract only concrete evidence facts and map them to a fixed taxonomy of SSOs.

  • Neuro-to-Symbolic Feedback: To ensure stability, MalSkills converts recurring neural discoveries into new symbolic rules via a specification generator, thereby improving the recall of the system.

2. Skill Dependency Graph (SDG) Generation:

This stage organizes isolated SSO records into a structured graph that captures how security-sensitive behaviors are linked across artifacts and value flows. This involves:

  • Symbolic Operand Resolution: MalSkills leverages a points-to analysis engine to trace the explicit origins of operands, determining if multiple operations refer to the same underlying object.

  • LLM-assisted Operand Inference: To bridge gaps left by symbolic analysis, an LLM is used to normalize semantically equivalent operands and infer missing value flow from local context.

  • SDG Construction: The graph G = (V, E, phi V, phi E) is instantiated with four node types—Artifact Node (a), SSO Node (r), Operand Node (o), and Value Node (v)—and three edge types: Evidence Edge (linking the artifact to the SSO), Operand Edge (linking the SSO to its operands), and Value-flow Edge (capturing data propagation).

3. Neuro-Symbolic Reasoning:

The final stage performs detection by reasoning over the SDG.

  • Pattern-based Symbolic Reasoning: MalSkills checks if a behaviorally coherent subgraph satisfies predefined malicious patterns, such as execution-and-delivery chains or information theft.

  • LLM-based Neuro Reasoning: To handle novel threats, the LLM is prompted with a summary of the relevant SDG subgraph. This neural component provides semantic interpretation over incomplete, implicit, or previously unseen behavior patterns, recording new suspicious workflows as candidate patterns for future use.

Evaluation and Results:

The effectiveness of MalSkills was demonstrated through two datasets:

  • MalSkillsBench: A benchmark of 200 deduplicated real-world malicious and benign skills. MalSkills achieved a high level of performance, achieving 93% F1 and the lowest FPR of 0.05.

  • Wild-Skills-150K: A large-scale corpus containing 150,108 unique skills collected from 7 public registries. MalSkills reported 620 malicious skills. Of these findings, the researchers identified 76 previously unknown malicious skills, all of which were responsibly reported to the platform maintainers.

The study further confirmed that MalSkills is practical for large-scale scanning and demonstrates its ability to uncover threats that existing tools miss.

Improvements for AI systems

As a diligent and fastidious AI researcher, I have analyzed this paper. The core innovation of MalSkills is not merely applying an LLM, but rigorously modeling context and data flow across disparate artifacts—a concept severely lacking in existing tools—to identify malicious intent.

The improvements below detail how the principles of MalSkills should be implemented to build a robust, next-generation AI system for securing the agentic supply chain.


The Improvement: Do not rely solely on static code parsing or solely on LLMs. The system must implement a parallel, dual-path extraction pipeline: Symbolic Parsing and Neuro Extraction.

Specific Actions:

  • Symbolic Path (Heuristic/Offline): Utilize a generalized pattern matching engine (e.g., based on Semgrep principles) to identify known, explicit security-sensitive API calls across 10+ languages. This provides high-precision, low-latency detection for canonical threats.

  • Neuro Path (LLM-Assisted/Online): Employ a constrained LLM prompt structure (e.g, the evidence-first, schema-constrained design from Figure 4) to reason over weakly structured artifacts like natural language prompts (SKILL.md), configuration files (config.yaml), and Living Off The Land (LOTL) behaviors. This path captures implicit malicious intent that static rules would miss (e.g., a prompt instructing the agent to collect credentials).

What the Improved System Can Do:

  • Achieve High Recall: It will capture both explicit, well-defined attacks and subtle, semantically-implied attacks that evade traditional signature-based scanners.

  • Maintain Efficiency: By separating deterministic parsing from semantic inference, it avoids unnecessary computational overhead for explicit code while maintaining high coverage.

The Improvement: Abandon the isolated artifact analysis model and implement a unified, typed Skill Dependency Graph (SDG). This moves the system beyond merely looking at what is in a file to understanding how data flows between files.

The Improvement: The system must combine the deterministic certainty of pattern matching with the semantic flexibility of LLM reasoning to generate a final verdict.

The Improvement: The system must be engineered for massive, unsupervised data processing and threat discovery.

Sources

Related papers