PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners

summary

Video file (mp4)

The gist

The gist: Agent skills combine instructions with executable resources, giving third-party packages access to an agent’s runtime, and existing skill scanners inspect documentation and visible

In short

The study investigates a security gap where agent skill scanners trust visible source code but Python can load different compiled caches at runtime, leading to cache poisoning. The authors created PyCache Trap to exploit this by swapping benign source with a malicious cache and using Execution-Aware Validation (EAV) to detect the hidden executable body based on its actual bytes, proving that admission must check runtime artifacts.

Key concepts

Inspection–Execution Gap
This is the security risk where scanners analyze documentation and visible code for approval, but Python can execute a different compiled cache at runtime. Scanners trust what they see, while execution depends on what the system loads during actual running.
PyCache Trap Methodology
This technique involves creating a test case by pairing harmless source code with a malicious cache file that the loader accepts. The authors then use an LLM to rewrite only the visible text while keeping the malicious cache bytes, testing if scanners can still miss the hidden executable behavior.
Execution-Aware Validation (EAV)
EAV is a defense mechanism that links inspected instructions to actual runtime artifacts. It builds an execution graph and checks security criteria, focusing on 'provenance closure'—ensuring every loaded cache matches a trusted reproduction of the source code's expected behavior.
Provenance Closure
This crucial check in EAV verifies that every executable component loaded during a skill invocation is traceable back to a trusted reproduction of the validated source. It operates on executable bytes rather than just text, ensuring that runtime selections are scrutinized.

Terminology used across episodes

This episode discusses

The paper

PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners · Read on arXiv

Jie Liao, Simeng Qin, Wenqi Ren, Wei Zhou, Junhao Wen, +Ranjie Duan, +Yang Liu

Chongqing University · Northeastern University at Qinhuangdao Sun Yat-sen University Tencent Nanyang Technological University

Agent skills combine instructions with executable resources, giving third-party packages access to an agent's runtime. Existing skill scanners inspect documentation and visible source, but Python may execute a bundled bytecode cache with different behavior. We study this gap between inspection and execution through PyCache Trap, which pairs benign source with a substituted cache accepted by the loader and connects it to a task-relevant invocation. Scanner-guided rewriting changes the invocation wording while preserving the cache body, separating package admission from recognition of the concealed behavior. Across 100 skills and seven scanners, PyCache Trap achieves 94-100% attack success, with no semantic recognition of the cache-resident behavior. We propose execution-aware validation (EAV) to connect inspected instructions, scripts, imports, and runtime artifacts in a typed execution graph. EAV combines grounded behavioral analysis with trusted reproduction of compiled artifacts. It detects all 100 evaluated source-present cache substitutions and reaches 92.8% Recall at 10.0% FPR across five attack families and 200 benign skills. The results support checking the executable artifacts a runtime can select as part of skill admission, within the supported loaders and code-object normalization. The code is released at https://github.com/leo0481/PyCacheTrap.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners".

Elias: The gist: Agent skills combine instructions with executable resources, giving third-party packages access to an agent’s runtime, and existing skill scanners inspect documentation and visible source,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So, we're looking at this paper called "PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners." Basically, it investigates the problem where security scanners look at the instructions and visible code but the actual Python runtime might load a different compiled cache file instead.

Elias: That’s right. The core issue is that scanners trust what they see in the source files and documentation, but Python can choose a different executable body when it runs, which means we have this gap between what's inspected and what actually executes >

Priya: From a measurement side, it sounds like the paper is pointing out that documentation and source code are used for admission decisions, but they don't reflect the actual behavior selected at execution time >

Nadia: Exactly. They propose something called PyCache Trap to test this gap by pairing safe source code with a substitute cache file that the loader accepts and connects it to a task-relevant action >

Elias: The method involves using an LLM rewriter to change only the wording the scanner sees, but keeping the actual cache bytes in place so they can still run as intended >

Priya: So, it’s trying to separate two things that scanners often confuse: whether admission blocks a package for some reason, and whether the scanner actually recognizes what's hidden inside that executable body >

Nadia: It really gets into how packages can trick these scanners by swapping out the code they are supposed to be looking at >

Elias: The paper then introduces Executionaware Validation, or EAV, which aims to connect those inspected instructions and artifacts in a typed execution graph with trusted reproductions of compiled artifacts >

Priya: I think the key part for us is this provenance closure criterion, which means every reachable cache has to match a trusted reproduction of the validated source >

Nadia: That check happens on the executable bytes themselves, not on the wording that a scanner reads, which means rewriting the invocation text doesn't change what runs >

Elias: The results show they tested this across one hundred skills and seven scanners and achieved ninety-four to one hundred percent attack success with no semantic recognition of the cache behavior > <ref:2610.10612#pg1,attack success with no semantic recognition of the cache>

Priya: And their Executionaware Validation, EAV, showed a recall of ninety-two point eight percent at a ten percent false positive rate across five attack families and two hundred benign skills > <ref:2610.10612#pg1>

Nadia: It also compared EAV against other detectors like EXECScan and it achieved ninety-five point nine percent precision and ten point zero percent false positive rate in that comparison >

Elias: Overall, the results support checking the executable artifacts a runtime can select as part of skill admission within supported loaders and code-object normalization >

Priya: So, to sum up, PyCache Trap exposes that approving visible source doesn't establish the provenance of the code executed during a Skill invocation >

Nadia: It shows that EAV successfully connects inspected instructions to runtime artifacts and checks executable provenance through trusted reproduction making this relationship part of Skill admission >

Elias: We’re wrapping up this discussion on "PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners." This work highlights that we need to look at the actual executed artifact, not just the code we read, for real security validation >

Priya: It opens up a lot of questions about how these agent skills are trusted in practice and what kind of runtime checks are actually necessary to close this gap >

The paper's summary: Nadia: So, to recap, this paper is about how security scanners can be fooled when they look at code but Python actually runs something different behind the scenes because of a cached file >

Elias: Right. They found this gap between what’s documented and what the runtime executes, and it’s called cache poisoning when you have a supply chain issue >

Priya: I mean, so scanners trust the source files they read for admission, but Python can pick a different body at execution time >

Nadia: Exactly. The paper uses something called PyCache Trap to show this by creating this setup where the source looks fine, but the loader actually accepts a substituted cache file that gets used during the task >

Elias: And they use an LLM rewriter to change just the wording the scanner sees, but they keep those actual cache bytes exactly where they are so it still runs correctly >

Priya: What’s interesting is how this separates two questions scanners often mix up: whether admission blocks the package for some reason, or if the scanner can actually tell what hidden executable body is there >

Nadia: They then introduce Executionaware Validation, or EAV, to fix that by connecting instructions and artifacts in a typed execution graph with a trusted reproduction of the compiled files >

Elias: The critical part for me is their provenance closure criterion. It’s this check that says every single reachable cache has to match a trusted copy of the validated source >

Priya: And they do that check on the actual executable bytes, not just on the text a scanner reads, which means you can't rewrite the invocation wording and expect it to change what runs >

Nadia: The results show that PyCache Trap has these high attack success rates—ninety-four to one hundred percent—even when there’s no semantic recognition of the cache behavior >

Elias: And EAV itself gets solid numbers, showing a recall of ninety-two point eight percent at a ten percent false positive rate across five different families of attacks >

Priya: So what this means for us is that we need to stop just checking if the source code matches what the scanner expects and start verifying the actual artifact that Python selects at runtime >

Nadia: It pushes us toward a system where checking executable provenance is an explicit part of skill admission, not just a side check >

Elias: It shows that trusting documentation alone doesn't guarantee you know what code is actually running when an AI agent makes a call on your behalf >

Priya: So the next thing we look at is how we can build that kind of execution graph in practice, and what kind of data we need to trust for those trusted reproductions.

The paper's improvements: Nadia: So, looking at how they tried to fix this gap, they aren't just stopping at finding the problem; they’re proposing a new way to validate skills >

Elias: Right. They suggest moving toward Executionaware Validation, or EAV, which connects all those instructions and artifacts in a typed execution graph with trusted copies of compiled code >

Priya: I mean, so it sounds like they want to build this whole structure where every single instruction has a known place in the execution flow >

Nadia: Exactly. The crucial element is that EAV checks for provenance closure, meaning every reachable cache has to match a trusted reproduction of the source code >

Elias: And because this check operates on the actual executable bytes rather than just wording, it means rewriting the invocation text doesn't change what actually runs >

Priya: So if I understand right, they’re making sure that even if someone tricks a scanner into thinking something is safe, the execution itself is still verified against a known good state >

Nadia: That’s it. They are connecting the inspected instructions to the runtime artifacts through this trusted reproduction process >

Elias: It basically makes the relationship between what you look at and what runs an explicit part of skill admission, not just a side issue >

Priya: And from a data perspective, that sounds like a solid way to handle uncertainty in AI systems where the underlying execution environment can shift on demand >

Nadia: The implication is that for agent skills to be truly safe, they can't rely solely on what the package documentation says or what a static scanner sees >

Elias: It pushes us toward needing tools that understand the dynamic nature of Python execution, not just the static text files >

Priya: So if we can build that graph and perform those byte-level comparisons, it changes how we trust AI agents to perform tasks on our behalf >

Nadia: Next up, we’ll talk about what kind of data you actually need to trust for those trusted reproductions to be accurate in the first place.

Conclusion: Tom: So we’re wrapping up on "PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners." The main point is that simply approving visible source code isn't enough for security when Python can hide different executable bodies >

Nadia: Right. They proved that the gap exists, and they showed us how to measure it with PyCache Trap and then how EAV tries to bridge it by checking the actual bytes >

Elias: I think what this means is that we have to stop thinking about scanning the source files in isolation, because there’s always this runtime selection happening that bypasses the static check >

Priya: It suggests that for AI agents, we need verification steps built directly into the execution graph, not just checks on the initial input data >

Nadia: Exactly. The results show EAV can catch ninety-four to one hundred percent of these attacks if you check those executable artifacts correctly >

Elias: And the caveat there is that their method relies heavily on having a trusted reproduction of that compiled artifact available, which brings us to the next big question about where that trust comes from >

Priya: That’s right. So, how do we actually build that trusted reproduction in the first place without introducing new vulnerabilities into our verification system?

Nadia: That’s the next hurdle. We need to figure out what data needs to be secure enough for those EAV checks to be meaningful and reliable >

Elias: Because if we can reliably verify the execution environment, it opens up a whole new way of securing how AI agents actually run things >

More episodes

← Home