Beyond Source: An Empirical Study of Python Bytecode Security Risks
Baihong Chen, Tian Xie, Wen Li
Utah State University
cs.CR
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/zrsx/pycdc
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper presents an empirical study of Python bytecode as a security artifact, motivated by the observation that "Python package security is largely source-centric, yet Python runtimes can execute
Terminology
Summary
This paper presents an empirical study of Python bytecode as a security artifact, motivated by the observation that Python package security is largely source-centric, yet Python runtimes can execute bytecode directly through.pyc files, compiled-only modules, and marshalled code objects, creating an inspection-execution gap.
The study is structured around four research questions (RQ1-RQ4) and uses a staged empirical design that separates ecosystem transparency from adversarial interpreter robustness.
RQ1: Bytecode Exposure and Packaging. The study collected 1,034,843 package-balanced PyPI artifacts (wheels and source distributions) and found that 7,388 artifacts contain at least one.pyc file, yielding a total of 228,578 bytecode files.
Notably, 28,193 of the 228,578.pyc files (12.33%) are artifact-local source-less.pyc files under our conservative artifact-local source check: they lack a corresponding.py file within the same distribution artifact.
Within the 21,541 in-scope CPython 3.8–3.14 artifact-local source-less.pyc files, roughly two-thirds are shipped import caches while one-third are bare compiled-only modules, the form most consistent with deliberate source omission.
The study also found that dynamic-loading indicators appear in 242,586 artifacts (23.44%).
RQ2: Practical Analyzability. The study evaluated version-aware tooling (marshal, dis, Decompyle++, PyLingual) on 204,904 in-scope.pyc files from CPython 3.8–3.14. Results show that "practical analyzability is effectively universal for the modern CPython bytecode observed in PyPI. PyLingual successfully decompiles 204,901 of 204,904 in-scope files, allowing nearly all files to reach L4, meaning that at least one selected decompiler emits source. However, the paper clarifies that
L4 records successful source emission, not verified functional equivalence. The tools are not robust:
across unmodified PyPI bytecode, we observe 50 robustness-failure tool results spanning marshal, dis, and PyLingual (six signatures, dominated by dis), caused by hangs or uncaught exceptions; Decompyle++ produces only controlled rejections."
RQ3: Interpreter Robustness Under Adversarial Bytecode. The study fuzzed version-matched CPython interpreters (3.8–3.14) with mutated bytecode seeds derived from CPython's own unittest suite. Across seven 24-hour campaigns, Honggfuzz retained 12,404 crash-triggering inputs... Stack-based deduplication reduced these crashes to 1,009 unique crash groups.
The findings are dominated by pointer-dereference symptoms
with 261 groups exhibit potential memory-corruption characteristics.
Critically, at least 925 groups (91.7%) reach interpreter execution beyond the ingestion boundary,
meaning they crash in post-ingestion execution contexts (frame evaluation, object runtime, garbage collection, instrumentation) rather than just during marshal deserialization, which is documented as unsafe.
RQ4: Source Reproduction. The study attempted to reproduce the 1,009 bytecode-level findings through recovered Python source using PyLingual and Decompyle++. The result is that None of the 1,009 stack-deduplicated runtime findings can be reproduced through recovered Python source.
The failures are categorized as: 662 findings produce no source reproducer, 285 produce source-like output that fails to compile, and 62 compile but do not reproduce the original bytecode behavior on source rerun.
Additionally, the reproduction attempts themselves exposed tool robustness failures: 23 tool runs (11 distinct signatures) drive decompilers into uncaught exceptions, timeouts, or signal-terminated subprocess failures.
The paper concludes that bytecode introduces a measurable gap between what package-security workflows inspect, what Python runtimes execute, and what source-level artifacts can faithfully represent.
The authors recommend that "Python package security should therefore inventory bytecode directly, analyze it with version-aware tools, triage runtime findings at the bytecode level, and avoid collapsing bytecode evidence into source-level vulnerability claims."
Improvements for AI systems
Improvements to AI Systems Based on This Paper:
- Bytecode-Aware Static Analysis Module
-
Improvement: Extend existing code-analysis AI (e.g., vulnerability scanners, code review assistants) to ingest
.pycfiles and marshalled code objects directly, not just.pysource. -
Capability: The AI can now flag malicious or vulnerable logic in compiled-only Python packages (e.g., obfuscated backdoors in bare
.pycmodules) that source-centric tools miss, closing the inspection-execution gap.
- Version-Aware Decompilation and Semantic Equivalence Checking
-
Improvement: Train a neural decompiler or post-processor that maps CPython 3.8–3.14 bytecode to source, but with a verification layer that checks functional equivalence (not just syntactic emission) using differential execution or symbolic traces.
-
Capability: The AI can reliably reconstruct source from bytecode and certify whether the recovered source truly matches runtime behavior, reducing false confidence from tools like PyLingual that only guarantee source emission.
- Adversarial Bytecode Robustness Triage for Interpreter Security
-
Improvement: Build an AI-based crash triage system that automatically classifies fuzzer-found interpreter crashes (e.g., from Honggfuzz) by execution stage (ingestion vs. post-ingestion) and memory-corruption potential, using the paper’s 1,009 deduplicated crash groups as training data.
-
Capability: The AI can prioritize which bytecode-level vulnerabilities to patch in CPython, focusing on the 91.7% that reach frame evaluation or GC, and can generate minimal reproducers directly from bytecode, bypassing unreliable source recovery.
- Source-Reproduction Failure Predictor
-
Improvement: Train a classifier that, given a bytecode file and a decompiler’s output, predicts whether the source will (a) fail to compile, (b) compile but diverge in behavior, or (c) faithfully reproduce the original runtime behavior—using the paper’s 662/285/62 failure distribution as labels.
-
Capability: The AI can automatically reject decompiled source that is unsafe for security analysis, preventing analysts from wasting time on non-reproducible findings and avoiding false vulnerability claims.
- Dynamic-Loading Indicator Enrichment for Supply-Chain Risk Scoring
-
Improvement: Integrate the paper’s finding that 23.44% of PyPI artifacts contain dynamic-loading indicators (e.g.,
ctypes,dlopen,importlibof bytecode) into an AI-based package risk scorer. -
Capability: The AI can flag packages that execute bytecode at runtime (even if no
.pycis shipped) as higher-risk, and can trace those dynamic loads to specific bytecode objects for deeper inspection, improving automated malware detection in CI/CD pipelines.
- Robustness-Aware Tool Orchestration
-
Improvement: Create an AI agent that dynamically selects and combines
marshal,dis,Decompyle++, andPyLingualbased on bytecode version and complexity, and that detects when a tool is about to hang or crash (using the 50 observed failure signatures). -
Capability: The AI can reliably analyze large-scale PyPI bytecode corpora without tool-induced downtime, and can fall back to alternative decompilers or raw disassembly when one tool fails, ensuring near-universal analyzability (matching the 204,901/204,904 success rate).
- Bytecode-Level Vulnerability Evidence Management
-
Improvement: Design an AI system that stores and queries security findings at the bytecode level (e.g., specific opcode sequences, marshal payloads) rather than collapsing them into source-level claims, as recommended by the paper.
-
Capability: The AI can link a runtime crash or malicious behavior directly to the exact
.pycfile and bytecode offset, enabling reproducible, version-pinned security advisories that do not degrade when source is unavailable or non-reproducible.
Sources
- An Empirical Analysis of the Python Package Index (PyPI)
- An Empirical Study of Malicious Code In PyPI Ecosystem
- The Art, Science, and Engineering of Fuzzing: A Survey
- SpellBound: Defending Against Package Typosquatting
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs