A Time Series Analysis of Malware Uploads to Programming Language Ecosystems
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "A Time Series Analysis of Malware Uploads to Programming Language Ecosystems".
Jane: The paper was written by Jukka Ruohonen and Mubashrah Saddiqa from University of Southern Denmark.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone! Tom here, and I've got Jane with me, and we are diving into a paper that's got a serious title and serious implications: "A Time Series Analysis of Malware Uploads to Programming Language Ecosystems."
Jane: And Tom, I have to say, when I first read that title, I thought, okay, this is a niche security paper. But the more we dug in, the more I realized this is about the plumbing of the entire internet. These are the ecosystems—npm, PyPI, RubyGems—where developers grab the building blocks for basically every modern app.
Tom: Exactly! And the authors, Jukka Ruohonen and Mubashrah Saddiqa from the University of Southern Denmark, they’ve done something pretty clever. They didn't just count malware; they looked at it over time. They built a time series. So instead of a snapshot, we get a movie of how these attacks have evolved.
Jane: And that movie has a pretty scary plot twist. The paper shows that, recently, the number of detected malware uploads has actually surpassed the number of traditional vulnerability reports. Think about that for a second.
Tom: It’s a huge shift. For years, we worried about "vulnerabilities"—like a crack in the foundation of a house. But now, we’re talking about active break-ins. Malware is not a flaw; it’s a burglar that’s already inside the house.
Jane: Right. And the title hints at the scope—"Programming Language Ecosystems." These aren't just random file-sharing sites. We’re talking about the official repositories for Python, JavaScript, Ruby, Java, and Go. The places where millions of developers get their code.
Tom: So when the paper talks about "malware uploads," it’s not just a theoretical risk. It’s about someone uploading a poisoned package that looks exactly like a legitimate one, and developers accidentally installing it. The paper calls it "typo-squatting," where you misspell a name like "requests" and get "requestz" instead.
Jane: And the authors are asking a really fundamental question with this research. They want to know if we can predict these trends. Can we see the wave coming before it crashes? That’s the promise of the time series analysis, and that’s what we’re going to dig into next.
Tom: Stay with us, because we’re about to look at the actual data and see just how bad the malware problem has gotten in these ecosystems.
Summary: Jane: So, Tom, we’ve set the stage with the title. Now let’s talk about what the paper actually found. The summary in the abstract is pretty stark: malware uploads have "recently surpassed" vulnerability reports in the OSV database.
Tom: And that OSV database—the Open Source Vulnerabilities database—that’s the key data source here. The authors pulled a snapshot from April two thousand twenty-five and looked at six ecosystems: CRAN, Go, Maven, npm, PyPI, and RubyGems.
Jane: And the numbers are wild. Let’s talk about npm first, because that’s the JavaScript ecosystem. The paper found over twenty thousand malware entries there. Twenty thousand! That’s compared to about four thousand vulnerability reports.
Tom: It’s a complete inversion. And PyPI, the Python ecosystem, is right behind it with nearly nine thousand malware entries. RubyGems is third with about eight hundred. Meanwhile, CRAN, Go, and Maven have almost none—less than ten total.
Jane: Which is fascinating, right? Why the difference? The paper suggests it might be about detection capabilities or the sheer size of these ecosystems. npm and PyPI are massive, so more packages mean more opportunities for bad actors to slip something through.
Tom: And the authors also looked at something called the "malware share"—the percentage of all entries that are malware. In early two thousand twenty-five that share hit up to eighty percent in the OSV database. That means four out of every five new entries about these ecosystems were about malware, not vulnerabilities.
Jane: That’s the headline number, Tom. It completely reframes how we think about supply chain security. We used to worry about accidentally using a package with a bug. Now we have to worry about intentionally malicious packages.
Tom: But here’s the thing—the paper doesn’t just stop at counting. They built statistical models to see if they could predict these trends. And that’s where it gets really interesting. They used something called an autoregressive distributed lag model.
Jane: Which sounds complicated, but the idea is simple. They’re asking: does the number of ecosystems reporting malware today predict the number of malware reports tomorrow? And does media attention or security advisories drive more discoveries?
Tom: And the answer is yes. The model found that the number of ecosystems is a strong predictor. If malware shows up in three ecosystems one week, you can expect a spike in total reports soon after. It’s like a contagion effect.
Jane: And the media articles—the publicity—also had a measurable effect. When there’s a big story about a malware campaign, more people go looking, and they find more. It’s a feedback loop.
Tom: So the summary is this: malware is not just a problem; it’s a growing, predictable problem. And the paper gives us the tools to see it coming. Let’s talk about what that means for the people actually defending these systems.
Improvements: Tom: So, Jane, we’ve established that malware is surging and that we can model the trends. But what does this paper actually suggest we *do* about it? What are the improvements?
Jane: Well, Tom, the paper doesn’t hand us a silver bullet, but it does point to some practical steps. One of the big ones is the idea of "curated lists" for safe packages. Think of it like a trusted grocery list for developers.
Tom: That’s a great analogy. Instead of letting developers pick any package off the shelf, you have a list of pre-vetted, signed packages. The paper mentions that Linux distributions used to do this kind of quality gating, but that got lost when these programming language ecosystems took over.
Jane: And it ties into the discussion about code signing. If a package is signed by a trusted entity, you know it hasn’t been tampered with. That’s a concrete improvement that could stop a lot of these attacks.
Tom: But there’s also a call for better detection and monitoring. The paper suggests that the current state of malware detection in these ecosystems might not be as good as we think. They even hint that we need better benchmark datasets to test our detection tools.
Jane: And that’s a really important point. The paper notes that some previous research tried to verify malware by checking if a package was absent from PyPI. But since PyPI has so much malware itself, that check is basically useless. We need better ground truth.
Tom: Right! And then there’s the regulatory angle. The paper brings up the EU’s Cyber Resilience Act, the CRA. It’s a new law that requires products to ship without known vulnerabilities. And the paper argues that this should extend to malware too.
Jane: That’s a big deal. If a company accidentally ships a product with a malware-ridden dependency, they could face regulatory sanctions. That’s a huge incentive to clean up the supply chain.
Tom: And the paper even suggests that regulators could use time series analysis like this to decide when to do "sweeps"—coordinated checks of specific product categories. If the data shows a spike in malware in npm, regulators could target that ecosystem.
Jane: So the improvements are really about shifting from reactive to proactive. Instead of waiting for malware to be discovered, we use these predictive models to anticipate where the next wave will hit.
Tom: And that’s the exciting part. This paper gives us a roadmap. It’s not just a warning; it’s a plan. Let’s wrap this up and see what the big picture is.
Conclusion: Jane: Alright, Tom, we’ve covered the title, the data, and the recommendations. Time to say goodbye to "A Time Series Analysis of Malware Uploads to Programming Language Ecosystems."
Tom: And what a journey it’s been. We started with a simple question—how much malware is out there?—and ended up with a full picture of a shifting threat landscape.
Jane: The key takeaway, and I want to make sure our listeners get this, is that malware has overtaken vulnerabilities as the primary concern in these ecosystems. It’s not about bugs anymore; it’s about active attacks.
Tom: And the paper showed us that npm and PyPI are the main battlegrounds, with tens of thousands of malware entries. But it also showed us that this is predictable. The time series models worked, and they gave us a way to forecast future trends.
Jane: Which means we can prepare. We can improve detection, we can push for curated lists and code signing, and we can use regulations like the CRA to force better hygiene.
Tom: And I think the most exciting implication is that this isn’t just an academic exercise. The authors are giving security teams, regulators, and even individual developers a tool to understand the risk.
Jane: Absolutely. And on a personal level, as a developer, this makes me want to double-check every single package I install. The risk is real, and it’s growing.
Tom: Well said, Jane. So, for Ruohonen and Saddiqa, thank you for this eye-opening research. We’ll be watching to see if these trends continue.
Jane: And that’s a wrap on this one. Stay tuned, because we’ve got another fascinating paper coming up next. Thanks for listening, everyone!
Tom: See you on the next episode!
Jukka Ruohonen, Mubashrah Saddiqa
University of Southern Denmark
cs.CR, cs.SE
Submitted: 2026-08-07
Comments: Proceedings of the 20th International Conference on Availability, Reliability and Security (ARES 2025) Workshops, Ghent, Springer, pp. 269-285
DOI: 10.1007/978-3-032-00633-2_16
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 58/100
The gist: The paper examines the previously overlooked longitudinal aspects of software ecosystem security, focusing on malware uploaded to six popular programming language ecosystems: CRAN, Go, Maven, npm,
Key concepts
- Programming Language Ecosystem
- These are official repositories for code blocks used by developers, such as npm (JavaScript) and PyPI (Python. These platforms allow millions of developers to acquire building blocks for modern applications.
- Time Series Analysis
- This method involves looking at data points collected over time. The paper uses it to track how the number of malware uploads has evolved in these ecosystems, allowing researchers to predict future trends.
- Typosquatting
- This is a malicious tactic where an attacker uploads a package with a name similar to a popular one, often by misspelling it (e.g, 'requestz' instead of 'requests'). Developers may accidentally install the wrong package.
Terminology
Summary
The paper examines the previously overlooked longitudinal aspects of software ecosystem security, focusing on malware uploaded to six popular programming language ecosystems: CRAN, Go, Maven, npm, PyPI, and RubyGems. The dataset examined is based on the new Open Source Vulnerabilities (OSV) database.
Research Questions:
-
RQ.1: How much detected malware uploads have popular programming language ecosystems seen compared to traditional vulnerability reports?
-
RQ.2: Which ecosystems have been particularly prone to malware uploads?
-
RQ.3: Can time series analysis provide insights into malware uploading trends?
Motivation: The paper notes that "most—if not all—malware recently discovered from the npm ecosystem come with the following warning: 'Any computer that has this package installed or running should be considered fully compromised. All secrets and keys stored on that computer should be rotated immediately from a different computer.' The paper also references the Cyber Resilience Act (CRA) recently enacted in the European Union, which
contains an obligation to only ship products without known vulnerabilities."
Data and Methods: The dataset was assembled from a bulk snapshot obtained in April 2025 from the OSV database. Five time series were constructed: MalFreqt (count of malware entries), MalSharet (percentage share of malware entries to all entries), Ecot (count of ecosystems contributing to malware entries), Advt (count of security advisories), and Artt (count of media articles and related information sources). These were operationalized into daily, weekly, and monthly aggregates with lengths T = 1195, T = 168, and T = 39, respectively, starting from January 2022. The paper uses an autoregressive distributed lag (ARDL) model with long-run multipliers and dynamic multipliers.
Key Results:
-
RQ.1 Answer:
malware uploads have surpassed the reporting of traditional software vulnerabilities in packages distributed in the ecosystems.
The paper states thatin the early 2025 even up to 80% of all entries in the OSV have been about malware.
-
RQ.2 Answer: "npm (over twenty thousand malware uploads), PyPI (nearly nine thousand malware uploads), and RubyGems (about eight hundred malware uploads) have been particularly prone to malware uploads, whereas CRAN, Go, and Maven have seen less than ten malware uploads in total." Specifically, the malware shares were: npm 82.46%, PyPI 56.29%, RubyGems 47.07%, Go 0.19%, Maven 0.02%, and CRAN 0.00%.
-
RQ.3 Answer: "time series analysis can reveal insights about malware uploads and their trends. The decent statistical performance obtained—the average coefficient of determination is 0.79—indicates that forecasting could be used also in practical foresight; a hypothesis is that the increasing trend of malware uploads continues also in the nearby future."
Regression Findings: All three explanatory time series indicate long-run effects, irrespective whether daily, weekly, or monthly aggregates are used.
The long-run multipliers show that a unit increase in Ecot increases MalSharet by 46.5 percentage points daily, 19.4 percentage points weekly, and 13.1 percentage points monthly.
The paper concludes that "the number of ecosystems provides particularly good predictive power; the more there are ecosystems, the more malware uploads are also reported. Smaller but still visible effects are present for security advisories and media and other articles; publicity seems to also affect the malware upload trends."
Limitations: The paper acknowledges that the OSV database has not been yet validated in research
and that only known and reported malware cases were observed,
meaning nothing can be deduced about unknown true positive cases. The paper also notes that an unconditional probability of picking a malware package is supposedly still tiny—even when dependencies are accounted for.
Implications: The paper suggests that improving monitoring and detection capabilities
may partially explain the results, and recommends curated lists for safe and secure packages
and code signing.
It also notes that somewhere in recent history this quality gating function was forgotten or overridden by the emergence of programming language software ecosystems.
Improvements for AI systems
Based on the paper, here are specific improvements I can make to AI systems, along with what the improved system can do:
Improvement: Implement an ARDL (Autoregressive Distributed Lag) model that uses ecosystem count, security advisory frequency, and media article frequency as predictors for malware upload rates.
What the improved system can do:
-
Forecast weekly malware upload counts and malware share percentages with 59–95% accuracy (R2 range reported)
-
Predict that a sustained increase of one additional ecosystem reporting malware leads to a 19.4 percentage-point increase in malware share (weekly aggregate)
-
Provide early-warning signals 2–4 weeks ahead of expected spikes, enabling proactive security scanning
Improvement: Build a risk-scoring engine that weights ecosystems based on their historical malware prevalence (npm: 82.5% of entries are malware, PyPI: 56.3%, RubyGems: 47.1%).
Improvement: Create a detection system that increases scanning intensity when security advisories or media articles spike, leveraging the finding that publicity correlates with (and possibly precedes) malware discovery.
Improvement: Integrate the finding that malware propagates through dependencies (not just direct installs) into a propagation-tracking module.
Improvement: Use the time series models to schedule coordinated security sweeps (analogous to CRA Article 60) at optimal intervals.
Improvement: Build a system that analyzes package name similarity and flags potential typo-squatting attempts, informed by the paper's finding that this remains a primary attack vector.
Improvement: Use the ARDL model's residual patterns (shown in Fig. 5) to detect unusual malware upload activity beyond what the predictors explain.
The improved AI system can:
-
Predict malware upload trends 1–4 weeks ahead with quantifiable accuracy
-
Prioritize security scanning across six ecosystems based on empirical risk
-
Trigger enhanced detection during publicity events
-
Trace malware propagation through dependency networks
-
Schedule sweeps at evidence-based intervals
-
Detect typo-squatting with contextual awareness
-
Alert on anomalies that deviate from established patterns
These improvements are directly grounded in the paper's empirical findings (Table 3 LRMs, Fig. 6 DMs, Table 1 ecosystem breakdowns) and address the practical gaps the authors identify, particularly the need for better monitoring, detection, and forecasting in programming language ecosystems.
Sources
- Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
- PsyScam: A Benchmark for Psychological Techniques in Real-World Scams
- Decoding Dependency Risks: A Quantitative Study of Vulnerabilities in the Maven Ecosystem
- A Mapping Analysis of Requirements Between the CRA and the GDPR
- The Popularity Hypothesis in Software Security: A Large-Scale Replication with PHP Packages
- Tracing Vulnerability Propagation Across Open Source Software Ecosystems
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs