A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds

arXiv:2609.03457 · cs.LG · Submitted 2026-09-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds".

Jane: The paper was written by Ashir Javeed, Anton Borg, Håkan Grahn, Lars Lundberg, Dhyey Patel et al. from Department of Computer Science, Blekinge Institute of Technology and Ericsson AB.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

The Core Mechanism: Tom: We’ve seen the initial concept and now we want to talk more about how the core mechanism of "A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds" works internally, Jane. How do these two stages actually interact?

Jane: The authors describe a specific architecture where they use a dedicated forecasting model to predict future customer service requests, which they quantify as Transactions Per Second (TPS), and then feeding that forecasted TPS into the second stage to estimate the subsequent CPU utilization.

Lu: It’s not just an input-output mapping; it’s this cascading relationship—a cascaded learning architecture that allows us to handle complex dynamics in a structured way, which is hard to achieve otherwise.

Meng: I'm interested in how the process accounts for the fact that CPU usage is a downstream consequence of service requests, making sure we aren't missing that dependency.

Lalam: This approach ensures we are anticipating needs before fulfilling them, Lalam. It’s like having a highly tuned predictive engine for our digital services.

Tom: The paper details this by establishing the first stage: predicting TPS using an expanding window, and the second stage: estimating CPU based on that prediction.

Jane: This is key because in dynamic cloud environments, you need to know what's coming—the demand—before you can even start figuring out the resource cost.

Meng: It’s a practical way to manage complexity; instead of trying to predict CPU using all the input variables simultaneously, they separate the concerns.

Lu: That separation allows us to handle those specific parts of causality that we might miss if we tried to model everything at once in a single layer.

Lalam: By focusing on demand first, it helps us build a more reliable and predictable foundation for our automated systems.

Tom: So, the core of this entire system is understanding the relationship between customer request volume and resource requirement before knowing the resource need itself.

Jane: It's all about building that causal link into making this approach so robust in dynamic cloud environments where things change every second.

Meng: It’s a functional solution that addresses a major pain point for any engineer trying to manage fluctuating cloud infrastructure.

Lu: Indeed, we are moving toward systems that understand the flow of data rather than just relying on patterns we've seen before.

Lalam: And with this integrated understanding, we can ensure our digital infrastructure supports the actual pace of human activity.

Technical Improvements: Tom: We’ve covered the core idea, so now let's talk about the technical improvements in "A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds," Jane. What makes this architecture technically superior to traditional methods?

Jane: The authors specifically chose XGBoost for both stages, which is a huge improvement over using computationally expensive deep learning models like LSTM or Transformer-based ones that require massive hardware.

Lu: It’s fantastic to see the combination of predictive power and computational efficiency; I can already picture how fast this running on thousands of servers simultaneously.

Meng: From an engineering standpoint, the fact that XGBoost is lightweight is a massive win for real-time streaming deployment; we don't need huge clusters just to run the predictor.

Lalam: This speaks directly to improving efficiency in our culture by enabling rapid, low-overhead decision-making in automated systems.

Tom: The paper also highlights how they managed concept drift using an expanding window, which is a big leap forward for continuous operation because the model stays adaptive.

Jane: It means the model can't get stuck predicting based on old patterns; it’s designed to keep learning from new data points as they arrive in the stream.

Meng: And I appreciate the integration of rolling maximum smoothing for the CPU target value—that helps filter out noise and find that peak demand, which is very practical.

Lu: It shows a deep understanding of the signal-to-noise ratio in real-world workload data, ensuring that we are predicting the peak moment rather than just averaging it out.

Lalam: By smoothing and adapting simultaneously, we ensure our digital infrastructure is stable and predictable for a more reliable user experience.

Tom: It looks like this is about making cloud resource management smarter, but also faster and easier to maintain for continuous deployment.

Jane: It’s all about achieving robust forecasting with controlled error accumulation over that sixty-step prediction horizon, which is crucial for long-term planning.

Meng: A manageable average runtime of around two hundred one milliseconds across the board is something we can actually integrate into live systems without any latency concerns.

Lu: That speed, combined with the structural integrity of the two stages in this paper, makes it a truly elegant solution for complex resource management problems.

Practical Implications and Outlook: Tom: We’ve seen so much about how it works, and now we want to talk about the practical implications of "A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds," Jane. What does this mean for cloud providers?

Jane: The main takeaway is that this two-stage approach offers a much more robust and efficient way to predict CPU needs than traditional methods, allowing us to anticipate demand rather than just reacting to it.

Lu: I’m excited about the potential, especially how we can use this kind of predictive power to drive innovation in cloud computing by fundamentally changing how we view resource allocation.

Meng: It's definitely practical; we have a tool that works fast and accurately enough for real-time auto-scaling decisions in production environments right now.

Lalam: Lalam believes this leads to a more conscientious relationship with technology, ensuring our digital footprint is both efficient and sustainable for the future.

Tom: So, as we wrap up the discussion on this study, what are your final thoughts on its real-world impact?

Jane: We've seen that the system achieves strong accuracy—a median SMAPE of five point nine percent—and it’s ready for real-time use in critical cloud applications.

Lu: And I think we can see further possibilities by integrating this into more complex, distributed systems, leveraging its predictive strength.

Meng: I just hope we can get to the point where this is used across cloud environments that are far more dynamic than our test data sets were, but it's a solid foundation.

Lalam: My final thought is that this enables a culture of proactive resource management where everyone benefits from foresight, not just from hindsight.

Tom: That’s a powerful way to conclude the discussion on "A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds." It's truly a clever piece of work.

Jane: Thank you all for this great conversation today, and we hope you enjoyed hearing about this impressive research!

Final Conclusion: Tom: Before we move on, I think it's worth taking one last look at the overall success of "A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds," Jane. How does this system measure up against the old methods?

Jane: It really is, Tom; by predicting customer demand first using TPS and then calculating CPU needs, they’ve made the whole process much more intuitive for real-time systems compared to direct estimation.

Lu: I find the concept of cascading predictions so exciting because it feels like we're finally moving beyond just a correlation and into a genuine understanding how these complex workloads operate.

Meng: From an engineering viewpoint, I think the practical implication is that this kind of low-latency, two-stage approach allows us to deploy auto-scaling systems without the massive overhead that traditional deep learning models demand.

Lalam: Lalam sees this as a vital step toward promoting a more sustainable and thoughtful use of cloud resources by aligning our infrastructure with actual human activity.

Tom: Exactly, Lalam, it's about being proactive rather than reactive, which is such a huge shift in mindset for resource provisioning.

Jane: And we've seen that the accuracy—a median SMAPE of five point nine percent—is solid enough to provide confidence in this framework for critical cloud applications.

Lu: I can't wait to see how researchers take this further, perhaps using these insights to train adaptive systems across different types of hardware configurations.

Meng: The real-world deployment capability is what makes this so important; it’s a practical solution that works within the constraints of time and resources.

Lalam: It's a model that understands the pulse, and we need models like this to improve our collective relationship with technology.

Tom: We hope to see how this impacts future cloud management systems, especially when we talk about scaling these solutions up globally.

Department of Computer Science, Blekinge Institute of Technology · Ericsson AB

cs.LG

Submitted: 2026-09-03

Updated: 2026-09-03

Importance score: 90/100

The gist: I apologize, but you have only provided a list of references (citations [12] through [36]) and not the actual content or body text of the arXiv paper titled "A Two-Stage Forecasting System for CPU

Key concepts

Two-Stage Forecasting System
This architecture separates the prediction process into two stages to handle complex causality. The first stage predicts customer demand (TPS), and the second stage uses that forecasted demand to calculate the resulting CPU utilization, ensuring a structured approach.
Transactions Per Second (TPS) Prediction
TPS is used as a proxy for future customer service request volume. The system forecasts this demand using an expanding window technique, establishing the initial causal link between incoming requests and necessary resource allocation.
Expanding Window and Concept Drift
The expanding window allows the model to continuously learn from new data points as they arrive in a stream. This method prevents the model from becoming outdated (concept drift), ensuring it remains adaptive for continuous operation in dynamic cloud environments.
XGBoost Efficiency
The system utilizes XGBoost, which is noted for its predictive power and computational efficiency. This lightweight choice avoids the need for massive hardware clusters required by deep learning models, making it ideal for real-time streaming deployment.

Terminology

Summary

I apologize, but you have only provided a list of references (citations [12] through [36]) and not the actual content or body text of the arXiv paper titled A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds.

To perform the detailed extraction and summary you require—adhering to the strict structural constraints, word count, and direct quotation rules—I need the full manuscript text.

Please provide the complete PDF or text content of the paper. Once I have access to the material, I will immediately generate a summary that is:

  1. Structured with one orienting paragraph followed by 3–5 bolded sections (e.g., "Methodology").

  2. Approximately 450–600 words long, using detailed quotes and lists as appropriate.

  3. Completely objective, containing no commentary or external analysis, starting directly with the paper's substance.

Improvements for AI systems

(Internal Monologue: The bibliography points to a strong area of research—resource prediction for cloud optimization. However, many papers treat prediction and optimization as separate stages. To prevent catastrophic failures in a real-world system, I must enforce mathematical rigor, uncertainty quantification, and true architectural integration. The improvements must elevate the system from merely predictive to proactively guaranteed.)


The current research primarily focuses on prediction using various time-series models (LSTM, ARIMA, XGBoost). The critical flaw is the lack of mandatory uncertainty quantification and the failure to integrate prediction directly into a real-time constrained optimization loop.

I propose upgrading the system architecture in three mandatory phases: Predictive Core Enhancement, Robustness Layer Implementation, and Actionable Optimization Integration.

Improvement: Implement a Hierarchical Ensemble Transformer Architecture for workload forecasting that explicitly separates trend components, seasonal cycles, and high-dimensional feature interactions. This moves beyond simple concatenation of models.

  • Mechanism: The system must use a modular design:
  1. A classical statistical module (e.g., SARIMA) to capture stable seasonality/periodicity (low variance component).

  2. A Transformer/LSTM module trained on exogenous variables and complex, non-linear interactions (high variance component).

  3. The two streams are combined using a weighted, attention-based mechanism that dynamically assigns trust scores to each stream based on the observed data volatility.

What the Improved System Can Do:

  • Predict resource demand (CPU t+k, Memory t+k) with significantly higher accuracy and robustness than single-model approaches, especially during periods of concept drift or sudden load spikes (e.g., identifying a novel user behavior pattern).

  • Crucially, it can predict the causal relationship between diverse inputs (e.g., a specific API call volume correlated with peak CPU usage) rather than just correlating time steps.

  • Mechanism: Instead of outputting CPU t+k = 60 units, the system must output a predicted interval: CPU t+k about [45, 75] with a 95% confidence level. The width of this interval becomes the primary metric for risk assessment.

  • Mechanism: The RL Agent receives three inputs:

  1. Predicted Load (mu).

  2. Uncertainty Bounds ([L, U]).

  3. Operational Constraints (Budget B, SLA threshold S).

  • The MIP solver then solves for the optimal resource allocation vector X (e.g., number of instances, CPU limits) that minimizes a defined cost function C:

Minimize C(X) = alpha times (Cost(X)) - beta times (SLA Compliance Score) + gamma times (Resource Waste Penalty)

Related papers