User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

arXiv:2608.11840 · cs.DC, cs.AI, cs.LG, cs.NI, cs.PF · Submitted 2026-08-12 · Read on arXiv

Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta

Stockholm University

cs.DC, cs.AI, cs.LG, cs.NI, cs.PF

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper proposes a collaborative distributed inference system that combines dedicated server infrastructure with resources contributed by service users' devices.

Terminology

Summary

The paper proposes a collaborative distributed inference system that combines dedicated server infrastructure with resources contributed by service users' devices. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. The system partitions inference computations between dedicated resources, such as centralized servers, and volunteered resources, such as users' devices. Unlike centralized cloud inference or approaches distributed exclusively across dedicated computing continuum infrastructure, the system does not require dedicated computational capacity to scale in proportion to demand. In contrast to conventional volunteer computing, it retains dedicated infrastructure to provide a reliable baseline capacity and support QoS guarantees despite the intermittent availability and heterogeneity of volunteered resources.

To capture stochastic and dynamic interactions among users, resources, tasks, and policies, the authors develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. The system state is defined as a tuple of high-dimensional variables: st = ⟨ot, qt, at, xt, yt, ct, dt, ut⟩, where each variable captures a specific aspect of the system state. The model includes availability states (ot), inference requests (qt), allocated resources (at), subtask readiness states (xt), subtask execution states (yt), resource consumption (ct), and decision-making policies (ut). The model uses parametric survival distributions for availability durations, a Pareto survival function for inter-request times and patience, generative ARMA models for resource allocations and consumptions, and discrete-time hazard models for subtask phase completion.

The system is controlled by two decision-making policies: a server-side scheduler policy π and a node-side executor policy ς. The executor policy enforces local feasibility by comparing total consumption with allocated resources, pausing or aborting downloads and executions under overflow conditions. Three scheduler policies are evaluated: a centralized scheduler that runs all executions on the server, a uniform scheduler that uses online volunteered nodes opportunistically and assigns subtasks uniformly at random among feasible nodes, and an affinity-based scheduler that prioritizes nodes by the number of active executions of the same subtask to improve batching.

The evaluation simulates trajectories of the generative model with 100,000 time steps (approximately 28 hours), varying user count 100, 1,000, 5,000, 10,000, user capacity 1×, 2× of base resource profile, server capacity 10×, 30×, 50×, 80×, 100×, and scheduler policy. The base resource profile is defined as 100 Mbit inbound/outbound bandwidth, 3 GB storage, 2.5 GB memory, and 200% CPU. Four task types are considered: light and heavy, each with 5 or 10 subtasks, sampled with probabilities 0.4, 0.3, 0.2, 0.1.

Results show that the centralized policy is unaffected by user capacity because all subtasks run on the server. Under the centralized policy, at 10× capacity, completions grow only through 1,000 users, then plateau and decline as cancellations surge. Higher capacities delay saturation, but gains diminish beyond 30×; at large user counts, cancellations still equal roughly one third to one half of completions. The uniform policy matches the centralized through 1,000 users and improves with 2× user capacity. At 5,000 users, it performs better at 10× and 30× server capacity and comparably above them. At 10,000, it yields more completions and fewer cancellations at every server capacity, including 100×. The affinity-based policy is competitive through 1,000 users and beats the centralized policy at 10× capacity, but otherwise generally yields the worst completion-to-cancellation balance.

Regarding P99 latency, the centralized policy performs best or comparably through 1,000 users, and at 5,000 as server capacity grows. At 10,000, the uniform policy remains substantially better. The uniform policy matches the centralized policy through 1,000 users, outperforms it at 5,000 with up to 30× server capacity, and remains comparable above that. At 10,000, it consistently achieves much lower P99 latency. Both distributed policies show little improvement beyond 30×–50× server capacity, indicating this range provides sufficient dedicated capacity to complement volunteered resources.

Regarding dedicated resource consumption, the centralized policy consistently uses more dedicated resources. CPU consumption reaches every allocation limit, indicating persistent saturation. Both distributed policies use substantially fewer dedicated resources and never exceed their allocations. Consumption changes little between 30× and 100×, indicating that approximately 30×–50× capacity sufficiently complements volunteered resources. The affinity-based policy is slightly more efficient than the uniform policy.

Regarding subtask completion rates, user-side rates under distributed policies are largely independent of server capacity. The uniform policy achieves approximately 0.75–0.95, improving with user capacity. Centralized server-side rates at 10× capacity fall from 0.9–1.0 for up to 1,000 users to approximately 0.2 at 10,000. Distributed policies reach near-perfect server-side rates with considerably less capacity, particularly for 2× users.

The paper concludes that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling. The paper notes that this work does not focus on designing a scheduling algorithm or resource-allocation optimizer; instead, the primary focus is to define the collaborative distributed inference system and develop a model of its dynamics. Future work will use this model to address optimization problems, including optimizing the scheduler policy to maximize QoS and determining the amount of dedicated capacity required to satisfy target QoS constraints.

Improvements for AI systems

Improvements to AI Systems:

  1. Hybrid Resource-Aware Inference Scheduler
  • Integrate a dual-policy scheduler (server-side and node-side) that dynamically partitions inference subtasks between dedicated cloud servers and volunteered user devices based on real-time availability, resource consumption, and QoS targets.

  • Use the generative Markov model (with survival, Pareto, ARMA, and hazard components) as a predictive digital twin to forecast user churn, request bursts, and node capacity fluctuations.

  • Capability: The improved system can autonomously scale inference capacity without proportional cloud infrastructure growth, maintaining P99 latency and completion rates under user surges (e.g., 10,000+ concurrent users) while cutting dedicated CPU/memory usage by up to 70% compared to centralized-only inference.

  1. QoS-Aware Autoscaling Controller
  • Replace static server capacity rules with a model-predictive controller that uses the paper’s findings (diminishing returns beyond 30×–50× server capacity) to set optimal dedicated resource baselines.

  • Incorporate the affinity-based batching heuristic (prioritizing nodes with active same-subtask executions) to improve GPU/CPU utilization on volunteered devices.

  • Capability: The system can preemptively adjust dedicated capacity in response to predicted request inter-arrival times (Pareto-distributed) and user patience, reducing cancellation rates by 30–50% during peak load while avoiding over-provisioning.

  1. Fault-Tolerant Distributed Inference Executor
  • Implement the node-side executor policy (ς) that enforces local feasibility by pausing/aborting subtasks under resource overflow, but enhance it with a fallback mechanism: if a volunteered node fails mid-execution, the subtask is re-routed to the nearest available dedicated server or another volunteer with similar affinity.

  • Use the discrete-time hazard models to predict subtask phase completion times, enabling proactive checkpointing and migration.

  • Capability: The improved system can maintain near-perfect server-side completion rates (0.9–1.0) even with 10,000 users and intermittent volunteer availability, while keeping user-side completion rates above 0.75 without dedicated capacity scaling.

  1. Generative Simulation-Based Policy Optimizer
  • Train a reinforcement learning agent on the high-dimensional Markov model (st = ⟨ot, qt, at, xt, yt, ct, dt, ut⟩) to learn optimal scheduling policies under varying user counts, capacities, and task mixes (light/heavy, 5/10 subtasks).

  • Use the simulation to generate synthetic trajectories for offline policy evaluation, avoiding costly real-world A/B testing.

  • Capability: The system can discover novel scheduling strategies that outperform the uniform and affinity-based baselines, achieving higher completion-to-cancellation ratios and lower P99 latency at 10,000 users, while automatically determining the minimum dedicated capacity (e.g., 30×) needed to meet target QoS constraints.

  1. Adaptive Task Partitioning Engine
  • Use the model’s resource consumption (ct) and allocation (at) variables to dynamically split inference tasks into subtasks (5 vs. 10) based on current node heterogeneity and network bandwidth (e.g., 100 Mbit baseline).

  • Optimize subtask granularity to balance communication overhead vs. batching efficiency, using the affinity heuristic to group similar subtasks on the same volunteer device.

  • Capability: The improved system can reduce end-to-end inference latency by 20–40% for heavy tasks under mixed user capacities, by choosing optimal subtask counts and node assignments that minimize idle time and data transfer bottlenecks.

  1. Real-Time QoS Monitoring and Alerting
  • Embed the generative model’s survival and hazard functions into a live dashboard that predicts imminent QoS violations (e.g., P99 latency exceeding thresholds) based on current availability states (ot) and request queues (qt).

  • Trigger preemptive actions—such as increasing server capacity or rebalancing subtasks—before cancellations surge.

  • Capability: The system can maintain QoS guarantees (e.g., <500ms P99) during flash crowds, automatically shifting load to dedicated resources only when volunteer capacity is predicted to fall below demand, reducing cloud costs by up to 50% during normal operation.

Abstract

Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. To capture stochastic and dynamic interactions among users, resources, tasks, and policies, we develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. We evaluate the system across user populations, resource capacities, and centralized and distributed scheduling policies. Simulations show that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling.

Sources

Related papers