Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing

arXiv:2610.00063 · cs.DC, cs.AI, cs.CL · Submitted 2026-09-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Pushing CPU Speech Synthesis to the Wall".

Jane: Instance-billed serverless platforms charge for CPU and memory held across an instance's lifetime, making billing-aware serving crucial for optimizing neural text-to-speech inference costs.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's look at the title again, "Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing." It sounds intense, like they’re pushing the limits of what these systems can handle before you start incurring massive costs.

Jane: It certainly suggests they're looking at an extreme tuning scenario where performance meets a strict billing constraint on serverless CPUs. The authors are trying to find the right balance between getting fast results and not overspending on idle time.

Lu: It’s fascinating that they tie the tuning directly to the billing architecture; it shows that cost isn't just an afterthought but a primary constraint in designing inference pipelines.

Meng: So, when you look at these authors, it seems like they are deeply embedded in both the hardware constraints of serverless environments and the specific resource consumption patterns of neural text-to-speech models.

Lalam: It suggests a very mature understanding of the operational costs associated with deploying large language models in production settings, focusing on efficiency beyond just model accuracy.

The paper's summary: Tom: In terms of the summary, they pinpoint two major issues that existing inference runtimes miss: first, how machine-wide parallelism causes CPU contention when requests run at the same time, and second, how warm instances keep large chunks of billable memory around even when there’s no traffic.

Jane: That's a clear way to frame it for us—the problems are oversubscribing the CPU with concurrent requests and leaving gigabytes of billable state sitting idle, which directly impacts the cost structure.

Lu: The paper proposes two mechanisms to fix this: request-sized concurrent inference to manage parallelism per request, and a reclaimable instance lifecycle to handle that idle memory.

Meng: So, the summary boils down to solving the busy time efficiency problem with one technique and then tackling the idle time cost component with another mechanism.

Lalam: That dual approach seems smart; it addresses both when things are busy and when things are quiet, which is crucial for maintaining a healthy operational budget.

The paper's improvements: Tom: They go into detail about the specific fixes, like using request-sized concurrent inference to bound per-request CPU parallelism by serving requests through a fixed pool of inference streams, maybe eight streams with two threads each on a thirty-two-vCPU container.

Jane: And they pair that up with the reclaimable instance lifecycle, which essentially lets the server release inference state and billable page cache memory after a quiet period while keeping the process ready for a quick restoration.

Lu: I see how this directly tackles the problem of idle inference state becoming a direct serving cost, which is what they call "idle inference state into a direct serving cost". It’s about explicitly managing the CPU-seconds and GB-seconds accumulated over the instance lifetime.

Meng: From an engineering perspective, getting that reclaimable lifecycle to work—releasing state and returning memory while keeping the compile cache for faster restoration—that sounds like a complex state management challenge.

Lalam: This mechanism could have major implications for how we design our core AI services; if we can effectively manage that released state, it means our systems become much more cost-effective over a long period.

Conclusion: Tom: To wrap things up, the paper shows that when you apply these two mechanisms—request-sized concurrent inference and the reclaimable instance lifecycle—the results are substantial, reducing the cost per audio-hour by four point one times compared to PyTorch defaults and significantly cutting idle billed memory.

Jane: It really demonstrates that this restructuring around the billed instance lifetime is what converts gains in busy-time efficiency into lower overall serverless costs, especially when dealing with bursty traffic where idle memory is a big part of the bill.

Lu: The conclusion that "further improvement is determined by the model rather than by the serving path" suggests that while this serving architecture helps a lot, optimizing the underlying text-to-speech model itself will still be necessary for further gains.

Meng: I'm glad they showed that their design conclusions transfer across different CPU architectures and even to an independently developed TTS system, which gives us confidence in its general applicability.

Lalam: It’s exciting to think about the cultural impact; if we can make these essential services significantly more cost-effective, it frees up resources that could be directed toward developing more complex and nuanced AI applications.

Pakorn Nathong, Kunat Pipatanakul

Paxa Labs

cs.DC, cs.AI, cs.CL

Submitted: 2026-09-04

Updated: 2026-09-04

Comments: 6 pages, technical report

Code: https://github.com/paxalabs/cpu-tts-to-the-wall

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: Instance-billed serverless platforms charge for CPU and memory held across an instance's lifetime, making billing-aware serving crucial for optimizing neural text-to-speech inference costs.

Key concepts

Billing-Aware Serving
This approach focuses optimization on minimizing the total CPU-seconds and GB-seconds accumulated over an instance's entire lifetime, rather than just optimizing for speed or throughput in isolation. It directly addresses how serverless platforms charge for held resources.
Request-Sized Concurrent Inference
Instead of allowing all incoming requests to run simultaneously, this technique limits the number of concurrent CPU threads per request. This ensures that independent utterances share a fixed execution budget efficiently, preventing them from competing through large machine-wide thread pools and reducing CPU contention.
Reclaimable Instance Lifecycle
This mechanism creates a state between being warm and being completely cold. When idle, the server releases billable inference state and page-cache memory while keeping the process ready for fast restoration. This prevents paying for unused memory during quiet periods.

Terminology

Summary

Instance-billed serverless platforms charge for CPU and memory held across an instance's lifetime, making billing-aware serving crucial for optimizing neural text-to-speech inference costs. The core finding is that organizing serving around the billed instance lifetime rather than always-on models, through request-sized concurrent inference and a reclaimable instance lifecycle, significantly reduces serverless costs by addressing both busy and idle components of the bill.

The Core Problem

Existing inference runtimes are poorly matched to billing-aware serverless settings because they fail to account for the cost structure where idle inference state into a direct serving cost. The optimization objective must shift from throughput or latency in isolation to minimizing the CPU-seconds and GB-seconds accumulated over an instance lifetime. This mismatch manifests in two structural problems: first, machine-wide per-request parallelism causes CPU contention when independent utterances execute concurrently, and second, warm instances retain gigabytes of billable inference state and page-cache memory during idle periods. Scaling to zero removes the idle charge but introduces a costly multi-second process cold start.

The Proposed Solution: Two Mechanisms

The paper addresses these inefficiencies with two primary mechanisms:

  1. Request-sized concurrent inference: This bounds per-request CPU parallelism so that independent requests efficiently share a fixed execution budget. This prevents concurrent requests from competing through overlapping machine-wide thread pools.

  2. A reclaimable instance lifecycle: This adds a released state between warm-idle and process cold, where after a quiet period, the server relinquishes inference state and billable page-cache memory while retaining the server process and compile cache for faster restoration.

Performance Results

On the Kokoro-82M model, these changes yielded substantial improvements. The resulting system sustained 2.71 audioseconds per CPU-second, compared to 0.89 for ONNX Runtime defaults, and reduced cost per audio-hour from a PyTorch default of 0.0631 to 0.0153—a 4.1× reduction relative to PyTorch. Furthermore, it reduces idle billed memory from 8.7 GB for PyTorch to 1.33 GB, a 6.5× reduction, and restores a released instance in 2.2 s, against 7.7 s for a PyTorch process cold start.

Cost Analysis Under Bursty Traffic

The paper demonstrates that while request sizing alone improves busy-time efficiency, it leaves a large warm-idle footprint. The reclaimable lifecycle is what converts busy-time inference efficiency gain into lower serverless cost. Figure 1 illustrates this: request sizing leaves a large warm-idle footprint, increasing the daily bill to 3.24 when compared to PyTorch's 2.41, whereas adding the lifecycle reduces the complete-system bill to 0.38.

Optimization Headroom and Transferability

After addressing these major serving-path inefficiencies, profiling suggests limited end-to-end headroom from these techniques for this model and hardware. Testing static-shape compilation or padding without masking failed to improve performance beyond repeat-run variability. However, the design conclusions transfer across architectures; on the Zen 5 host, the system achieved 3.26 audio-seconds per CPU-second against PyTorch defaults, and the benefits of request sized concurrent inference and the reclaimable instance lifecycle remain. The same lifecycle structure was shown to be motivating for an independently developed model (pocket-tts), confirming its general applicability.

Conclusion

Serverless CPU platforms bill an integral that current inference frameworks were not designed to minimize. Restructuring the serving path around this integral, with request-sized concurrent inference and a reclaimable instance lifecycle that releases billed memory including page cache and restores from compile caches in seconds, reduces the cost of StyleTTS2-family TTS by 4.1× per audio-hour relative to PyTorch defaults, and by far more on bursty traffic where warm-idle memory dominates the bill. The final conclusion is that further improvement is determined by the model rather than by the serving path.

The gist: Serverless CPU inference should be organized around the billed instance lifetime rather than the always-on execution model assumed by conventional inference runtimes.

How it works

  1. Request-sized concurrent inference: This mechanism bounds per-request CPU parallelism using a fixed pool of inference streams (e.g., eight streams with two threads each on a 32-vCPU container). This ensures that independent requests efficiently share a fixed execution budget rather than competing through machine-wide thread pools, leading to improved busy efficiency and reduced median first-audio latency.

  2. Reclaimable instance lifecycle: This introduces a state between warm-idle and process cold.

Improvements for AI systems

Based on the provided research, here are specific improvements that can be made to AI speech synthesis systems, along with what those improved systems can achieve:

  1. Improvements in Serverless CPU Inference Management: Implement a serving architecture that explicitly targets billing-aware objectives by managing inference state based on the billed instance lifetime rather than throughput or latency in isolation.

  2. Implementation of Request-Sized Concurrent Inference: Design the system to bound per-request CPU parallelism using a fixed pool of inference streams (e.g., 8 streams with two threads each on a 32-vCPU container) instead of giving each request machine-wide parallelism, ensuring independent requests share a fixed execution budget efficiently.

  3. Implementation of Reclaimable Instance Lifecycle: Introduce a mechanism where the server relinquishes billable inference state (including allocator memory and page cache files) after a quiet period (e.g., 45 seconds), while retaining only the process and compile cache for faster restoration upon the next request.

These improvements result in an improved AI system that can achieve the following:

  1. Significant Cost Reduction: Reduce the cost per audio-hour by up to 4.1 times compared to PyTorch defaults and 3.1 times compared to ONNX Runtime defaults, by minimizing idle billed memory (reducing warm-idle memory from 8.7 GB down to 1.3 GB).

  2. Improved Efficiency Under Bursty Traffic: Convert busy-time inference efficiency gains into lower overall serverless cost by eliminating the dominant cost component of idle memory during bursty traffic, where request sizing alone is insufficient.

  3. Faster Recovery for Interactive Use Cases: Restore an instance to first audio in 2.2 seconds after a quiet period, which is significantly faster than the 7.7-second cold start required by standard scaling-to-zero models, making them suitable for interactive applications despite scaling to zero capabilities.

  4. Maintained or Improved Latency: Achieve sustained efficiency of up to 2.71 audio-seconds per CPU-second while keeping first-audio latency comparable to or better than existing defaults (e.g., p50 latency of 0.94s for the complete system vs 0.69s for sized OpenVINO).

Abstract

Instance-billed serverless platforms charge for CPU and memory over the lifetime of a warm instance, making idle inference state a direct serving cost. We present billing-aware neural text-to-speech (TTS) serving on serverless CPUs, optimizing CPU-seconds and GB-seconds rather than throughput or latency alone. Conventional runtimes are poorly suited to this setting: per-request parallelism causes CPU contention under concurrency, while warm instances retain gigabytes of billable inference and page-cache state. We address these costs with request-sized concurrent inference, which bounds per-request CPU parallelism, and a reclaimable instance lifecycle, which releases inference state and page-cache memory after idle periods while retaining the server process and compile cache. On Kokoro-82M, our system achieves 2.71 audio-seconds per CPU-second versus 0.89 with ONNX Runtime defaults and reduces cost per audio-hour from 0.0631 with PyTorch to 0.0153, a 4.1x reduction. Idle billed memory falls from 8.7 GB to 1.33 GB, while restoration reaches first audio in 2.2 s versus 7.7 s for a PyTorch cold start. Under bursty traffic, lifecycle reclamation is essential for translating inference efficiency into lower serverless cost.

Sources

Related papers