Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

arXiv:2608.12123 · cs.DC, cs.AI, cs.OS · Submitted 2026-08-12 · Read on arXiv

cs.DC, cs.AI, cs.OS

Submitted: 2026-08-12

Updated: 2026-09-12

Comments: 14 pages, 4 figures. Includes formal proofs, trace provenance, and a reproducibility appendix. Code and artifacts: https://github.com/josefchen/ready-cohorts ; processed evidence: https://huggingface.co/datasets/josefchen/ready-

Code: https://github.com/josefchen/ready-cohorts

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control, by Josef Chen, addresses the question of when deterministic control transitions between model and tool

Terminology

Summary

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control, by Josef Chen, addresses the question of when deterministic control transitions between model and tool calls in LLM-agent services can be grouped and executed profitably on a GPU. The paper formalizes the ready-cohort boundary using several quantities: the hardware threshold K (the start of a measured safe suffix where every tested cohort size beats a baseline), the fixed-partition share F (eligibility under a frozen time partition), the exact offline share P⋆ (maximum schedulable share with future knowledge), the local upper bound U (necessary-overlap bound), and the online achieved share A (achieved accelerated share by a finite online runtime). The paper states: For zero service time, unlimited capacity, and equal relative launch deadlines, a specialized dynamic program computes P ⋆ exactly.

In the trace study, using a prospectively frozen stationary Poisson replay of one pinned 851-session public trace panel, the primary condition at 100,000 target active sessions, K = 256, and a 50 ms launch deadline gives F = 30.19%, P ⋆ = 43.00%, and U = 45.85%. The paper notes: Exact packing recovers 81.83% of the opportunity lost at fixed window boundaries. It also emphasizes that The outcome-derived route key is a conditioning proxy, not proof of executable identity. The trace study also finds that cohort supply collapses below the boundary: Under route-key grouping at K = 256, P ⋆ is zero for every tested deadline when C ≤ 10,000. Even at C = 100,000, it remains zero at 10 and 25 ms.

In a separate mechanism study, the paper keeps a GPU-computed binary decision on device instead of returning four bytes to the host and redispatching. Across four named GPU placements (local GTX 1660 Ti, Modal L4, RunPod L4, Lambda H100 SXM5), the device-resident path is faster in all 36 configurations; within-placement row-median ratios range from 1.19× to 2.39×. At the primary cell (N = 256, H = 32), ratios are 1.71×, 2.39×, 2.06×, and 1.84× respectively. The paper also reports: Across both admissible mechanisms, all 14,557,440 tested batched invocations match a separately implemented host oracle. A negative control with a fixed nested device graph that removes no host decision is slower in all 60 configurations across five placements, ruling out device launch alone as the explanation.

The paper concludes: "Together, the studies establish two measurable gates for GPU agent control: deadline-feasible cohort supply and observation placement. They expose schedulable work missed by fixed windows. A joined finite online runtime is required to measure A, CPU displacement, and service-level benefit. The proposed next step is a finite-capacity online route compactor that receives typed completion events, updates compact state, queues events by a verified executable route and deadline, launches a route only above its measured safe-suffix threshold Kr, and falls back to a tuned CPU path when a deadline or capacity limit prevents batching." The paper defines achieved share A as A = (1/E) Σ zi, where zi = 1 only when the frozen online policy accelerates event i with an exact output and observed device start no later than its launch deadline, and defines opportunity recovery RA = (A − F)/(P ⋆ − F). The paper explicitly warns that The current numbers do not form a service-level acceleration estimate and must not be multiplied, because the trace threshold is a swept candidate rather than a measured crossover, and the mechanism horizon H = 32 is absent from the trace model.

Improvements for AI systems

Improvements to AI Systems:

  1. GPU-Resident Decision Execution: Modify LLM-agent runtimes to keep binary control decisions (e.g., tool-call eligibility, route selection) on the GPU, avoiding host round-trips. This yields 1.19×–2.39× faster per-decision latency across diverse hardware (GTX 1660 Ti, L4, H100), with verified output correctness against host oracles.

  2. Deadline-Aware Cohort Batching: Implement a dynamic-programming-based scheduler that groups tool/model calls into “ready cohorts” based on launch deadlines and GPU capacity, replacing fixed time-window batching. This recovers up to 81.83% of lost scheduling opportunity (F=30.19% → P⋆=43.00% at 100k sessions, K=256, 50ms deadline).

  3. Adaptive Safe-Suffix Thresholding: Use per-route, measured safe-suffix thresholds (K r) to decide when to launch a GPU batch, rather than a global static batch size. This prevents cohort supply collapse (e.g., P⋆=0 for C≤10,000) by dynamically falling back to CPU when demand is too low or deadlines too tight.

  4. Finite-Capacity Online Route Compactor: Build a runtime that receives typed completion events, maintains compact state, queues events by verified executable route and deadline, and launches only above K r. This enables real-time measurement of achieved share (A) and opportunity recovery (R A) without assuming infinite capacity or zero service time.

  5. Conditioning-Proxy-Aware Routing: Treat outcome-derived route keys (e.g., predicted tool identity) as a proxy, not proof of executable identity. The improved system must verify executable identity on-device before batching, preventing incorrect cohort grouping and ensuring exact outputs.

  6. Cautious Multiplier-Free Acceleration Metrics: Avoid conflating trace-level scheduling gains with mechanism-level latency gains. The improved system reports A and R A separately, and only claims service-level acceleration after a joined runtime measures both under real capacity and deadline constraints.

What the Improved AI System Can Do:

  • Execute LLM-agent control loops with sub-millisecond GPU-resident decisions, eliminating host communication bottlenecks for binary choices.

  • Schedule tool calls and model invocations into deadline-feasible cohorts that exploit GPU parallelism beyond fixed windows, recovering significant lost throughput at scale (e.g., 100k concurrent sessions).

  • Gracefully degrade to CPU execution when cohort supply is insufficient (low session counts or tight deadlines), maintaining correctness and latency bounds.

  • Provide real-time, measurable guarantees on accelerated share (A) and opportunity recovery (R A) under finite capacity, enabling safe deployment without overstating benefits.

  • Handle heterogeneous hardware (from local GPUs to cloud H100s) with consistent speedups, verified against host oracles for output exactness.

Abstract

LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on device. We formalize the ready-cohort boundary using fixed-partition share F, exact offline share P*, local upper bound U, and online achieved share A. Under zero service time, unlimited capacity, and equal relative launch deadlines, a specialized dynamic program computes P* exactly. In a stationary Poisson replay of one pinned 851-session public trace panel, the primary condition at 100,000 target active sessions, K=256, and a 50 ms launch deadline gives F=30.19%, P*=43.00%, and U=45.85%. Exact packing recovers 81.83% of the opportunity lost at fixed window boundaries. The outcome-derived route key is a conditioning proxy, not proof of executable identity. A separate mechanism study keeps a GPU-computed binary decision on device instead of returning four bytes to the host and redispatching. Across four named GPU placements, the device-resident path is faster in all 36 configurations; within-placement row-median ratios range from 1.19x to 2.39x. Across both admissible mechanisms, all 14,557,440 tested batched invocations match a separately implemented host oracle. A fixed nested device graph that removes no host decision is slower in all 60 configurations across five placements. Together, the studies establish two measurable gates for GPU agent control: deadline-feasible cohort supply and observation placement. A joined finite online runtime is required to measure A, CPU displacement, and service-level benefit.

Sources

Related papers