RW-TTT: Batched Serving for Request-Owned Test-Time Training State
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "RW-TTT: Batched Serving for Request-Owned Test-Time Training State".
Jane: The paper was written by Jian Yang, Zhizhuo Kou, Yao Tian, Hao Zhang, Han Chen et al. from Hong Kong University of Science and Technology (HKUST) and Chinese University of Hong Kong (CUHK) and National University of Singapore (NUS).
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary: Jane: The summary highlights that we have to manage this request-owned mutable state—it needs to be read and updated during decoding—so we can't just treat it like a static parameter, which is the core problem.
Tom: It’s about managing that complexity, so how did they solve the issue of having an update happen while other requests are still reading old data?
Lu: They use this concept of "read-write TTT serving," which is essentially defining a contract for state management—it's not just any random update; it's a formalized process.
Meng: The summary emphasizes that by using an owner map, even if multiple requests are running at once, the updates only go into the correct state slot, which prevents race conditions from corrupting data.
Jane: That’s right; this is how they handle the complexity—they're ensuring every request is accountable to its own versioned history so that a READ never sees a half-written WRITE.
Lalam: It’s about building a digital infrastructure where personalized intelligence doesn't mean performance degradation; we can have unique, evolving intelligence for each user while keeping the lights on.
Tom: That clarity is the key, but understanding how to handle failure and edge cases is just as important as defining the summary, leading into the technical improvements of this groundbreaking work.
Improvements: Tom: Now we're looking at what's actually in the code; the authors of RW-TTT have introduced a sophisticated batching planner that is completely unique and necessary for this system.
Jane: This planner doesn't just group requests based on how much they need to read; it groups them based on whether they are performing compatible READ or WRITE operations, which is a massive improvement over naive batching.
Meng: It’s about compatibility—if the state types and the backend operators match, we can run them together in a single batch, which is exactly what unlocks that massive throughput gain we saw in the results.
Lu: And this planner also incorporates a bounded waiting mechanism; if all compatible requests aren't ready right on cue, it waits for a short period before moving to the next step.
Jane: That waiting period is key because it allows us to gather more requests into those optimal batches without letting the latency get too high for the end user.
Lalam: It’s a calculated patience that ensures we are maximizing the potential of our hardware while still delivering a smooth, fast experience for every single user.
Tom: The authors also emphasize that RW-TTT doesn't just about batching; they’re showing how to handle failure and rollback safely—this is huge for reliability in production environments.
Meng: When a write operation fails, we have to know exactly which state version is the last one that was actually committed so we can fall back correctly without data corruption.
Lu: That rollback-safe commit mechanism ensures that the integrity of the adapted model's behavior is maintained even if things get chaotic during generation.
Jane: It’s a safety net that allows us to push boundaries with TTT while guaranteeing consistency when things go wrong, which is a huge win for confidence.
Tom: We are talking about a system where two hundred seventy-four point six one tokens per second is achievable, which is a monumental achievement for the serving layer in this research "RW-TTT: Batched Serving for Request-Owned Test-Time Training State."
Lalam: That speed allows us to think bigger, enabling cultural and creative applications that were simply too slow to be efficient before.
Technical Deep Dive: Tom: We've looked at the conceptual core and the technical improvements, but what we are left with is a clear picture of an entire new serving paradigm for dynamic models.
Jane: It’s truly impressive how this has managed to restore batching capability for mutable state, making the adapted model efficient and reliable while keeping all its unique behaviors.
Lu: I think the biggest impact here is that we are moving beyond merely simulating adaptation; we are actually running it efficiently at scale using these mechanisms.
Meng: From my side, it's a proof of concept that operational viability meets cutting-edge research, providing a clear path to high-throughput AI deployment in the industry.
Lalam: I feel this work guarantees that our future LLMs will be not only smart but also highly accessible and responsive for everyone who uses them.
Tom: So, before we wrap up and move on to the next topic, I want each of you to give us one final thought on the overall impact of this paper.
Jane: It is a foundational piece that proves dynamic adaptation versus static serving is perfectly possible in service architecture today.
Lu: I hope this opens the door for even more complex, evolving AI agents in future research endeavors that rely on continuous adaptation.
Meng: I'm just relieved to know that practical efficiency has a clear, scalable path forward thanks to this framework.
Lalam: The idea of RW-TTT gives us hope that personalized intelligence can coexist with massive scale for every user.
Conclusion: Tom: We’ve spent a lot of time unpacking how the authors built this framework, but it's really worth circling back to what the paper ultimately achieves in its conclusion.
Jane: It proves that you can have dynamic, evolving intelligence in an LLM while still running it at a massive scale that makes sense for deployment.
Lu: I think the biggest conceptual leap is realizing that we don't have to sacrifice batching just because the state is mutable; it’s a complete paradigm shift in how we view server architecture.
Meng: From an engineering standpoint, seeing this approach, it shows us a viable path to handling complex, personalized workloads without the memory bottlenecks that plague current independent replicas.
Lalam: I feel that this work guarantees our future LLMs will be not only smart but also highly accessible and responsive for everyone who uses them.
Tom: It’s truly a monumental achievement in keeping adaptability and efficiency go hand-in-hand, which is what the title, RW-TTT: Batched Serving for Request-Owned Test-Time Training State, encapsulates so well.
Jane: That system allows us to build trust in this technology because it accounts for every single step through its versioning and rollback safeguards.
Lu: It opens the door to even more complex, evolving AI agents that were previously unthinkable due to how we managed state before this architecture exists.
Meng: And since the hardware actually supports two hundred seventy-four tokens per second, it’s not just theoretical; we can actually ship this performance and deliver real-world impact.
Lalam: This architecture allows us to deliver a personalized intelligence experience that feels seamless and consistent for every user, which is a huge win for cultural interaction.
Tom: It's an incredible feat of design, and it's a massive relief to see the technical solutions matching the research goals in this work.
Jane: It gives us so much more confidence in this new direction than simply running things one-by-one, which is a huge step forward for efficiency.
Lu: I’m really excited to see how far this takes us in terms of modeling complex, highly adaptive behavior across diverse tasks.
Meng: We just have to be ready for the next challenge though, as these systems always grow bigger and more demanding on our infrastructure.
Hong Kong University of Science and Technology (HKUST) · Chinese University of Hong Kong (CUHK) · National University of Singapore (NUS)
cs.LG
Submitted: 2026-05-27
Updated: 2026-09-03
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: The paper introduces RW-TTT (Request-Owned Test-Time Training State), a novel serving framework designed to efficiently handle the complex state management required during Test-Time Training (TTT)
Key concepts
- Request-Owned Mutable State
- This is the unique, evolving data state that must be read and updated for each individual request during decoding. The core problem addressed by the paper is managing this complexity, as it cannot be treated like a static parameter.
- Read-Write TTT Serving
- This concept defines a formalized contract for state management. It uses an owner map to ensure that even when multiple requests run concurrently, updates only go into the correct state slot, preventing race conditions from corrupting data.
- Sophisticated Batching Planner
- This unique planner groups requests based on whether they perform compatible READ or WRITE operations. This compatibility matching maximizes throughput by allowing multiple operations to run together in a single batch, boosting efficiency.
Terminology
Summary
The paper introduces RW-TTT (Request-Owned Test-Time Training State), a novel serving framework designed to efficiently handle the complex state management required during Test-Time Training (TTT) within large language model inference. TTT necessitates that each incoming request maintains its own mutable, evolving state, which traditional LLM serving architectures struggle to manage efficiently. RW-TTT addresses this by providing a robust, batched mechanism that allows multiple requests to simultaneously read from and write to their unique, isolated states while maintaining high throughput and memory efficiency.
Mechanism-Level Comparison and State Management
RW-TTT fundamentally changes how mutable state is handled compared to existing serving abstractions. The system is designed as an owner-indexed read/write state,
allowing for the management of multiple, distinct TTT states concurrently. This capability distinguishes it from:
-
Static LLM serving, which uses
Paged KV / continuous batching.
-
Adapter serving, which typically involves a single mutable stream at a time.
The core contract of RW-TTT is that it supports both read and write operations on the TTT state, making it significantly more capable than methods that are limited to only reading or only writing.
The TTT-Aware Serving Loop
The operational flow is governed by a comprehensive serving loop (Algorithm 2). This algorithm manages queueing, wait budgets, and fallback handling in addition to the core processing steps. The process involves the following key stages:
-
State Extraction: For all active requests (r in R), the system retrieves a versioned state view (v r) from the request's TTT State.
-
Operation Execution: Each request executes its operation, resulting in an intermediate set of operations E. This step handles both
READ or WRITE
actions. -
Batch Planning: The system uses a function to plan TTAWARE BATCHES (G) from the accumulated operations E, considering the target batch size and wait budget.
-
Execution and Commit: For each group g in G, an operator group is executed, yielding outputs (Y, S). Crucially, if the group's effect is a
WRITE,
the system mustCOMMIT VERSIONS(S, g.owner map)
to persist the updated state.
Efficiency and Optimization of State Updates
The framework demonstrates significant opportunities for optimization within its write path. The read-write serving contract fixes the operator boundary, requiring separate handling for reading a versioned state view and writing an updated state. On representative H800 BF16 shapes, Triton kernels reveal substantial WRITE-side operator opportunities,
including:
-
Selective commit (ActiveStateUpdateInplace), which showed a best speedup of 3.07×.
-
Fused update + writeback (StateUpdateFused), achieving a median speedup of 2.49×.
-
Checkpoint write, providing a best speedup of 1.20×.
Furthermore, the system's performance is validated across various benchmarks, including the Full 64K RULER task scores,
confirming its scalability and capability to handle large-scale testing environments while maintaining high throughput metrics (e.g., achieving an aggregated token rate of 36.13 tokens/s with 7 streams).
Improvements for AI systems
Based on this research, which details advancements in state management, multi-step reasoning pipelines, and memory efficiency for large-context Language Model serving, I propose the following critical architectural improvements:
1. Implementation of a Unified Read/Write Time Transformer (RW-TTT) Serving Kernel:
-
Improvement: Integrate the
RW-TTTmechanism directly into the core inference engine, replacing sequential or segregated read/write passes. This requires developing specialized Triton kernels that treat state updates (S) and state reads (V) as atomic, owner-indexed operations within a single execution group. -
Capability: The system can execute complex, multi-turn reasoning tasks (like those in the RULER benchmarks) with guaranteed consistency across read and write steps without requiring intermediate checkpointing or flushing the entire state. This significantly reduces latency overhead associated with state commitment, enabling real-time adaptation in interactive agents.
2. Developing an Advanced State Management Layer for Mutable Weights (DeltaAdapterState):
-
Improvement: Formalize the
DeltaAdapterStatecontract as a primary serving abstraction layer, moving beyond simple LoRA attachment. This layer must manage and serialize any mutable weight delta (W) alongside the standard KV cache and adapter weights. The system must enforce versioning and ownership mapping (g.owner map) at this low level. -
Capability: Enables
on-the-fly, expert-specific adaptation
during inference without reloading model weights or requiring external orchestration. An agent can dynamically switch its internal knowledge base (e.g., switching from a medical knowledge module to a financial compliance module) by committing a new S state in milliseconds, making the model contextually fluid and highly specialized for domain tasks.
3. Optimizing the Serving Loop with Predictive Batching and Wait-Budget Management:
-
Improvement: Fully implement Algorithm 2's
PLAN TTTAWARE BATCHES(E, B, w)function. This requires a scheduler that analyzes the dependency graph of all active requests (R) based on their required state reads/writes (E) and groups them into optimal execution batches (G) considering both hardware constraints (peak GiB) and request deadlines (wait budget w). -
Capability: Maximizes GPU utilization far beyond standard continuous batching. Instead of processing requests purely by token count or arrival time, the system processes requests in dependency-aware
compute groups.
This ensures that when one request is waiting on a write operation from another, the GPU immediately switches to executing a necessary read operation for a third request, minimizing idle cycles and maximizing aggregate throughput (Agg. tok/s).
4. Implementing Context-Aware Memory Offloading and Scaling:
-
Improvement: Extend the 64K/128K token support by integrating the memory management techniques proven effective in the 16K stress tests. The system must dynamically scale state storage, treating the TTT state (V, S) as a first-class citizen that can be selectively paged out to slower, high-capacity memory tiers (e.g., HBM2e/DDR) when the active working set falls below a critical threshold, while retaining immediate rollback capability.
-
Capability: Allows the model to reliably handle unprecedented context lengths (e.g., 100K+ tokens) for document analysis or long-form reasoning without hitting catastrophic Out-of-Memory (OOM) errors, maintaining high performance even at extreme scale.
5. Establishing a Standardized, Efficient State Projection Interface:
-
Improvement: Formalize the
Read-path projection
mechanism (e.g.,MutableWeightLinear,FusedReadResidual) into the standard inference API. These projections must be benchmarked against native cuBLAS/torch primitives to ensure that the overhead of maintaining state ownership does not negate computational gains. -
Capability: Provides a clean, optimized pathway for the model to read specific, versioned views of its internal state or external knowledge bases during inference. This allows for high-fidelity retrieval-augmented generation (RAG) where the retrieval component isn't just a prompt injection but an integral part of the weight computation graph.
Sources
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Test-Time Learning for Large Language Models
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
- DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- S-LoRA: Serving Thousands of Concurrent LoRA Adapters
- Efficient LLM Serving on Hybrid Real-time and Best-effort Requests
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- In-Place Test-Time Training
- DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
- End-to-End Test-Time Training for Long Context
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks