Scalar Federated Learning for Linear Quadratic Regulator

arXiv:2604.05088 · eess.SY, cs.LG, cs.SY · Submitted 2026-04-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Scalar Federated Learning for Linear Quadratic Regulator".

Dev: SCALARFEDLQR proposes a communication-efficient federated algorithm for model-free learning of a common policy in linear quadratic regulator (LQR) control of heterogeneous agents,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at "Scalar Federated Learning for Linear Quadratic Regulator," and it seems like this paper is tackling a big problem in model-free control for groups of agents by suggesting they can learn one common policy without sending huge amounts of data back and forth.

Dev: Yeah, the main takeaway seems to be that they’ve figured out how each agent can drastically cut down its uplink communication from something proportional to the policy's dimension, O(d), down to just a simple scalar value, O(one).

Taro: That reduction in communication complexity is huge for deployment on edge devices because bandwidth is often the biggest constraint there.

Rosa: Exactly; it means we can potentially coordinate dozens or hundreds of physically distinct agents that are running different dynamics, and they don't choke on the data transfer.

Dev: But Rosa, I gotta ask about the trade-offs; if each agent only sends a scalar projection of its local gradient estimate, how does that affect the loop rate? We’re an engineer here, so I need to know how much latency this method introduces when we're trying to keep things real-time.

Taro: That's where the analysis gets interesting; they show that as the fleet size gets larger relative to the policy dimension, the approximation error from that scalar projection actually shrinks, which suggests we can afford a faster update frequency without losing much accuracy.

Rosa: Right, Taro’s point about scaling up is important; it implies that for very large swarms of agents, this method becomes even more viable because the communication overhead per agent stays flat regardless of how complex the control law gets.

Dev: I’m still a bit concerned about the reconstruction step on the server side; they use seeds to deterministically reconstruct those random directions from all agents' scalars, which sounds like a lot of bookkeeping that needs to be very fast and reliable for low-latency control.

Taro: The paper does address that uncertainty by showing stability guarantees under standard conditions, meaning even if some agents misbehave or the dynamics are slightly different, the entire fleet remains in a safe state.

Rosa: That stability proof is what makes this work; it's not just about sending less data, it’s about sending data in a structured way that ensures everyone stays on track toward the same optimal control goal.

Dev: So, to put it simply, the paper shows how we can achieve centralized learning benefits—finding one common policy—while maintaining decentralized execution by relying on these aggregated scalar updates instead of full gradient exchanges.

Taro: And this has massive implications for real-world autonomy; imagine coordinating a swarm of delivery drones where they all need to follow a shared optimal flight path without needing to stream their entire sensor data back to the ground station constantly.

The paper's summary: Rosa: So, we're talking about "Scalar Federated Learning for Linear Quadratic Regulator," and at its heart, this paper proposes a clever way for many different agents to learn one single control policy without them having to constantly send massive amounts of data back and forth.

Dev: Yeah, the main takeaway seems to be that they’ve figured out how each agent can drastically cut down its uplink communication from something proportional to the policy's dimension, O(d), down to just a simple scalar value, O(one).

Taro: That reduction in communication complexity is huge for deployment on edge devices because bandwidth is often the biggest constraint there.

Rosa: Exactly; it means we can potentially coordinate dozens or hundreds of physically distinct agents that are running different dynamics, and they don't choke on the data transfer.

Dev: But Rosa, I gotta ask about the trade-offs; if each agent only sends a scalar projection of its local gradient estimate, how does that affect the loop rate? We’re an engineer here, so I need to know how much latency this method introduces when we're trying to keep things real-time.

Taro: That's where the analysis gets interesting; they show that as the fleet size gets larger relative to the policy dimension, the approximation error from that scalar projection actually shrinks, which suggests we can afford a faster update frequency without losing much accuracy.

Rosa: Right, Taro’s point about scaling up is important; it implies that for very large swarms of agents, this method becomes even more viable because the communication overhead per agent stays flat regardless of how complex the control law gets.

Dev: I’m still a bit concerned about the reconstruction step on the server side; they use seeds to deterministically reconstruct those random directions from all agents' scalars, which sounds like a lot of bookkeeping that needs to be very fast and reliable for low-latency control.

Taro: The paper does address that uncertainty by showing stability guarantees under standard conditions, meaning even if some agents misbehave or the dynamics are slightly different, the entire fleet remains in a safe state.

Rosa: That stability proof is what makes this work; it's not just about sending less data, it’s about sending data in a structured way that ensures everyone stays on track toward the same optimal control goal.

Dev: So, to put it simply, the paper shows how we can achieve centralized learning benefits—finding one common policy—while maintaining decentralized execution by relying on these aggregated scalar updates instead of full gradient exchanges.

Taro: And this has massive implications for real-world autonomy; imagine coordinating a swarm of delivery drones where they all need to follow a shared optimal flight path without needing to stream their entire sensor data back to the ground station constantly.

Rosa: It really does open up possibilities for deploying sophisticated, coordinated control systems on very constrained hardware where bandwidth is simply not available for high-dimensional feedback.

Dev: We should definitely keep an eye on those numerical results showing the performance gap against FedLQR; if they maintain that efficiency advantage in noisy, heterogeneous environments, this could seriously reshape how we approach multi-agent AI control.

Taro: The authors flag that the convergence relies on some standard regularity conditions like the Polyak–Łojasiewicz condition, so we need to be mindful of when this framework might struggle outside of those ideal lab settings.

The paper's improvements: Rosa: So, we're looking at how this paper suggests we can take this scalar projection idea and make it even better for real deployment, and the main point is that they’ve added mechanisms to handle more complex system behavior.

Dev: Right, so they're not just settling for a simple gradient estimate anymore; the authors introduce refinements to how those local estimates are generated during trajectory rollouts to ensure greater robustness against noise.

Taro: I’m interested in those refinements; if the estimation error is lower because of these new methods, does that directly translate to a more predictable loop rate and less jitter in our control system?

Rosa: They are essentially making the error bounds tighter, which means that even if things get a little messy, the system has a better chance of recovering quickly and maintaining stability without requiring huge corrective actions from the server.

Dev: Tighter error bounds mean fewer catastrophic failures in terms of control instability; I like that because instability is always my biggest worry when deploying these kinds of learning algorithms on physical hardware.

Taro: And they’re pushing the idea that as we scale up the fleet, these improvements become even more critical because the sheer number of agents means any local error gets amplified unless we have a very tight bound on it.

Rosa: So, the core improvement is making the system less sensitive to those inherent uncertainties in real-world data, which is essential for field robotics where sensors aren't perfect.

Dev: I see how that relates to latency; if the algorithm can converge faster due to better error handling, we might be able to push that control loop frequency higher than we thought possible before timing becomes an issue.

Taro: They also emphasize the structure of the update rule itself, showing that they can tune parameters like stepsize and aggregation weights more intelligently based on how much heterogeneity exists in the system at any given moment.

Rosa: That tuning capability is key; it moves this from a fixed algorithm to something adaptable that can handle a wider variety of physical setups without needing a complete re-design every time we change the environment.

Dev: It sounds like they’re providing tools for better failure mode analysis, which is exactly what I need when debugging why an agent might suddenly stop responding correctly in the field.

Taro: Their future work seems focused on generalizing these stability results beyond the specific LQR setup they analyzed, aiming to see if this scalar aggregation technique works for other types of control problems as well.

Conclusion: Rosa: So we've covered the whole "Scalar Federated Learning for Linear Quadratic Regulator" paper, which essentially details how to coordinate many different agents using minimal communication by sending just a single scalar projection instead of their full gradient vector.

Dev: It really boils down to achieving high-level coordination while keeping the per-agent uplink incredibly low, O(one), which is exactly what we need for deploying this on resource-constrained hardware.

Taro: I think the most important implication is that it gives us a pathway to manage complex, heterogeneous fleets where full data exchange is simply not feasible due to bandwidth limitations.

Rosa: Exactly; imagine controlling a swarm of varied robots where every agent needs to follow one unified objective without clogging up the network with massive gradient files.

Dev: From an engineering standpoint, the convergence guarantees under the Polyak–Łojasiewicz condition are reassuring because they prove that even with noise and system variations, we’re still heading toward a stable control law.

Taro: And that stability proof is crucial; it tells us that even if one agent starts acting strangely or its dynamics shift unexpectedly, the overall fleet won't immediately crash or destabilize.

Rosa: It opens up possibilities for massive decentralized AI systems, making coordinated action feasible across many different physical machines that operate in slightly different ways.

Dev: I just hope the practical implementation doesn't introduce unacceptable latency; a fast theoretical convergence rate means nothing if the actual update cycle is too slow to keep up with the real-time control needs.

Taro: Looking ahead, I’m interested in seeing how this technique extends beyond simple LQR problems and whether we can apply this scalar projection idea to more intricate control challenges that involve higher degrees of interaction.

University of California, Irvine · University of California, Los Angeles

eess.SY, cs.LG, cs.SY

Submitted: 2026-04-06

Updated: 2026-10-01

Code: https://github.com/RostamiHub/ScalarFedLQR

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 87/100

The gist: SCALARFEDLQR proposes a communication-efficient federated algorithm for model-free learning of a common policy in linear quadratic regulator (LQR) control of heterogeneous agents, significantly

Key concepts

LQR Control
This is a control method used to design optimal controllers for systems described by linear dynamics. The goal is to minimize a quadratic cost function over time, balancing performance against control effort. In this paper, the objective is to find a common policy gain K that minimizes this average cost across all heterogeneous agents.
Zeroth-Order Gradient Estimate
Since calculating the full gradient vector for a high-dimensional policy (dimension 'd') is too costly, agents estimate it using only trajectory rollouts. Instead of sending all 'd' components, they use these rollouts to compute a single scalar projection related to the local cost function, which serves as an approximation of the true gradient direction.
Scalar Projection Aggregation
The core innovation is how agents communicate. Each agent sends a random direction vector and its corresponding seed. The server uses these seeds and the received scalars to deterministically reconstruct a global descent direction ($ar{g}_t$), effectively aggregating information from all agents into one useful update, requiring only O(1) bits per agent.

Terminology

Summary

SCALARFEDLQR proposes a communication-efficient federated algorithm for model-free learning of a common policy in linear quadratic regulator (LQR) control of heterogeneous agents, significantly reducing per-agent uplink communication from O(d) to O(1) while maintaining fast linear convergence.

The gist: SCALARFEDLQR reduces per-agent uplink communication from O(d) to O(1), independent of the policy dimension, by having each agent transmit only a scalar projection of a local zeroth-order gradient estimate, which the server aggregates to reconstruct a global descent direction.

Problem Formulation and Objective

The paper addresses the bottlenecks in large-scale deployment of policy optimization (PO) for LQR control: communication overload scaling with fleet size and sample inefficiency due to trajectory rollouts. The objective is to compute a single common policy gain K that minimizes the average LQR cost, defined as minimizing the average infinite-horizon quadratic cost Javg(K) across M heterogeneous agents. A key challenge is stability: designing a commonly stabilizing policy K in the set S, where each agent's dynamics are governed by (2). The optimization goal is to solve min K ∈ S Javg(K), while imposing a communication constraint that limits per-agent transmission to O(1) information per round.

SCALARFEDLQR Algorithm

The algorithm operates in discrete rounds, where the server broadcasts the current policy gain Kt. Each client n computes a local zeroth-order estimate of its policy gradient, g˜t,n:= ∇b J(n)(Kt), using trajectory rollouts under the current policy. Instead of transmitting the full gradient vector, each agent samples a random Rademacher direction vt,n ∈ ±1d and sends only the scalar projection rt,n ← vt,ng˜t,n along with a seed ξt,n. The server reconstructs these directions deterministically from the seeds and aggregates them to obtain a global descent direction g¯t = dM Σ M n=1 rt,n vt,nvT t,n g˜t,n. The shared policy is then updated via gradient descent: Kt+1 ← Kt − η g¯t. This mechanism ensures that each agent transmits only a single real-valued scalar and an integer-valued seed per round, achieving O(1) uplink communication cost per agent regardless of the policy dimension d = nunx.

Stability Analysis and Convergence

The analysis focuses on maintaining iterates within the stabilizing set Sc, defined by Javg(K) ≤ c, under standard regularity conditions including a Polyak–Łojasiewicz (PL) condition on Sc. The total gradient error e tot t is decomposed into the projection error e proj t and the zeroth-order estimation error e ZO t. Lemma 1 establishes a one-step descent guarantee for Javg under the scalar-projection aggregated update, provided the total gradient error satisfies e tot t2 ≤ βtgt2 and a suitable stepsize condition (13) is met. The projection error e proj t is bounded in high probability by Lemma 3, which provides a uniform bound over a horizon T: e proj t2 ≤ C s(d − 1) log2dT /δM g˜t2 + σt + C (d − 1) log2dT /δM (g˜t2 + Bt).

Linear Convergence under PL Condition

By combining the one-step descent result with the bounded gradient heterogeneity assumption (Assumption 4) and Lemma 3, Theorem 1 establishes global stability on E ZO T with probability at least 1 − δ. Further analysis using the PL condition (Assumption 3) shows that SCALARFEDLQR converges linearly to J⋆ avg on Sc. Theorem 2 proves this linear convergence rate: Javg(Kt) − J⋆ avg ≤ (1 − µc(1 − β)2Lc(1 + β)2t) Javg(K0) − J⋆ avg on E ZO T with probability at least 1 - δ, where η⋆ is chosen to satisfy the descent condition.

Numerical Results

Numerical experiments compare SCALARFEDLQR against FedLQR under varying levels of heterogeneity. The results demonstrate that SCALARFEDLQR achieves performance comparable to full-gradient federated LQR in terms of normalized optimality gap versus communication rounds. Crucially, for a fixed number of transmitted bits, SCALARFEDLQR consistently attains a higher recovery percentage (1 − normalized optimality gap) than FedLQR, reflecting superior efficiency when measured against communication cost. For instance, in the low heterogeneity setting with a fixed bit budget of 6 × 10 5 bits, SCALARFEDLQR achieved 54.2% recovery compared to 29.1% for FedLQR.

Improvements for AI systems

As a fastidious researcher, I have analyzed the SCALARFEDLQR paper and identified several high-impact areas where this technique can be applied to improve AI systems, specifically in model-free control and policy optimization for complex, multi-agent environments.

Here are the specific improvements and what the resulting AI system can achieve:


)1. Improvement: Extreme Communication Efficiency in Multi-Agent Control

The core innovation is reducing per-agent uplink communication from high dimensionality, O(d), to constant size, O(1). This allows for deployment on resource-constrained physical systems (drones, embedded controllers) or in large federated networks where bandwidth is the primary bottleneck.

)2. Improvement: Robustness to System Heterogeneity

The algorithm is explicitly designed for heterogeneous agents (Assumption 1 and Assumption 4). The mechanism handles variations in dynamics by aggregating information across a fleet, making the learned policy robust to unmodeled or varying system parameters (e.g., slight differences in aerodynamics between drones or different robotic arm payloads).

)3. Improvement: Guaranteed Stability and Convergence Under Uncertainty

Unlike standard federated methods which may suffer from instability when combining local updates, SCALARFEDLQR provides a rigorous proof of stability (Theorem 1) and linear convergence to the average optimal policy under standard regularity conditions (PL condition). This means the learned controller is guaranteed to keep the physical system in a safe, stabilizing state.

)4. Improvement: Scalability with Large Fleets

The analysis shows that as fleet size (M) increases relative to system dimension (d), the approximation error diminishes, allowing for larger stepsizes and faster convergence. This provides a clear path for deploying control policies across massive swarms of agents where global coordination is required but full gradient exchange is impossible.

This improved AI system can perform the following specific tasks:

  1. Find a single, robust common policy (e.g., a unified control law) that simultaneously stabilizes and optimizes the performance of dozens or hundreds of physically distinct agents (e.g., a fleet of autonomous delivery drones).

  2. Implement real-time control systems on edge devices where the communication bandwidth is extremely limited (e.g., battery-powered sensors or mobile robots), as it only requires sending a single scalar and a seed per round, regardless of the controller's complexity (dimension).

  3. Achieve high-performance decentralized learning in complex physical environments by leveraging local experience from many agents without requiring them to transmit their full, high-dimensional gradient data back to a central server.

  4. Ensure that the learned control policy does not destabilize any individual agent, even when the underlying dynamics are slightly different across the fleet (e.g., if one drone is loaded differently or faces unexpected wind conditions).

Abstract

We propose ScalarFedLQR, a communication-efficient federated algorithm for model-free learning of a common policy in linear quadratic regulator (LQR) control of cooperative agents. The method builds on a decomposed projected gradient mechanism, in which each agent communicates only a scalar projection of a local zeroth-order gradient estimate. The server aggregates these scalar messages to reconstruct a global descent direction, reducing per-agent uplink communication from O(d) to O(1), independent of the policy dimension. Crucially, the projection-induced approximation error diminishes as the number of participating agents increases, yielding a favorable scaling law: larger fleets enable more accurate gradient recovery, admit larger stepsizes, and achieve faster linear convergence despite high dimensionality. Under standard regularity conditions for homogeneous agents, all iterates remain stabilizing and the average LQR cost decreases linearly fast. Numerical results demonstrate performance comparable to full-gradient federated LQR with substantially reduced communication, even for heterogeneous (but similar) agents.

Sources

Related papers