MLTCP: Congestion Control for DNN Training

summary

Video file (mp4)

The gist

MLTCP is a technique to augment today’s congestion control algorithms to accelerate DNN training jobs in shared GPU clusters.

This episode discusses

The paper

MLTCP: Congestion Control for DNN Training · Read on arXiv

Sudarsanan Rajasekaran, Sanjoli Narang, Anton A. Zabreyko, Manya Ghobadi

Massachusetts Institute of Technology

We present MLTCP, a technique to augment today's congestion control algorithms to accelerate DNN training jobs in shared GPU clusters. MLTCP enables the communication phases of jobs that compete for network bandwidth to interleave with each other, thereby utilizing the network efficiently. At the heart of MLTCP lies a very simple principle based on a key conceptual insight: DNN training flows should scale their congestion window size based on the number of bytes sent at each training iteration. We show that integrating this principle into today's congestion control protocols is straightforward: by adding 30-60 lines of code to Reno, CUBIC, or DCQCN, MLTCP stabilizes flows of different jobs into an interleaved state within a few training iterations, regardless of the number of competing flows or the start time of each flow. Our experiments with popular DNN training jobs demonstrate that enabling MLTCP accelerates the average and 99th percentile training iteration time by up to 2x and 4x, respectively.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MLTCP: Congestion Control for DNN Training".

Jane: The paper was written by Sudarsanan Rajasekaran, Sanjoli Narang, Anton A. Zabreyko and Manya Ghobadi from Massachusetts Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we're digging into a paper that's been making waves in the networking world, and it's called "MLTCP: Congestion Control for DNN Training." Jane, I've got to say, the title alone got me excited because it's tackling a problem that's been nagging at anyone who runs large AI models.

Jane: Oh, absolutely, Tom. And for our listeners who might not be networking folks, let me break down what we're looking at. When you train a big neural network across multiple GPUs, those GPUs have to talk to each other constantly, sending millions of parameters back and forth. That traffic goes over the network, and when multiple training jobs share the same network, they start fighting for bandwidth. This paper is about fixing that fight.

Tom: Right, and the clever part is in the name. MLTCP stands for Machine Learning TCP, so they're taking the standard internet protocol that moves data around and tweaking it specifically for AI training. I mean, that's a pretty bold move, right? Messing with the fundamental protocol?

Jane: It is, but here's the thing, Tom. The authors, Sudarsanan Rajasekaran, Sanjoli Narang, Anton Zabreyko, and Manya Ghobadi from MIT, they noticed something fascinating. When two training jobs collide on the network, they both slow down. But if you could somehow make their communication phases alternate, like two people taking turns using a shared printer, everything speeds up.

Tom: That's the interleaving concept, and I love how they describe it in the paper. They're not just trying to make the network fairer. They're trying to make it smarter by recognizing that DNN training is periodic. Every iteration, the same amount of data gets sent. So they can predict and shape the traffic.

Jane: Exactly. And the beauty is, they don't need a central controller telling everyone what to do. Each flow just adjusts its own behavior based on how many bytes it has already sent in the current iteration. It's like each job is saying, "Hey, I'm almost done with my turn, let me grab a bit more bandwidth to finish up."

Tom: So we're talking about a distributed solution that just works by tweaking the congestion control algorithm. No new hardware, no central scheduler, just a change in how aggressive each flow is. That's the kind of elegance that gets me excited.

Jane: And the results are pretty dramatic. We're going to get into the numbers in a bit, but they're seeing up to two times faster average training times and four times faster for the worst-case iterations. That's huge for anyone running large models in shared clusters.

Tom: I can't wait to dig into the details. But first, let's give a quick shout-out to the authors for tackling this real-world problem. So, stay tuned because next we're going to look at the summary of the paper and how they actually pulled this off.

Summary: Jane: So, Tom, we've established that "MLTCP: Congestion Control for DNN Training" is about making network traffic smarter for AI workloads. Now let's talk about the core idea in the summary, because it's deceptively simple.

Tom: It really is. The whole trick is based on a principle that sounds almost too easy. They change how the congestion window grows based on how much data has already been sent in the current training iteration. So instead of every flow increasing its window at the same rate, the flow that's closer to finishing its communication phase gets a boost.

Jane: And why does that help? Well, imagine two jobs are sharing a link. If they both try to send their gradients at the same time, they collide and slow each other down. But if one job is slightly ahead, it gets more bandwidth, finishes its communication faster, and then the other job gets the whole link to itself. They naturally fall into a rhythm where they take turns.

Tom: Right, and the paper calls this the job favoritism policy. It's basically a distributed way of implementing something called Shortest Remaining Processing Time, or SRPT. In simple terms, the flow that has the least amount of data left to send gets priority. That's a classic scheduling idea that works great in theory, but usually requires a central scheduler to know everything.

Jane: But here's the kicker, Tom. They figured out how to do it without any central knowledge. Each flow just looks at its own bytes sent ratio, which is the number of bytes it has successfully sent divided by the total bytes in the iteration. And because DNN training is so predictable, that total number is known and constant.

Tom: That predictability is the secret sauce. In normal internet traffic, you never know how big a flow is going to be. But in DNN training, every iteration sends the same amount of data, every single time. So the flow can calculate its own progress and adjust its aggressiveness accordingly.

Jane: And they show this works across different congestion control algorithms. They modified Reno, CUBIC, and even DCQCN, which is used in RDMA networks. And they only had to add about thirty to sixty lines of code for each one. That's a remarkably small change for such a big payoff.

Tom: The paper reports that the system stabilizes into this interleaved state within about ten training iterations. That's fast. And once it's there, it stays there, which is what you want. No oscillation, no thrashing.

Jane: Exactly. And the improvements are not just in average time. They're seeing massive reductions in packet drops and ECN marks, which means the network is just healthier overall. Less congestion, less retransmission, less wasted work.

Tom: So the summary is clear: a simple, distributed tweak to congestion control can unlock significant speedups for AI training. But I'm curious about the practical side. How did they actually test this? That's what we're going to dig into next.

Improvements: Tom: Welcome back. We've talked about the core idea of "MLTCP: Congestion Control for DNN Training," and now I want to bring in our panel to discuss the actual improvements and the experiments. Lu, you're our AI researcher, what stands out to you?

Lu: Tom, the thing that excites me most is the robustness. The paper doesn't just show it works in a perfect lab setting. They tested it against three real-world challenges. First, stragglers, which are servers that randomly slow down. Second, partially compatible jobs, where the communication patterns don't perfectly line up. And third, circular dependencies, where job A competes with job B on one link, B competes with C on another, and C competes with A on a third.

Jane: And that circular dependency is a nightmare for any centralized scheduler. It's like a rock-paper-scissors game where you can't pick a winner. But the paper shows MLTCP just handles it naturally because it's not trying to find a global schedule. Each flow just reacts locally.

Meng: As an engineer, I'm always skeptical of things that work in simulation but fall apart in the real world. But they built a real testbed with twelve servers, each with an A100 GPU and a fifty Gbps NIC. That's a serious setup. And they tested with real models like GPT-two VGG16, and RoBERTa. That gives me confidence.

Tom: And the numbers are pretty wild. They saw up to fourteen point five nine times fewer ECN-marked packets with their MLQCN variant. That's a massive reduction in congestion signals, which means the network is just not being stressed as much.

Lu: Right, and that translates directly to training speed. The average iteration time improved by up to two times, and the 99th percentile tail latency improved by up to four times. That tail latency is critical because in distributed training, you're only as fast as your slowest worker. If one iteration takes forever, it stalls the whole job.

Jane: And I love that they compared against prior work. They showed that the Static approach, which just forces unfair bandwidth sharing, falls apart when jobs are only partially compatible. It actually makes things worse for one job. But MLTCP gracefully degrades and still provides speedups.

Meng: That's the key difference. The Static approach is like setting a thermostat to a fixed temperature and hoping it works. MLTCP is like having a smart thermostat that constantly adjusts based on the actual conditions. It's adaptive, and that's why it's robust.

Tom: So we've got a solution that's simple, robust, and tested on real hardware. But what does this mean for the broader world? Lalam, I know you have thoughts on the cultural and practical impact.

Lalam: Tom, I think this paper represents a shift in how we think about network fairness. Traditionally, we've wanted every flow to get equal bandwidth. But this paper shows that for periodic workloads like AI training, a little bit of temporary unfairness can lead to better outcomes for everyone. It's a more nuanced view of fairness, one that considers the overall job completion time rather than instantaneous bandwidth share.

Jane: That's a profound point. It's not about being fair in the moment; it's about being efficient over time. And that's a lesson that could apply beyond networking, to resource allocation in general.

Tom: Absolutely. So we've got a paper that's technically brilliant, practically tested, and philosophically interesting. Let's wrap up our thoughts in the conclusion.

Conclusion: Tom: Alright, we've spent some time with "MLTCP: Congestion Control for DNN Training," and I think it's fair to say this is one of those papers that makes you wonder why nobody thought of it sooner.

Jane: It really is. The core insight is so elegant: because DNN training is periodic, each flow knows exactly how much data it needs to send per iteration. So it can adjust its own aggressiveness based on how far along it is. That simple change causes competing jobs to naturally interleave their communication phases, avoiding congestion altogether.

Tom: And the results speak for themselves. Up to two times faster average training, four times faster tail latency, and a massive reduction in packet drops. All from a thirty to sixty line code change to existing congestion control algorithms.

Lu: And the robustness is what really sells it. They didn't just show it works in a perfect world. They showed it handles stragglers, partially compatible jobs, and even circular dependencies. That's the kind of real-world readiness that makes a paper impactful.

Meng: From an engineering standpoint, the fact that they built a real testbed and tested with real models gives me confidence that this could actually be deployed. It's not just a theoretical curiosity. It's a practical tool.

Lalam: And I think the broader implication is that we need to rethink what fairness means in shared infrastructure. Sometimes, a little bit of short-term unfairness, when guided by the right principles, can lead to better outcomes for everyone. That's a powerful idea that could influence how we design all sorts of shared systems.

Jane: Well said, Lalam. So, as we say goodbye to this paper, I want to thank our listeners for joining us. We've covered the title, the summary, and the improvements, and I hope you're as excited about this as we are.

Tom: Definitely. "MLTCP: Congestion Control for DNN Training" is a fantastic piece of work, and we're looking forward to seeing how it evolves. Maybe we'll see it integrated into real clusters soon. Until next time, keep those GPUs busy and those networks happy.

Jane: Thanks for listening, everyone. We'll be back with another paper soon. Take care.

More episodes

← Home