Federated Self-Supervised Modulation Classification under Non-IID and Imbalanced Data

summary

Video file (mp4)

The gist

conditions typical in AMC tasks.

In short

The episode discusses a paper titled 'Federated Self-Supervised Modulation Classification under Non-IID and Class-Imbalanced Data.' The hosts explain how this system works—training a shared encoder on unlabeled data first, then personalizing a small classifier. They conclude that this approach is highly robust to real-world data issues like heterogeneity and scarcity.

Key concepts

Federated Learning
A training method where raw data is not collected in one central location. Instead, local devices learn independently and only share model updates with the central server, which protects privacy and saves bandwidth.
Self-Supervised Learning
A technique used to train models without needing many expensive labeled examples. The system learns patterns from raw, unlabeled data first (e's encoder), then uses a small amount of labeled data to finish the classification task.

Terminology used across episodes

This episode discusses

The paper

Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data · Read on arXiv

Usman Akram, Yiyue Chen, Haris Vikalo

University of Texas at Austin · Qualcomm Technologies Inc.

Training automatic modulation classification (AMC) models on centrally aggregated data raises privacy concerns, incurs communication overhead, and often fails to confer robustness to channel shifts. Federated learning (FL) avoids central aggregation by training on distributed clients but remains sensitive to class imbalance, non-IID client distributions, and limited labeled samples. We propose FedSSL-AMC, which trains a causal, time-dilated CNN with triplet-loss self-supervision on unlabeled I/Q sequences across clients, followed by per-client SVMs on small labeled sets. We establish convergence of the federated representation learning procedure and a separability guarantee for the downstream classifier under feature noise. Experiments on synthetic and over-the-air datasets show consistent gains over supervised FL baselines under heterogeneous SNR, carrier-frequency offsets, and non-IID label partitions.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Federated Self-Supervised Modulation Classification under Non-IID and Imbalanced Data".

Jane: The paper was written by Usman Akram, Yiyue Chen and Haris Vikalo from University of Texas at Austin and Qualcomm Technologies Inc..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a real mouthful of a title — "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." Jane, I'm going to need you to break that down for me, because that's a lot of jargon in one sentence.

Jane: Happy to, Tom. So imagine you've got a bunch of cell phones or radios scattered around a city, and each one is picking up wireless signals. Automatic modulation classification is basically figuring out what kind of signal it is — is it a simple BPSK signal, or a more complex sixteen-QAM? That matters for managing the spectrum efficiently.

Tom: Right, and the "federated" part means we're not collecting all that raw data in one central server. Each device learns locally and only shares the model updates. That protects privacy and saves bandwidth.

Jane: Exactly. But here's the catch — the paper's title mentions "Non-IID and Class-Imbalanced Data." In plain English, that means each device sees a different slice of reality. One phone might mostly see BPSK signals, another mostly sees QPSK. And some signal types are rare. That throws traditional training methods off.

Tom: And that's where "self-supervised" comes in, right? Instead of needing tons of labeled examples — which are expensive to get for radio signals — the model learns patterns from raw, unlabeled data first. Then it only needs a tiny bit of labeled data to finish the job.

Jane: You got it. The team at UT Austin, led by Usman Akram, Yiyue Chen, and Haris Vikalo, built a system that does exactly that. They call it FedSSL-AMC. And the results are pretty impressive — on their synthetic dataset, they hit over fifty-five percent accuracy where the best supervised baseline only managed around forty-two percent.

Tom: That's a big jump. But I'm curious about the real-world part — they also tested on something called the MIGOU dataset, which is actual over-the-air signals. Jane, what happened there?

Jane: The gap got even wider. Under heavy data heterogeneity, FedSSL-AMC hit over eighty-two percent accuracy while the standard federated learning baselines were stuck in the 30s and 40s. It's a clear win for the self-supervised approach.

Tom: So the title is basically advertising the solution to a real problem — wireless networks are messy, data is uneven, and labels are scarce. This paper says, hey, we can still make this work. Stick around, because next we're going to talk about how they actually built this thing.

Summary: Tom: Welcome back. We're still on "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." Last time we covered the big picture — why this matters for wireless networks. Now let's get into the meat of it. Jane, how does this thing actually work?

Jane: So the clever part is they split the problem into two stages. First, they train a shared "encoder" — think of it as a feature extractor — using a triplet loss on unlabeled I/Q data. That's the raw in-phase and quadrature signal samples. The encoder learns to recognize patterns: "these two signal snippets look similar, those two look different."

Tom: And that's the self-supervised part. No labels needed. But what kind of neural network are they using for this encoder?

Jane: They use a causal convolutional neural network with time dilation. That's a fancy way of saying the network looks at the signal over a long window of time, but it does it efficiently. Each layer skips ahead exponentially, so it can see far back in time without needing a huge number of parameters.

Tom: And after that encoder is trained collaboratively across all the clients, each client trains its own tiny classifier on its own labeled data. That's the personalization step.

Jane: Right. They use a support vector machine — an SVM — which is lightweight and works well with small labeled sets. So the heavy lifting is done in the unsupervised phase, and the labeled data just fine-tunes the final decision.

Tom: Now, the paper also has some serious math in it. They prove that this training process converges — meaning it reliably reaches a good solution even when clients have very different data. Lu, you're our theory person. What did you make of that?

Lu: I was actually quite impressed. They establish a convergence bound for the federated representation learning, and they show that the variance of the stochastic gradients stays bounded. That's the kind of guarantee you want before deploying something in the real world. They also prove a separability condition — basically, how much signal-to-noise ratio you need for the downstream SVM to work reliably.

Tom: So it's not just "we tried it and it worked" — there's actual theory backing it up.

Lu: Exactly. And the bound they get in the high-SNR limit is clean. The average squared gradient norm goes to something like 64βM plus 2m2R2B over W, plus a small projection error term. It's a solid result.

Jane: And the practical payoff is that this approach is robust to all kinds of heterogeneity — different SNRs across clients, different carrier frequency offsets, even different quantization levels. We'll dig into those experiments next.

Tom: Great, because I want to see how it holds up when you really stress-test it. That's coming up right after this.

Improvements: Tom: We're back, still talking about "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." We've covered the setup and the theory. Now, Jane, what did they actually test, and what improvements did they show?

Jane: They ran a whole battery of experiments. First, on their synthetic dataset, they compared FedSSL-AMC against four supervised baselines: FedAVG-CNN, FedeAMC, FedProx, and FedDyn. And they also compared against a SimCSE-style contrastive baseline. Under standard label imbalance, FedSSL-AMC hit fifty-five point four one percent accuracy, while the best baseline, SimCSE, only got fifty-one point five five percent.

Tom: And when they cranked up the heterogeneity — different SNR ranges per client — the gap stayed. FedSSL-AMC got forty-one point four two percent on the worst client, while FedAVG-CNN only managed thirty-one point six four percent there.

Jane: Right. And they also tested mobility-induced carrier frequency offset. They simulated four mobility regimes, from ultra-low to high Doppler shift. Even with that extra distortion, FedSSL-AMC stayed on top.

Meng: I want to jump in here, because as an engineer, I care about whether this thing can actually run on real hardware. The paper reports the encoder has zero point two four seven million parameters — that's tiny compared to the one point seven eight million in the baseline CNN. But it does need more compute: four hundred seventy-three MFLOPs versus seventeen point seven six. That's a real tradeoff.

Tom: So it's more compute per inference, but way fewer parameters. Is that a problem for edge devices?

Meng: It's manageable. The compute is dominated by the contrastive loss during training, not inference. And the communication cost during federated training is actually lower because the model is smaller. For a radio on a drone or a sensor node, that's a win.

Jane: And on the MIGOU dataset — that's real over-the-air data — the improvements were even more dramatic. Under heavy heterogeneity, FedSSL-AMC hit eighty-two point five nine percent accuracy, while FedeAMC only got thirty-three point two zero percent. That's a massive jump.

Tom: Meng, what about the quantization experiment? They tested with clients using different numerical precisions — float32, float16, and int8.

Meng: That's the model heterogeneity test. And even with that mess, FedSSL-AMC stayed ahead of all the supervised baselines. That tells me the self-supervised representation is robust to the noise introduced by quantization, which is really useful for real deployments where devices have different capabilities.

Tom: So the improvements aren't just incremental — they're substantial across every heterogeneity scenario they threw at it.

Jane: Exactly. And that's the story of this paper. It's not just a new algorithm; it's a new way of thinking about how to train models when your data is messy and your labels are scarce.

Tom: Alright, we're going to wrap this up with our final thoughts and what this means for the future. Stay with us.

Conclusion: Tom: Alright, we're at the end of our discussion on "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." Jane, give us the final summary.

Jane: So the big takeaway is that by decoupling representation learning from classification — training a shared encoder on unlabeled data first, then personalizing a tiny classifier per client — you get a system that's far more robust to the messy realities of wireless networks. Non-IID data, class imbalance, SNR variation, even quantization differences — none of them break it.

Tom: And the theory backs it up. Lu, you want to give us the one-sentence version?

Lu: They proved convergence for the federated contrastive training and derived a separability condition for the downstream classifier. That's the kind of rigor that makes this more than just a clever hack.

Meng: And from a practical standpoint, the model is small enough to deploy on edge devices, and the communication cost is manageable. The compute is higher, but it's a fair trade for the accuracy gains.

Tom: Now, Lalam, you're our in-house model. What's the bigger picture here? Where does this go?

Lalam: This framework points toward a future where wireless networks are self-organizing and adaptive. Instead of needing a central authority to label every signal, devices can learn from the raw spectrum around them. That could enable smarter spectrum sharing, better interference management, and more resilient IoT networks. And the same self-supervised federated approach could apply to other domains — sensor networks, autonomous vehicles, even healthcare monitoring — wherever data is distributed and labels are scarce.

Jane: That's a nice way to think about it. The paper is really about making machine learning work in the wild, not just in the lab.

Tom: Well said. We've covered the title, the method, the experiments, and the implications. I think we've done this paper justice. Thanks to everyone for listening, and we'll see you next time with another exciting paper from arXiv.

Jane: Take care, everyone.

More episodes

← Home