Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data

arXiv:2510.04927 · cs.LG, cs.AI, eess.SP · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Federated Self-Supervised Modulation Classification under Non-IID and Imbalanced Data".

Jane: The paper was written by Usman Akram, Yiyue Chen and Haris Vikalo from University of Texas at Austin and Qualcomm Technologies Inc..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a real mouthful of a title — "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." Jane, I'm going to need you to break that down for me, because that's a lot of jargon in one sentence.

Jane: Happy to, Tom. So imagine you've got a bunch of cell phones or radios scattered around a city, and each one is picking up wireless signals. Automatic modulation classification is basically figuring out what kind of signal it is — is it a simple BPSK signal, or a more complex sixteen-QAM? That matters for managing the spectrum efficiently.

Tom: Right, and the "federated" part means we're not collecting all that raw data in one central server. Each device learns locally and only shares the model updates. That protects privacy and saves bandwidth.

Jane: Exactly. But here's the catch — the paper's title mentions "Non-IID and Class-Imbalanced Data." In plain English, that means each device sees a different slice of reality. One phone might mostly see BPSK signals, another mostly sees QPSK. And some signal types are rare. That throws traditional training methods off.

Tom: And that's where "self-supervised" comes in, right? Instead of needing tons of labeled examples — which are expensive to get for radio signals — the model learns patterns from raw, unlabeled data first. Then it only needs a tiny bit of labeled data to finish the job.

Jane: You got it. The team at UT Austin, led by Usman Akram, Yiyue Chen, and Haris Vikalo, built a system that does exactly that. They call it FedSSL-AMC. And the results are pretty impressive — on their synthetic dataset, they hit over fifty-five percent accuracy where the best supervised baseline only managed around forty-two percent.

Tom: That's a big jump. But I'm curious about the real-world part — they also tested on something called the MIGOU dataset, which is actual over-the-air signals. Jane, what happened there?

Jane: The gap got even wider. Under heavy data heterogeneity, FedSSL-AMC hit over eighty-two percent accuracy while the standard federated learning baselines were stuck in the 30s and 40s. It's a clear win for the self-supervised approach.

Tom: So the title is basically advertising the solution to a real problem — wireless networks are messy, data is uneven, and labels are scarce. This paper says, hey, we can still make this work. Stick around, because next we're going to talk about how they actually built this thing.

Summary: Tom: Welcome back. We're still on "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." Last time we covered the big picture — why this matters for wireless networks. Now let's get into the meat of it. Jane, how does this thing actually work?

Jane: So the clever part is they split the problem into two stages. First, they train a shared "encoder" — think of it as a feature extractor — using a triplet loss on unlabeled I/Q data. That's the raw in-phase and quadrature signal samples. The encoder learns to recognize patterns: "these two signal snippets look similar, those two look different."

Tom: And that's the self-supervised part. No labels needed. But what kind of neural network are they using for this encoder?

Jane: They use a causal convolutional neural network with time dilation. That's a fancy way of saying the network looks at the signal over a long window of time, but it does it efficiently. Each layer skips ahead exponentially, so it can see far back in time without needing a huge number of parameters.

Tom: And after that encoder is trained collaboratively across all the clients, each client trains its own tiny classifier on its own labeled data. That's the personalization step.

Jane: Right. They use a support vector machine — an SVM — which is lightweight and works well with small labeled sets. So the heavy lifting is done in the unsupervised phase, and the labeled data just fine-tunes the final decision.

Tom: Now, the paper also has some serious math in it. They prove that this training process converges — meaning it reliably reaches a good solution even when clients have very different data. Lu, you're our theory person. What did you make of that?

Lu: I was actually quite impressed. They establish a convergence bound for the federated representation learning, and they show that the variance of the stochastic gradients stays bounded. That's the kind of guarantee you want before deploying something in the real world. They also prove a separability condition — basically, how much signal-to-noise ratio you need for the downstream SVM to work reliably.

Tom: So it's not just "we tried it and it worked" — there's actual theory backing it up.

Lu: Exactly. And the bound they get in the high-SNR limit is clean. The average squared gradient norm goes to something like 64βM plus 2m2R2B over W, plus a small projection error term. It's a solid result.

Jane: And the practical payoff is that this approach is robust to all kinds of heterogeneity — different SNRs across clients, different carrier frequency offsets, even different quantization levels. We'll dig into those experiments next.

Tom: Great, because I want to see how it holds up when you really stress-test it. That's coming up right after this.

Improvements: Tom: We're back, still talking about "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." We've covered the setup and the theory. Now, Jane, what did they actually test, and what improvements did they show?

Jane: They ran a whole battery of experiments. First, on their synthetic dataset, they compared FedSSL-AMC against four supervised baselines: FedAVG-CNN, FedeAMC, FedProx, and FedDyn. And they also compared against a SimCSE-style contrastive baseline. Under standard label imbalance, FedSSL-AMC hit fifty-five point four one percent accuracy, while the best baseline, SimCSE, only got fifty-one point five five percent.

Tom: And when they cranked up the heterogeneity — different SNR ranges per client — the gap stayed. FedSSL-AMC got forty-one point four two percent on the worst client, while FedAVG-CNN only managed thirty-one point six four percent there.

Jane: Right. And they also tested mobility-induced carrier frequency offset. They simulated four mobility regimes, from ultra-low to high Doppler shift. Even with that extra distortion, FedSSL-AMC stayed on top.

Meng: I want to jump in here, because as an engineer, I care about whether this thing can actually run on real hardware. The paper reports the encoder has zero point two four seven million parameters — that's tiny compared to the one point seven eight million in the baseline CNN. But it does need more compute: four hundred seventy-three MFLOPs versus seventeen point seven six. That's a real tradeoff.

Tom: So it's more compute per inference, but way fewer parameters. Is that a problem for edge devices?

Meng: It's manageable. The compute is dominated by the contrastive loss during training, not inference. And the communication cost during federated training is actually lower because the model is smaller. For a radio on a drone or a sensor node, that's a win.

Jane: And on the MIGOU dataset — that's real over-the-air data — the improvements were even more dramatic. Under heavy heterogeneity, FedSSL-AMC hit eighty-two point five nine percent accuracy, while FedeAMC only got thirty-three point two zero percent. That's a massive jump.

Tom: Meng, what about the quantization experiment? They tested with clients using different numerical precisions — float32, float16, and int8.

Meng: That's the model heterogeneity test. And even with that mess, FedSSL-AMC stayed ahead of all the supervised baselines. That tells me the self-supervised representation is robust to the noise introduced by quantization, which is really useful for real deployments where devices have different capabilities.

Tom: So the improvements aren't just incremental — they're substantial across every heterogeneity scenario they threw at it.

Jane: Exactly. And that's the story of this paper. It's not just a new algorithm; it's a new way of thinking about how to train models when your data is messy and your labels are scarce.

Tom: Alright, we're going to wrap this up with our final thoughts and what this means for the future. Stay with us.

Conclusion: Tom: Alright, we're at the end of our discussion on "Federated Self-Supervised Learning for Automatic Modulation Classification under Non-IID and Class-Imbalanced Data." Jane, give us the final summary.

Jane: So the big takeaway is that by decoupling representation learning from classification — training a shared encoder on unlabeled data first, then personalizing a tiny classifier per client — you get a system that's far more robust to the messy realities of wireless networks. Non-IID data, class imbalance, SNR variation, even quantization differences — none of them break it.

Tom: And the theory backs it up. Lu, you want to give us the one-sentence version?

Lu: They proved convergence for the federated contrastive training and derived a separability condition for the downstream classifier. That's the kind of rigor that makes this more than just a clever hack.

Meng: And from a practical standpoint, the model is small enough to deploy on edge devices, and the communication cost is manageable. The compute is higher, but it's a fair trade for the accuracy gains.

Tom: Now, Lalam, you're our in-house model. What's the bigger picture here? Where does this go?

Lalam: This framework points toward a future where wireless networks are self-organizing and adaptive. Instead of needing a central authority to label every signal, devices can learn from the raw spectrum around them. That could enable smarter spectrum sharing, better interference management, and more resilient IoT networks. And the same self-supervised federated approach could apply to other domains — sensor networks, autonomous vehicles, even healthcare monitoring — wherever data is distributed and labels are scarce.

Jane: That's a nice way to think about it. The paper is really about making machine learning work in the wild, not just in the lab.

Tom: Well said. We've covered the title, the method, the experiments, and the implications. I think we've done this paper justice. Thanks to everyone for listening, and we'll see you next time with another exciting paper from arXiv.

Jane: Take care, everyone.

Usman Akram, Yiyue Chen, Haris Vikalo

University of Texas at Austin · Qualcomm Technologies Inc.

cs.LG, cs.AI, eess.SP

Submitted: 2026-08-15

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 42/100

The gist: conditions typical in AMC tasks.

Key concepts

Federated Learning
A training method where raw data is not collected in one central location. Instead, local devices learn independently and only share model updates with the central server, which protects privacy and saves bandwidth.
Self-Supervised Learning
A technique used to train models without needing many expensive labeled examples. The system learns patterns from raw, unlabeled data first (e's encoder), then uses a small amount of labeled data to finish the classification task.

Terminology

Summary

Summary

The paper introduces FedSSL-AMC, a federated self-supervised learning framework for automatic modulation classification (AMC) under non-IID and class-imbalanced data conditions. The authors motivate the work by noting that training AMC models on centrally aggregated data raises privacy concerns, incurs communication overhead, and often fails to confer robustness to channel shifts. While federated learning (FL) avoids central aggregation, the standard FedAvg algorithm presumes (near-)IID client data and degrades under non-IID distributions and class imbalance—conditions typical in AMC tasks. This is further compounded by the scarcity of labeled I/Q data and the relative abundance of unlabeled I/Q streams.

The proposed method decouples representation learning from downstream classification. Specifically, it trains a causal, time-dilated CNN with a triplet-loss self-supervision objective on unlabeled I/Q sequences across clients, followed by per-client SVMs on small labeled sets. The encoder architecture consists of stacked convolutional layers where a neuron in the k-th layer connects to previous-layer neurons with a spacing of 2(k−1), providing long-range temporal modeling, scalability, and online inference capability. The triplet loss is defined as:

L(x ref, x pos, x neg k) = −log σ(f(x ref) T f(x pos)) − Σ k=1 K log σ(−f(x ref) T f(x neg k))

where x pos is a randomly sampled subsequence within x ref, and each x neg k is a subsequence drawn from a different I/Q sequence.

The federated training procedure (Algorithm 1) follows FedAvg: in each round, clients download the global encoder parameters, update them using local unlabeled data with the triplet loss, and upload updates to the server, which aggregates them weighted by the number of local examples. After encoder training, each client fits a local SVM classifier on its labeled dataset.

Theoretical contributions include:

  1. Convergence analysis: The authors establish convergence of a time-smoothed federated representation learning procedure. They model the local contrastive loss as f c(Θ) = −(1/2)E[r T Θ T Θ r] + (λ/2)Tr(Θ T Θ), with input r = x + w′ where w′ is Gaussian noise. Theorem 1 provides a bound on the variance of stochastic gradients, which in the high-SNR and low-regularization limit simplifies to Var((∇Θ f c(Θ)) i,j) ≤ m2R2B. A second theorem bounds the average squared gradient norm of the global smoothed objective, showing convergence to a small value with appropriate choices of smoothing window size, learning rate, and SNR.

  2. Separability guarantee: Theorem 2 establishes that for a clean dataset that is (µ, ρ)-separable under the causal CNN encoder, for any ϵ > 0, there exists a threshold δ(ϵ, L) > 0 such that if the encoder output SNR γ enc > δ(ϵ, L), the noisy dataset remains (µ, ρ)-separable with probability at least 1 − ϵ.

  3. Mobility-induced frequency offset analysis: The paper provides a theoretical connection between CFO and relative motion, showing that the residual frequency offset introduced by mobility is ∆f = ν r/(f s c), which manifests as a phase rotation accumulating linearly over time.

Experimental results are reported on two datasets:

  1. Custom synthetic dataset: Generated with received baseband signal r[n] = A e(j(∆θ+2π∆f n/N)) s[n] + w[n], with four modulation types (BPSK, QPSK, 8-PSK, 16-QAM), four clients with uneven label distributions, and SNR sampled from U(−10, 10). Results show FedSSL-AMC achieves 55.41% client-averaged test accuracy, significantly outperforming baselines: FedAVG-CNN (41.61%), FedeAMC (27.34%), FedProx-CNN (40.82%), FedDyn-CNN (40.74%), and SimCSE-CNN+SVM (51.55%). The method maintains its advantage across varying label budgets (2800 to 14000 labeled examples per client), under SNR heterogeneity across clients, under mobility-induced CFO heterogeneity, and under model heterogeneity (client-specific quantization levels).

  2. MIGOU dataset: Contains over-the-air measurements from 11 modulation classes transmitted via a USRP B210 at distances of 1m and 6m (average SNRs of 37dB and 22dB). Label heterogeneity is simulated using Dirichlet sampling with concentration parameter α̃. Results across varying α̃ values (0.1 to 1.25) show FedSSL-AMC consistently outperforms all baselines, with the largest margins under high heterogeneity (e.g., 82.59% vs. 78.47% for SimCSE-CNN+SVM at α̃=0.1). In a larger-scale setting with 16 clients grouped into 5 clusters, FedSSL-AMC achieves 71.44% accuracy versus 66.14% for SimCSE-CNN+SVM and 62.76% for FedDyn-CNN.

Resource footprint analysis: The FedSSL-AMC encoder contains significantly fewer parameters than the baseline supervised CNN (0.247M vs. 1.78M) but requires more computation (473.56 MFLOPs vs. 17.76 MFLOPs) due to the contrastive loss and larger receptive field. The authors note this remains practical for edge deployment and is offset by the ability to learn from unlabeled data and communication efficiency during training.

The paper concludes that FedSSL-AMC outperforms supervised learning baselines, particularly in scenarios where unlabeled data is abundant and labels are scarce and unevenly distributed across clients. An interesting direction for future work is exploring whether clustering clients based on their data distributions can further enhance performance via group-wise contrastive learning or adaptive aggregation strategies.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:

Implementation:

  • Replace supervised FedAvg training with a two-stage pipeline: (1) unsupervised triplet-loss pretraining on raw I/Q streams using a causal, time-dilated CNN (10 layers, kernel size 3, dilation factor 2 k), and (2) per-client SVM classification on the frozen encoder's 320-dimensional features.

  • Use 10 negative samples per anchor in the triplet loss, with positive samples drawn as random subsequences within the same I/Q sequence.

  • Apply FedAvg only to the encoder weights; keep SVMs local.

Resulting Capability:

  • Achieves 55.41% accuracy on synthetic data vs. 41.61% for FedAVG-CNN and 27.34% for FedeAMC, with only 2,800 labeled samples per client.

  • Maintains performance with as few as 2,800 labels (vs. 14,000), showing 3–5× label efficiency improvement.

  • Handles non-IID label distributions (Dirichlet α=0.1) with 82.59% accuracy on MIGOU, vs. 37.19% for FedAVG-CNN.

Implementation:

  • Use the causal CNN's exponentially dilated receptive field to capture long-range temporal dependencies that are invariant to client-specific SNR, CFO, and quantization noise.

  • Train with triplet loss on unlabeled data, which aligns feature spaces across clients before any label-driven adaptation, reducing representation drift.

Implementation:

  • Apply time-smoothed stochastic projected gradient descent with exponential decay (κ→1) and window size w.

  • Use learning rate η = 1/β (inverse smoothness constant).

  • Bound gradient variance as ν2 = m2R2B in the high-SNR, low-regularization limit.

Implementation:

  • Use only 10 communication rounds (vs. 1,000 for baselines) with 2,500 local steps per round.

  • Encoder has 0.247M parameters (vs. 1.78M for FedeAMC's CNN), reducing uplink/downlink payload per round by 86%.

Implementation:

  • Use causal convolutions (parallelizable, O(χΨΛ) complexity) instead of transformers (O(ΨΛ2)) or RNNs (sequential).

  • Keep SVM as a lightweight classifier (no additional neural network layers).

  1. Learn from unlabeled wireless signals across distributed clients without sharing raw data, using only 10% of the labels required by supervised methods.

  2. Maintain high classification accuracy (70–83% on over-the-air MIGOU data) even when clients have highly skewed label distributions (Dirichlet α=0.1–0.5) and heterogeneous channel conditions.

  3. Guarantee convergence to a bounded regret under non-IID data, noisy inputs, and client drift, with formal bounds on gradient variance and downstream separability.

  4. Operate in real-time on edge devices with limited compute, memory, and bandwidth, thanks to the causal CNN's linear complexity and 99% communication reduction.

  5. Adapt to new clients by fitting a local SVM on a small labeled set (e.g., 200–2,800 samples) without retraining the shared encoder, enabling rapid personalization.

Abstract

Training automatic modulation classification (AMC) models on centrally aggregated data raises privacy concerns, incurs communication overhead, and often fails to confer robustness to channel shifts. Federated learning (FL) avoids central aggregation by training on distributed clients but remains sensitive to class imbalance, non-IID client distributions, and limited labeled samples. We propose FedSSL-AMC, which trains a causal, time-dilated CNN with triplet-loss self-supervision on unlabeled I/Q sequences across clients, followed by per-client SVMs on small labeled sets. We establish convergence of the federated representation learning procedure and a separability guarantee for the downstream classifier under feature noise. Experiments on synthetic and over-the-air datasets show consistent gains over supervised FL baselines under heterogeneous SNR, carrier-frequency offsets, and non-IID label partitions.

Sources

Related papers