Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance

arXiv:2608.12929 · cs.LG, cs.NI · Submitted 2026-08-13 · Read on arXiv

Chukwunonso Henry Nwokoye, Blessing Oluchi Iloka, Chikwue V. Umeugoji, Christopher Anene Egemba, Nnenna D. Duroha

York University · University of Hertfordshire · Joint Admissions and Matriculation Board · Lightenet Technologies Limited

cs.LG, cs.NI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 9 pages

Code: https://github.com/ChiNonsoHenry16/Multi-perspectiveImbalace-Aware-6G-Beatforming-Optimization

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 51/100

The gist: The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches.

Terminology

Summary

The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches. We compared the predictive power of network, environmental, device, and vision feature groups for 6GBO. Additionally, it addressed other unsupervised perspectives that can enhance 6GBO, including clustering network scenarios using methods such as K-means, DBSCAN, and hierarchical clustering. Several imbalance-aware experiments revealed that network features possess better prediction power than device, environmental and vision feature groups, as evidenced by their recall, F1-score and ROC-AUC values. For unsupervised ML exploration (assessed using Elbow, Silhouette score and Davies Bouldin Index methods), the results indicate that the deployment environment and type of device primarily influence clustering, rather than mobility-based attributes. Furthermore, the explainability analysis showed that bandwidth, IoT sensors, and mobility possess higher global feature importance across the feature groups. In the future, we would apply deep and reinforcement learning techniques to predict throughput/latency or to optimize rewards determined by performance indicators like SNR enhancement.

The study aims to conduct a comparative analysis of 6G beamforming optimization (6GBO) using network, environmental, device and vision factors with ML classifiers. The 6G IoT intelligent management dataset used for the study was accessed from Kaggle. Input features for the study include network (frequency (GHz), transmit power (dBm), bandwidth (MHz), codebook size, and interference level (dB)); environment (obstacle density, mobility (m/s), and environment type (outdoor)); and device (number of antennas and device type (IoT sensor and smartphone)). These input features are available before the occurrence of optimization. While running the experiments, the features were defined as groups, i.e., network parameters (NP), environmental factors (EF), device characteristics (DC), vision attributes (VA), and all features (AF). Note that VAs are Scale-Invariant Feature Transform (SIFT) keypoints, which are also part of the 6G IoT intelligent management dataset. Essentially, the target variable is the column called optimized, which contains 0s and 1s for unsuccessful and successful 6GBO. To evaluate the performance of the models, the following metrics were employed: precision (Pre), recall (Rec), F1, ROC-AUC, and PR-AUC.

Prior to model training, missing values were replaced with the median and standardized, whereas categorical variables were substituted with the mode and subjected to one-hot encoding. Boolean attributes were transformed into floating-point representations. Additionally, a stratified division of 80:20 for training and testing was implemented with a constant random seed (42) to maintain class distributions. On class balance, the positive rate is 0.168 (i.e., 168 positives out of 1000). In every aspect of preprocessing—such as feature normalization, median imputation, one-hot encoding, Boolean feature transformation, and SMOTE oversampling (when applicable)—were integrated into an imblearn pipeline and implemented separately within each cross-validation training fold, consequently averting the issue of data leakage. A 5-fold cross-validation (CV) technique was utilized for out-of-fold (OOF) threshold tuning (TT). The aggregated OOF prediction probabilities were utilized to establish the ideal classification threshold, which improved the F1-score in a situation of significant class imbalance. GridSearchCV (GSCV) was utilized largely for the optimization of hyperparameters and the selection of models by determining the optimal parameter combinations according to cross-validation scoring measures.

The initial experiment reveals that class imbalance dominates the observed performance. The dataset's positive rate (optimized = 1) is small, so many models converge to predicting the majority (non-optimized) class. As a consequence, overall accuracy is high (≈0.83) but provides a misleading impression of success: several classifiers (notably those trained only on device-level features) report precision = 0, recall = 0, and F1 = 0, indicating that they never detect any positive examples. ROC-AUC values across the experiments lie near random to weakly informative (approximately 0.44–0.57), with the best AUCs reaching roughly 0.56–0.57 for a few models trained on environmental or network features. Correspondingly, only a handful of models (e.g., RF and XGBoost on network/environmental features) produce non-zero recall and modest F1 scores (examples: Environmental RF F1 ≈ 0.21, Environmental XGBoost F1 ≈ 0.15, and Network XGBoost F1 ≈ 0.08). These non-zero values indicate there is some discriminative signal in network and environmental features, but it is weak and insufficient for reliable detection with the current data and training setup.

For Model 1 (Pre-optimization Factors Only), at a 0.5 threshold, the network and environmental groups demonstrated a comparably better performance, with AdaBoost and XGBoost attaining the highest F1 values of 0.327 and 0.324, respectively. After TT, recall significantly enhanced every FG, with some models attaining recall values around or equal to 1.0. LightGBM achieved the highest F1-score of 0.301 for NP, while SVM (RBF) attained the same score of 0.301 for vision attributes. The results demonstrate that network attributes offer strong predictive signals for 6GBO, whereas environmental, vision, and device features supply supplementary data for ML classification. No individual classifier consistently excelled throughout all feature domains, indicating significant disparities in the fundamental data distributions across the feature spaces. TT significantly enhanced recall among all FG, affirming the significance of imbalance-conscious classification for 6GBO predictions. NP consistently demonstrated a comparably better overall performance, whereas environmental and device features yielded identical F1-scores, suggesting that contextual and device-specific information significantly contributes to beamforming efficacy. The utilization of all features did not substantially exceed the performance of network parameters alone, indicating that supplementary feature groups offer minimal additional predictive capability. Various feature groups preferred distinct ML techniques, with LightGBM, XGBoost, LR, MLP, and SVM (RBF) identified as the most effective models for specific feature domains.

For the GridSearchCV-tuned models, across all groups, DC attained the best performance, with AdaBoost attaining the highest F1 score of 0.307 and an ROC-AUC of 0.564. The SVM (RBF) classification algorithm attained the highest F1 score of 0.256 for NP, while LR excelled for EF with a 0.266 F1 score. Considering the VA category (with SIFT keypoints), AdaBoost attained the highest F1-score of 0.240. These findings indicate that device-related characteristics yield a better predictive signal for 6GBO, but environmental, network, and VA-based parameters demonstrate moderate predictive efficacy when assessed in isolation. While the highest standalone F1 value in the GridSearchCV studies was attained by AdaBoost utilizing device-related characteristics (F1 = 0.307), network attributes consistently yielded superior ROC-AUC and PR-AUC values across many experimental configurations. Thus, NP-related properties seem to yield better discriminatory information for predicting 6GBO, whereas device characteristics present a desirable precision-recall equilibrium at specific classification levels.

For the ensembles (Model 1: Ensembles), at the default TT of 0.5, numerous models attained commendable accuracy levels, especially within the AF and NP categories; yet, recall and F1-scores were rather low, signifying inadequate identification of the minority class due to the imbalance. Following TT, recall significantly enhanced across the majority of feature groups, with numerous models attaining recall values around or equal to 1.000. The network FG exhibited the best predictive performance, with the VE attaining the largest ROC-AUC and PR-AUC of 0.585 and 0.224, respectively. Also, the SE for network parameters recorded the highest optimized F1-score (0.302). The environment and device features yielded a somewhat moderate performance, whereas the vision (SIFT) attributes exhibited relatively poor discriminative ability. The findings indicate that network-related characteristics are the primary contributors to the prediction of adaptive and dynamic 6GBO. Additionally, beyond the ML-based 5G beam selection analysis in Klautau, et al., the ensembles (SE and VE) in the study achieved higher accuracies. However, it is noteworthy that both studies were conducted using different datasets. Interestingly, the device attributes achieved the strongest F1-score (0.307) when GridSearchCV was employed.

For Model 2 (Performance Metrics Alone), the objective was to assess the extent to which post-hoc outcome factors independently forecast the optimized label. The performance metrics include the following: Beamforming Gain, Latency, Energy Consumption, Throughput, Beam Training Time, SNR Improvement, Processing Time, and Memory Usage. The findings demonstrated good classification performance across almost all ML models, with GB, AdaBoost, DT, HG, XGBoost, and LightGBM attaining excellent scores of 1.000 for all evaluation metrics. Likewise, RF attained nearly flawless performance, exhibiting an accuracy and ROC-AUC of 0.995 and 1.000, respectively. Even relatively worse models like KNN and logistic regression yielded impressive ROC-AUC values of 0.935 and 0.944, respectively. The findings indicate a good correlation between the performance measures as well as the target label, thereby providing significant post-hoc information regarding beamforming performance. Consequently, employing these factors for real-time forecasting will artificially enhance classification performance and inadequately represent actual decision-making scenarios in adaptive 6GBO. To a large extent, the results thus confirm the necessity of limiting deployable prediction pipelines of 6GBO to pre-decision (pre-optimization) data.

For Model 3 (All Features), the objective was to assess the impact of integrating pre-decision attributes with post-optimization result metrics to examine the effects of target leakage. Multiple ensemble models, notably GB, AdaBoost, DT, XGBoost, HGB, and LightGBM, attained excellent performance for every evaluation metric (1.000). Likewise, RF attained an exceptional result with accuracy and ROC-AUC of 0.985 and 1.000, respectively. Classical ML models, including LR and SVM (RBF), exhibited good prediction performance, attaining ROC-AUC values over 0.93. While these results first imply a highly superior prediction of 6GBO, the perfect performance clearly suggests the existence of target leakage inside the feature space. The incorporation of post-hoc performance indicators furnishes the models with outcome-related data that is inherently linked to the optimized label. The decision to remove SMOTE balancing and TT is to clearly show that this model achieved perfect performance without these imbalance-aware approaches. Thus, the classifiers are proficiently learning optimization results instead of forecasting circumstances based on actionable pre-decision parameters. The entire feature configuration yields remarkably high classification values; nevertheless, these outcomes do not accurately represent practical deployment scenarios, as several included metrics are accessible only post-beamforming optimization.

For explainability using SHAP, the network FG performed better than other groups: AdaBoost (F1 = 0.327), LightGBM (F1 = 0.327), and SVM (RBF) (F1 = 0.256). The AdaBoost model was employed for explainability. Among these synergistic explainability methods, codebook size, bandwidth, and frequency repeatedly surfaced as the most significant determinants of 6G beamforming enhancement. PI revealed Codebook Size (0.042) and Bandwidth (0.033) as the paramount attributes; however, AdaBoost's MSFI prioritized Codebook Size (0.304), Transmit Power (0.286), and Bandwidth (0.207). Conversely, SHAP analysis for global importance revealed bandwidth (0.073), frequency (0.029), and codebook size (0.020) as the primary influences. The persistent significance of the network feature group through various explainability techniques suggests that codebook arrangement, bandwidth availability, and operating frequency are crucial in influencing 6GBO decisions. For the DC explainability model, the most significant feature was actually Device Type (IoT Sensor) with a SHAP value of 0.035, followed by Device Type (Smartphone) at 0.025 and Number of Antennas at 0.023, suggesting that device type exerts a more substantial influence on 6GBO than the number of antennas. In the EF explainability model, Mobility was identified as the primary predictor (SHAP = 0.177), succeeded by Obstacle Density (0.099) and Environment Type (Indoor/Outdoor) (0.033). These results indicate that, in terms of XGBoost, user mobility and environmental factors have a more significant impact on device-specific traits than do device-specific traits, with mobility being the paramount environmental element. The GB model utilizing the Model 2 (Performance Metrics) attained flawless classification results. PI, MSFI, and SHAP analysis consistently recognized latency, throughput, and beamforming gain as the primary predictors, but the other performance metrics had minimal influence on model decisions. Latency demonstrated the greatest significance among all three techniques, succeeded by Throughput and Beamforming Gain. The nearly flawless prediction accuracy indicates a correlation between these variables and the optimization objective, serving as post-decision performance metrics, which results in target leakage when employed as predictive attributes.

For clustering network scenarios, unsupervised ML clustering approaches were employed to elucidate environmental and operational changes in 6GBO systems with the goal of aiming to identify various profiles of network scenarios. The analysis employed essential contextual characteristics, such as mobility, obstacle density, device type, and environment (outdoor), which jointly define evolving communication contexts and user scenarios. Three clustering methods—K-means, DBSCAN, and hierarchical clustering—were applied to categorize network scenarios based on feature similarity and spatial density correlations. For the initial set of features, k=4 was preferred, as it aligned with the elbow point, preserved a commendably high SS, and produced comprehensible operational scenarios. The SS analysis for the SelectKBest-based feature set revealed that k=2 yielded the most pronounced cluster distinctness; hence, it was used for the unsupervised analysis. K-Means and hierarchical clustering demonstrated superior performance for the initial features, attaining an SS of 0.352 at k=4, whereas the SelectKBest-based features yielded a moderately distinguishable 2-cluster arrangement at k=2 with an SS of 0.221. With the performance-based parameters, DBSCAN failed to identify any clusters, categorizing each of the 1000 data points as noise. The examination of the cluster centroids identified four unique deployment scenarios, largely distinguished by environmental type and device category. The clusters precisely corresponded to (cluster 0) outdoor non-IoT devices, (cluster 1) indoor non-IoT devices, (cluster 2) indoor IoT sensors, and (cluster 3) outdoor IoT sensors. Conversely, OD and mobility demonstrated minimal fluctuations among clusters, indicating that the clustering pattern was primarily influenced by the installation environment and type of device instead of mobility-based attributes. Obstacle density, while theoretically pertinent to mmWave obstruction, was uninformative in this dataset. Its dynamic range may be insufficient, inaccurately quantified, or associated with other variables that have already accounted for its impact. The unsupervised structure within the dataset is predominantly contextual as opposed to performance-oriented. Segmenting initially by environment and device type produces stable, interpretable cohorts; within each cohort, performance exhibits smooth variation and is more effectively modeled using regression or soft clustering techniques. Context serves as the primary determinant of separable network/user states, whereas KPI variables constitute a continuum that is unable to inherently divide into discrete clusters. For the performance-based features, two moderately distinguishable operating profiles (OPs) were observed. Cluster 0 shows a non-smartphone profile, defined by reduced latency (5.41 ms), elevated throughput (549.28 Mbps), enhanced beamforming gain (17.80 dB), and superior SNR enhancement (12.69 dB). Conversely, Cluster 1 aligns with a smartphone-centric profile, demonstrating marginally elevated latency (5.47 ms), diminished throughput (526.36 Mbps), worse beamforming gain (17.19 dB), and a lesser SNR enhancement (12.42 dB). Note that the device (smartphone)-type split is clean (0.00 vs. 1.00). These results suggest that the chosen performance and device-specific attributes inherently divide the 6G network into two OPs with modest differences, offering a comprehensible post-hoc analysis of beamforming optimization performance.

Finally, it is noteworthy that TT significantly enhanced recall, with multiple models attaining values near 1.000, signifying that almost all 6GBO prospects were identified. Nonetheless, this enhancement was achieved at the cost of accuracy, which remained comparatively low, around 0.17–0.19. This indicates that numerous samples identified as needing optimization were incorrect positives, leading to an elevated false-alarm rate. In adaptive beamforming, such erroneous detections could lead the network to inappropriately initiate beam reconfiguration and training, codebook searches, or other optimization processes for links that do not need modification or adjustment. These additional beam updates may elevate computational demands, energy usage, signaling traffic, and processing delays, thereby diminishing overall system efficiency. In extremely dynamic 6G settings, emphasizing recall could be advantageous when the expense of overlooking a legitimate 6GBO opportunity surpasses that of executing an unwarranted update. A false negative can result in diminished link quality, lowered throughput, and heightened communication latency. Thus, the criterion for decision-making ought to be contingent upon the application and must weigh the operational expenses of superfluous beam adjustments against the potential for missed 6GBO prospects.

The paper concludes that the study presents a systematic ML study of beamforming optimization for 6G-IoT, comparing the predictive power of network, environmental, and device feature groups and demonstrating how ML pipelines (with explainability and imbalance handling) can guide adaptive, low-latency beam management. Generally, the network parameters have stronger F1-score and ROC-AUC/PR-AUC values. However, the device attributes achieved the strongest F1-score when GridSearchCV was employed. Despite the substantial enhancement in detecting optimized 6G beamforming scenarios through threshold tuning, the resultant rise in false positives underscores the necessity for tailored threshold preference that reconciles detection efficacy with operational demands. In the future, the authors would compare pre-decision feature groups against a simple majority-class baseline and a calibrated probability baseline. Also, their comparative analysis would extend to predicting beamforming gain, latency, energy consumption, throughput, and SNR improvement using deep and reinforcement learning techniques.

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Imbalance-Aware Classification with Threshold Tuning: Implement out-of-fold threshold tuning (TT) within imbalanced pipelines to maximize recall for minority-class detection (e.g., identifying successful beamforming opportunities). The improved AI system can dynamically adjust decision thresholds based on application-specific costs (e.g., false alarms vs. missed optimizations), reducing unnecessary beam reconfigurations while capturing nearly all true optimization cases.

  2. Feature-Group-Specific Model Selection: Train separate classifiers for distinct feature groups (network, environmental, device, vision) and select the optimal algorithm per group (e.g., LightGBM for network features, AdaBoost for device features). The improved system can automatically route input data to the best-performing model for that feature subset, maximizing F1-score and ROC-AUC across heterogeneous data sources.

  3. Target Leakage Prevention via Pre-Decision Feature Filtering: Restrict predictive models to pre-optimization features only (e.g., frequency, bandwidth, codebook size) and exclude post-hoc performance metrics (latency, throughput, SNR gain) to avoid artificial performance inflation. The improved system can detect and flag target leakage in training data, ensuring real-world deployability where outcome metrics are unavailable at decision time.

  4. Explainability-Driven Feature Prioritization: Use SHAP, permutation importance, and mean squared feature importance to rank features (e.g., codebook size, bandwidth, frequency) and prune low-value inputs. The improved system can generate transparent, actionable insights for network operators, highlighting which tunable parameters (e.g., bandwidth allocation) most strongly influence beamforming success, enabling proactive configuration.

  5. Context-Aware Clustering for Scenario Profiling: Apply K-means or hierarchical clustering on environmental and device attributes (e.g., environment type, device type) to identify distinct operational profiles (e.g., indoor IoT sensors vs. outdoor smartphones). The improved system can segment users or network regions into homogeneous groups, then apply group-specific optimization policies—reducing computational overhead by avoiding one-size-fits-all beam management.

  6. Unsupervised Anomaly Detection for Degraded Scenarios: Use DBSCAN or density-based clustering on performance metrics to flag outlier network states that do not conform to typical operating profiles. The improved system can detect abnormal beamforming conditions (e.g., unexpected latency or SNR drops) in real time, triggering targeted diagnostics or fallback mechanisms.

  7. Cost-Sensitive Decision Framework: Integrate a tunable cost matrix that weighs false positives (unnecessary beam updates) against false negatives (missed optimizations) based on operational context. The improved system can adapt its classification threshold in real time—e.g., favoring recall in high-mobility, low-latency-critical environments, while favoring precision in stable, resource-constrained settings—balancing energy efficiency and link quality.

  8. Post-Hoc Performance Monitoring with Leakage-Aware Validation: Build a two-tier evaluation pipeline: one for pre-decision prediction (using only pre-optimization features) and another for post-hoc performance assessment (using outcome metrics). The improved system can continuously validate that deployed models do not inadvertently learn from future information, maintaining honest performance estimates and avoiding overconfident deployment.

  9. Hierarchical Model Ensembling: Combine predictions from feature-group-specific models (e.g., network-only, device-only) using stacking or voting, with meta-learners that weight each group’s contribution based on context. The improved system can dynamically fuse heterogeneous data sources (e.g., network stats + device type) to improve robustness when individual feature groups are weak or noisy.

  10. Reinforcement Learning for Adaptive Beam Management: Replace static classifiers with RL agents that learn optimal beamforming actions (e.g., codebook selection, power adjustment) by maximizing long-term rewards tied to SNR improvement or throughput. The improved system can continuously adapt to changing environmental conditions (mobility, obstacle density) without retraining, reducing latency and energy consumption compared to rule-based or ML-classification approaches.

Related papers