A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWAN
Robert Bitterling, Christian Nettersheim, Jörn Hees, Michael Rademacher
Fraunhofer FKIE · Hochschule Bonn-Rhein-Sieg
cs.NI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Accepted at IEEE Conference on Local Computer Networks (LCN)
Code: https://github.com/mclab-hbrs/lora-bonn-ml-pathloss
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 48/100
The gist: This paper investigates how training-set size affects the accuracy of machine learning (ML) models for LoRa path loss prediction in urban Low Power Wide Area Network (LPWAN) deployments.
Terminology
Summary
This paper investigates how training-set size affects the accuracy of machine learning (ML) models for LoRa path loss prediction in urban Low Power Wide Area Network (LPWAN) deployments. The authors use real-world measurements from an urban deployment in Bonn, Germany, where uplink Received Packet Power (RPP) was recorded from outdoor end devices mounted on garbage trucks over 230 days, using nine gateways citywide at 868 MHz, 125 kHz bandwidth, and Spreading Factor (SF) 12.
After preprocessing and cleaning, the dataset was reduced to 64,294 samples across eight gateways.
The study compares two simple ML models against traditional empirical and terrain-based propagation models. The ML models are: (1) a Random Forest (RF) using LiDAR-derived terrain features (including antenna heights above mean sea level, Fresnel zone clearances, downtilt angles, azimuths, free-space path loss, and distance) generated via SignalServer
software, and (2) k-Nearest Neighbors (k-NN) using only transmitter and receiver coordinates, with k=7 and k=31 variants selected via grid search. The baseline models include Okumura-Hata, COST-231 Hata, ECC33, ITM, ITWOM, and specialized LPWAN models (Dortmund, Ghent, Oulu, and Bonn).
The primary evaluation uses a 75/25 random pooled split: a test set with a fixed size of 16,074 samples was held out, and training subsets of increasing size were drawn at random from the remaining data.
Training-set sizes ranged from 100 to 48,220 samples. The authors also performed a leave-one-gateway-out (LOGO) evaluation at the largest training size to assess transferability to unseen gateways.
Key results under random pooled evaluation show that both ML models consistently outperform the considered baseline models across the evaluated training-set sizes.
At maximum training size, they achieve RMSE values below 6.5 dB compared to 9.7 dB for the best baseline.
Specifically, at 48,220 training samples, the k-NN with k=31 achieves the lowest RMSE at 6.15 dB, followed by k-NN with k=7 at 6.32 dB and RF at 6.34 dB. The baseline models' RMSE values remain nearly constant: Bonn around 9.68 dB, COST-231 Hata at 12.39 dB, ITM at 21.75 dB, and ITWOM at 22.58 dB.
The learning curve analysis reveals that for training-set sizes between 100 and 500, the RF model consistently outperforms both k-NN models.
Up to roughly 3,250 samples, RF maintains the lowest RMSE, after which both k-NN models overtake it. At around 4,750 samples, k-NN with k=31 overtakes k=7. The authors note a clear improvement trend emerges with diminishing returns: for RF, RMSE improves by 0.62 dB from 100 to 500 samples, by 1.08 dB from 500 to 15,000 samples, and by only 0.31 dB from 15,000 to 48,220 samples.
Error distribution analysis shows that ML models have more compact, with narrower and darker central bins
error distributions compared to baselines. Among baselines, most tend to underestimate path loss, except ITM, whose distribution is centered near zero but has a wider spread.
The Bonn model's error distribution closely matches that of the ML models trained with 500 samples in both spread and central tendency.
The LOGO evaluation reveals important limitations of the pooled-split results. Both k-NN variants degrade sharply, roughly doubling their RMSE from about 6.3 dB to about 13.5 dB,
which the authors attribute to k-NN acting as a spatial lookup without nearby receiver-side neighbors for held-out gateways. The RF model degrades less, from 6.34 dB to 8.92 dB (an increase of 2.58 dB),
indicating that its terrain-derived features transfer partially to unseen gateways.
However, this aggregate hides substantial variation: For five of the eight gateways, the RMSE remains between 6.6 dB and 7.6 dB, close to the random pooled error of about 6.3 dB,
while the remaining three gateways reach 10.4 dB, 11.2 dB, and 14.2 dB. These high-error gateways stand out by their site characteristics: an unusually large antenna height of 55 m, a mean link distance of 4.9 km, and an antenna height of 30 m.
The authors provide practical guidance: "With only a few hundred labeled links, RF yields the best accuracy, at the cost of deriving terrain features. With a few thousand samples or more, coordinate-only k-NN becomes equally competitive for within-deployment interpolation while requiring only link coordinates. They caution that
beyond roughly 3,250 samples the two lie within about 0.2 dB, a gap that should be interpreted with caution given the limited hyperparameter tuning and the differing feature sets."
The paper acknowledges limitations: results are based on one urban region, environment, technology, and frequency,
and no extensive hyperparameter tuning was performed. Future work directions include designing small measurement campaigns that achieve maximum tolerable RMSE and replicating the study in other cities, frequencies, and technologies to test external validity.
Improvements for AI systems
Improvements to AI Systems:
-
Sample-Efficient Model Selection Module: Implement an adaptive model-selection mechanism that automatically chooses between Random Forest (RF) and k-NN based on available training-set size. The system would use RF when labeled samples are fewer than 3,250 and switch to k-NN (k=31) beyond that threshold, optimizing accuracy while minimizing computational overhead and feature-engineering costs.
-
Terrain-Feature Transfer Learning for Sparse Data: Enhance RF-based path loss predictors with a feature-importance weighting scheme derived from LiDAR-derived variables (e.g., Fresnel zone clearance, antenna height above mean sea level). This allows the AI to generalize to unseen gateway locations by prioritizing physically meaningful features over purely spatial correlations, reducing RMSE degradation from 6.34 dB to 8.92 dB (vs. 13.5 dB for k-NN) in leave-one-gateway-out scenarios.
-
Spatial-Awareness Regularization for k-NN: Add a gateway-identity penalty term to k-NN training that discourages reliance on receiver-side spatial lookup alone. This would force the model to incorporate secondary features (e.g., distance, azimuth) when predicting for unseen gateways, mitigating the sharp RMSE doubling observed (6.3 dB → 13.5 dB) and improving cross-deployment robustness.
-
Learning-Curve-Aware Early Stopping: Integrate a dynamic training-size scheduler that monitors RMSE improvement rates (e.g., 0.62 dB per 400 samples initially, dropping to 0.31 dB per 33,000 samples later). The system automatically stops data collection or training when marginal gains fall below a user-defined threshold (e.g., <0.1 dB per 1,000 samples), saving measurement effort and compute resources.
-
Hybrid Ensemble with Baseline Fallback: Create an ensemble that blends ML predictions with the best-performing empirical model (Bonn, RMSE 9.68 dB) when ML confidence is low (e.g., high prediction variance or when input features fall outside training distribution). This prevents catastrophic failures in unseen environments, as the Bonn model’s error distribution closely matches ML models trained on only 500 samples.
-
Site-Characteristic Anomaly Detection: Train a separate classifier to flag gateway locations with extreme features (e.g., antenna height > 55 m, mean link distance > 4.9 km) that historically cause high prediction errors (RMSE 10.4–14.2 dB). The AI system would then automatically switch to a more conservative model or request additional local calibration data for such sites.
What the Improved AI System Can Do:
-
Deploy in new cities with minimal data: Automatically choose RF with terrain features when only 100–500 labeled links exist, achieving RMSE < 8 dB without extensive measurements.
-
Scale efficiently within a deployment: Transition to coordinate-only k-NN once 5,000 samples are collected, reducing feature-engineering overhead while maintaining RMSE ≈ 6.2 dB.
-
Predict reliably for unseen gateways: Maintain RMSE < 9 dB for 5 out of 8 new gateway locations by leveraging terrain features, while flagging high-risk sites (e.g., tall antennas, long links) for manual intervention.
-
Optimize measurement campaigns: Stop data collection when learning curves flatten, saving up to 70% of measurement effort (e.g., collecting 15,000 instead of 48,220 samples) while sacrificing only 0.31 dB accuracy.
-
Provide uncertainty-aware predictions: Output confidence intervals based on training-set density and feature similarity, allowing network planners to identify areas where empirical models (like Bonn) are safer fallbacks.
Related papers
- HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation
- Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification
- EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
- SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
- What is Normal? A Big Data Observational Science Model of Anonymized Internet Traffic