A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks".
Jane: The paper was written by Kostas Hatalis, Alberto J. Lamadrid, Katya Scheinberg and Shalinee Kishore from Lehigh University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper with a real mouthful of a title: "A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks."
Jane: And I promise, Tom, it's a lot more exciting than that title sounds. This is about wind power forecasting, which is one of those areas where getting the uncertainty right matters just as much as getting the point prediction right.
Tom: Right, because if you're a grid operator and you only know the expected wind output, you're flying blind. You need to know the range of possibilities.
Jane: Exactly. And the authors, Kostas Hatalis and his team at Lehigh University, they're tackling a specific problem called the quantile crossover problem. That's when you try to predict multiple quantiles at once, and the lower ones end up above the higher ones.
Tom: Which makes no sense mathematically. If you're saying there's a ten percent chance wind output will be below X, and a twenty percent chance it'll be below Y, then Y has to be bigger than X. Otherwise your probabilities are just nonsense.
Jane: Right. And that's been a persistent headache for anyone doing quantile regression with neural networks. The standard approach just doesn't guarantee that ordering.
Tom: So what did they do differently? I'm guessing the "smooth" part of the title is the key.
Jane: It is. The pinball loss function, which is what quantile regression uses, has a sharp kink at zero. It's not differentiable there, which makes gradient-based training of neural networks really awkward.
Tom: So they smoothed it out. Like rounding off the sharp corner so the math flows better.
Jane: Precisely. They use a logistic-based approximation that's smooth everywhere, so you can train the network with standard backpropagation and the Adam optimizer. And then they add a penalty term that kicks in whenever quantiles cross, pushing them back into the right order.
Tom: That's clever. You're not just hoping the network learns the ordering on its own. You're actively punishing it when it gets the order wrong.
Jane: And the results, at least on the GEFCom2014 wind data, show that this approach beats linear quantile regression and support vector quantile regression, and it's competitive with the top teams in that forecasting competition.
Tom: So we're talking about real-world impact here, not just academic neatness. When you're integrating more wind into the grid, having reliable prediction intervals can make the difference between a smooth operation and a costly emergency.
Jane: Exactly. And that's what we'll dig into more as we go through the paper. But first, let's just appreciate the fact that they got their quantile scores down to around zero point zero four two, which is right up there with the winning team's zero point zero three eight.
Tom: And they did that with raw wind speed data and time features, not the dozens of hand-crafted features the competition winners used. That's a pretty strong statement about the model's ability to learn on its own.
Jane: It really is. So stick around, because next we're going to look at how they actually built this network and why the smooth approximation matters so much in practice.
Summary: Tom: So Jane, we've established that this paper, "A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks," is about making wind power forecasts more useful. But let's get into the nitty-gritty of what they actually did.
Jane: Let's do it. So the core idea is pretty straightforward. They built a feedforward neural network that outputs multiple quantiles at once. Instead of training one model for the ten percent quantile, another for the twenty percent, and so on, you train a single network that gives you all of them simultaneously.
Tom: And that's where the crossover problem usually shows up. When you train each quantile separately, nothing stops the ten percent quantile from being higher than the twenty percent quantile at some time step.
Jane: Right. And the authors' solution has two parts. First, they replace the pinball loss with a smooth approximation. The pinball loss has that kink at zero that I mentioned, and that kink makes gradient descent tricky.
Tom: Because the gradient isn't defined at that point. You'd have to use subgradients or some workaround.
Jane: Exactly. Their smooth version, which they call Sτ,α, is a logistic-based function that gets closer and closer to the true pinball loss as a smoothing parameter α goes to zero. But for any positive α, it's differentiable everywhere.
Tom: So you can just train the network like any other neural network. No special tricks needed.
Jane: And then the second part is the penalty term. They add a term to the loss function that measures how much the quantiles cross. If the ten percent quantile is above the twenty percent quantile, the penalty is proportional to the square of that difference.
Tom: So the network gets penalized for crossing, and over time it learns to keep the quantiles in order.
Jane: Exactly. And they set the penalty parameter pretty high, one thousand in their experiments, so the network really doesn't want to violate that ordering constraint.
Tom: Now, I've seen other approaches to this problem. Some people just sort the quantiles after the fact, which feels like cheating.
Jane: It is a bit of a hack. The authors point out that reordering doesn't have a strong theoretical foundation. You might fix the ordering but mess up the actual quantile values in the process.
Tom: And there are also approaches that add constraints to the network weights, like the monotonic composite QRNN that's mentioned in the paper. But that adds complexity and more parameters.
Jane: Right. The penalty approach is simpler. You just add one term to your loss function and let the optimizer figure it out. No architectural changes, no extra constraints on the weights.
Tom: That's the kind of solution that actually gets adopted in practice. If you can implement it with a few lines of code in your existing training loop, people will use it.
Jane: And that's what makes this paper valuable. It's not just a theoretical contribution. It's a practical tool that forecasters can pick up and use.
Tom: So next, let's talk about what improvements they actually saw in their experiments. Because a clever idea is one thing, but does it actually work?
Improvements: Tom: So Jane, we've talked about the method. Now let's talk about the results. The paper "A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks" reports some pretty impressive numbers.
Jane: They do. And the first thing that stands out is the comparison with support vector quantile regression, or SVQR. That's a nonlinear method that's been used for this kind of forecasting.
Tom: And it didn't do well here, right?
Jane: Not at all. In the first case study, SVQR had coverage errors as bad as negative forty percent. That means its prediction intervals were way too narrow. It was overconfident in a bad way.
Tom: So it was saying, "I'm ninety percent sure the wind output will be in this range," and then the actual output fell outside that range forty percent of the time more than it should have.
Jane: Exactly. And the authors attribute that to overfitting. SVQR with a radial basis function kernel just couldn't extract meaningful features from the raw wind data.
Tom: Meanwhile, their smooth pinball neural network, SPNN, had coverage errors in the range of negative three percent to zero point three percent. That's a huge improvement.
Jane: And it's not just about coverage. They also looked at sharpness, which is how narrow the prediction intervals are. You want them narrow enough to be useful, but wide enough to actually capture the observations.
Tom: The Goldilocks problem.
Jane: Exactly. And SPNN found that balance. Quantile regression had wider intervals, which is why it had decent coverage but was less useful. SVQR had narrower intervals, but they were too narrow to be reliable.
Tom: So SPNN sits in that sweet spot. And when they looked at the quantile verification skill score, which measures improvement over a reference model, SPNN was consistently positive.
Jane: Right. They used linear quantile regression as the reference, and both SPNN1 and SPNN2, which are the one-hidden-layer and two-hidden-layer versions, showed clear improvements. SVQR, on the other hand, had negative skill scores, meaning it was actually worse than the linear baseline.
Tom: That's a pretty damning result for SVQR. But it also highlights that the smooth approximation is doing real work here.
Jane: And then in the second case study, they scaled up to ninety-nine quantiles across all ten wind farms in the GEFCom2014 dataset. That's eighty-seven thousand six hundred test observations total.
Tom: That's a serious test. And what did they find?
Jane: SPNN2 had a mean quantile score of about zero point zero four two. The winning team in the competition, kPower, had zero point zero three eight. So they're close, and they got there without all the feature engineering the competition winners used.
Tom: That's remarkable. The competition winners spent months crafting features like wind shear and direction differences between heights. And SPNN just took the raw wind speed components and time features and got within striking distance.
Jane: And that's the story of this paper. It's not about squeezing out every last bit of performance through clever feature engineering. It's about building a model that can learn the patterns on its own, while guaranteeing that the quantiles stay in the right order.
Tom: So the improvement isn't just in the numbers. It's in the simplicity and robustness of the approach.
Jane: Exactly. And that's what makes it a practical contribution. But let's dig into the first page of the paper next, because there's some context there that's worth unpacking.
First Page: Tom: So Jane, we've talked about the results, but let's go back to the beginning. The first page of "A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks" sets up the problem really well.
Jane: It does. And the key point is that wind power has grown rapidly over the last thirty years, and in some countries it's now the most used form of renewable energy. But that growth brings challenges.
Tom: Because wind is chaotic. It's not like a coal plant where you can just dial up the output. The weather drives everything, and the weather is unpredictable.
Jane: Right. And the authors list some of the specific problems. Grid operators need to manage power flow, operating reserves, and unit commitment. If you don't know how much wind is coming, you have to keep more reserves spinning, which costs money.
Tom: And for wind farm operators, they need forecasts for bidding into energy markets. If they overpromise, they get penalized. If they underpromise, they leave money on the table.
Jane: Exactly. And that's where probabilistic forecasting comes in. Instead of just saying "we expect fifty megawatts," you say "there's a ninety percent chance it'll be between thirty and seventy megawatts."
Tom: And that range is what the paper is all about. Getting those ranges right, and making sure they're consistent.
Jane: The authors also mention that traditional approaches often assume a specific distribution, like the Beta distribution, for wind power errors. But that assumption doesn't hold up for short-term forecasting.
Tom: Because wind power is bounded between zero and capacity, and it's often skewed. The errors aren't symmetric.
Jane: Right. So they argue for a nonparametric approach, where you don't assume any particular shape for the distribution. You just estimate the quantiles directly.
Tom: And that's what quantile regression does. But the authors point out that standard quantile regression has that crossover problem when you try to estimate multiple quantiles.
Jane: Which brings us back to their contribution. They're providing a way to do nonparametric probabilistic forecasting with a neural network, without the crossover problem, and with a loss function that's amenable to gradient-based training.
Tom: And they're doing it in a way that's practical. The first page also mentions that their method is evaluated on the GEFCom2014 dataset, which is the standard benchmark for this kind of work.
Jane: So it's not just theoretical. They're showing it works on real data from ten wind farms.
Tom: And that's what gives me confidence that this could actually be adopted. When you can show a clear improvement over existing methods on a public benchmark, people pay attention.
Jane: Absolutely. And the fact that they're competitive with the top teams in the competition, without all the feature engineering, is a strong signal that the method itself is doing the heavy lifting.
Tom: So we've covered the problem setup, the method, and the results. Let's wrap this up and think about what it all means.
Conclusion: Tom: So Jane, we've spent this whole episode on "A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks." Let's pull it all together.
Jane: Let's do it. The paper tackles a real problem in wind power forecasting: how to get reliable prediction intervals that don't violate basic mathematical consistency.
Tom: And their solution has two parts. A smooth approximation to the pinball loss that makes neural network training straightforward, and a penalty term that keeps quantiles from crossing.
Jane: The results on the GEFCom2014 data show that this approach beats linear quantile regression and support vector quantile regression, and it's competitive with the top teams in the competition.
Tom: And it does that with minimal feature engineering. Just raw wind speed components and time features.
Jane: Right. The model learns the patterns on its own, which is a testament to the power of neural networks when you give them a loss function that works.
Tom: Now, what does this mean for the world? I think the biggest impact is on grid operations. As we integrate more wind power, having reliable uncertainty estimates becomes critical.
Jane: And it's not just wind. The authors mention that the method could be applied to solar power, wave power, electricity pricing, and load forecasting. Anywhere you need quantile estimates.
Tom: So it's a general tool, not just a wind power tool.
Jane: Exactly. And the fact that it's implemented with standard gradient descent and the Adam optimizer means it's easy to adopt. You don't need specialized solvers or complex constraint handling.
Tom: That's the kind of thing that gets picked up by practitioners. If it's easy to implement and it works better, people will use it.
Jane: And the authors also mention future work. They want to test it on very short-term forecasting using only past wind power data, without the numerical weather predictions.
Tom: That would be interesting. Because sometimes you don't have access to weather forecasts, especially for small wind farms.
Jane: Right. And they also want to expand the model to provide full predictive densities, not just quantiles. That would be the next step in making these forecasts even more useful.
Tom: Well, I think we've given this paper a thorough look. It's a solid contribution with practical implications.
Jane: And it's a great example of how a small mathematical tweak, smoothing a loss function, can have a big impact on real-world forecasting.
Tom: Alright, that's a wrap on this one. Thanks for joining us, and we'll see you next time for the next paper.
Jane: Take care, everyone.
Kostas Hatalis, Alberto J. Lamadrid, Katya Scheinberg, Shalinee Kishore
Lehigh University
eess.SP, cs.LG, econ.EM
Submitted: 2019-09-24
Comments: 12 pages, IEEE Transactions Journal format. arXiv admin note: substantial text overlap with arXiv:1710.01720
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
Key concepts
- Quantile Crossover Problem
- This occurs when trying to predict multiple quantiles simultaneously, where the lower quantiles end up being higher than the higher ones. This is mathematically nonsensical because probabilities must be ordered correctly.
- Pinball Loss Function
- This is a loss function used in quantile regression that has a sharp kink at zero, making it non-differentiable at that point. The authors smoothed this function using a logistic-based approximation to allow for easier training with gradient descent.
- Smoothed Loss and Penalty Function
- The solution involves replacing the sharp pinball loss with a smooth, differentiable version (Sτ,α). Additionally, they add a penalty term to the loss function that increases if the predicted quantiles cross each other, forcing the network to maintain correct ordering.
- Coverage Errors
- This measures how often the actual wind output falls outside of the predicted prediction interval. The paper shows that their method significantly reduced coverage errors compared to support vector quantile regression, indicating more reliable prediction intervals.
Terminology
Summary
Summary
This paper, titled A Novel Smoothed Loss and Penalty Function for Noncrossing Composite Quantile Estimation via Deep Neural Networks,
proposes a new method for nonparametric probabilistic forecasting of wind power, addressing the challenge of uncertainty quantification in renewable energy integration. The authors introduce a neural network model called the smooth pinball neural network
(SPNN) that estimates multiple quantiles simultaneously while preventing the quantile crossover problem,
where lower quantiles overlap higher ones, violating the monotonicity principle of cumulative distribution functions.
The paper's motivation is that point forecasts lack uncertainty information, and traditional parametric assumptions (e.g., Beta distribution) may not hold for short-term wind power forecasting. The authors state: "To address the problem of dealing with nonlinearity in wind data, we propose a novel neural network model which we call the smooth pinball neural network (SPNN). This network is able to provide probabilistic forecasts in the form of multiple monotonically increasing quantiles estimated simultaneously."
The main contributions are explicitly listed:
-
We propose and investigate a new objective function which is a logistic based smooth approximation of the pinball loss function for multiple quantile regression.
-
We introduce a smooth penalty scheme to prevent the quantile crossover problem.
-
We showcase how a multiple quantile based neural network can be used for probabilistic forecasting of wind.
-
We design experiments to validate our model using publicly available data from 10 wind farms from the Global Energy Forecasting Competition 2014 and benchmark performance with common and advanced methods.
-
We show our method improves the skill, reliability, and sharpness of forecasts over various benchmarks.
The paper reviews related work in probabilistic wind power forecasting, including extreme learning machines, hybrid intelligent methods, ensemble approaches, and the Lower Upper Bound Estimation (LUBE) method. It also covers nonlinear quantile regression, including local QR, spline-based QR, quantile regression forests, support vector quantile regression (SVQR), and quantile regression neural networks (QRNN). The authors note that previous QRNN approaches either did not address optimization or used the Huber norm for smooth approximations, whereas their approach uses a logistic-based smooth function.
The mathematical foundation is detailed. The pinball loss function is defined as: ρτ(u) = τu if u ≥ 0, and (τ−1)u if u < 0, where 0 < τ < 1. The proposed smooth approximation is given by: Sτ,α(u) = τu + α log(1 + exp(−u/α)), where α > 0 is a smoothing parameter. The authors state: Zheng proves that in the limit as α → 0+ that Sτ,α(u) = ρτ(u).
This smooth function allows the direct application of gradient-based optimization methods, which are preferable for training neural networks.
The SPNN architecture is a feedforward neural network with an input layer of nx features, one or more hidden layers, and an output layer of M quantile estimates. For a single hidden layer, the input-to-hidden weights are W[1] with bias b[1], and hidden-to-output weights are W[2] with bias b[2]. The hidden layer uses a tanh activation function, and the output layer uses the identity activation. The objective function combines the smooth pinball loss summed over all quantiles with L2 regularization on the weights, given by: E = (λ1/(2NM))W[1]2 F + (λ2/(2NM))W[2]2 F + (1/(NM)) Σ Σ [τm(yt − q̂t(τm)) + α log(1 + exp(−(yt − q̂t(τm))/α))].
To prevent quantile crossover, the authors introduce a penalty term: p = c Σ Σ [max(0, ε − (q̂t(τm−1) − q̂t(τm)))]2, where ε is the least amount that two quantiles should differ by, and c is a high penalty parameter. This penalty is added to the cost function, and "If the constraints are not violated no penalty is added to the cost function. If a lower quantile exceeds the value of a higher one, the squared difference of these two quantiles is added to the cost function as a penalty."
The paper details the gradient calculations for backpropagation, showing the chain rule derivations for weights W[2] and W[1] and biases b[2] and b[1]. The Adam optimizer is used for training, as it has been shown to yield superior results compared to other gradient-based optimizers.
Evaluation metrics include the quantile score (QS), which is the pinball loss averaged over all test observations and quantiles; the quantile verification skill score (QVSS) relative to a reference model; the average coverage error (ACE) for interval reliability; and interval sharpness, measured as the mean interval size. The interval score (IS) is also used to evaluate prediction intervals, rewarding narrow intervals and penalizing observations that miss the interval.
Two case studies are conducted using GEFCom2014 wind data from 10 wind farms. The first case study uses Zones 1 and 2, estimating quantiles to form prediction intervals with nominal coverage from 10% to 90% in increments of 10%. The second case study estimates 99 quantiles across all ten wind farms for all 12 test months of 2013, totaling 87,600 test observations. Training uses a sliding window of the previous twelve months to forecast the next month. Input features are raw wind speed data at 10m and 100m for U and V directions, plus four time features based on hour and day of year.
Benchmark methods include persistence (normal distribution from last 24 hours), climatology (based on all past wind power), uniform distribution, multiple quantile regression (QR) with L2 regularization, and support vector quantile regression (SVQR) with a radial basis function kernel. The authors note that unlike the winning GEFCom2014 teams who used dozens of engineered features, their model uses raw data, stating: The goal of our study is not custom feature engineering, which might result in better scores, but to highlight the effectiveness of SPNN in creating its own latent features via its hidden layers.
Results from Case Study 1 show that SPNN2 (two hidden layers) has the lowest deviation from nominal coverage for Zone 1, with deviations ranging from-3% to 0.3%, while SVQR has very poor coverage with deviations as high as-40%. For sharpness, QR has the widest intervals and SVQR the narrowest, but SVQR's narrow intervals fail to capture observations. QVSS analysis shows SPNN1 and SPNN2 provide clear performance increases over QR, with SPNN1 leading for quantiles with τ 0.7 in Zone 1, while SVQR shows negative performance.
In Case Study 2, box plots of QS show SPNN2 has the lowest range (0.036 to 0.047), with SPNN1 close second, while other benchmarks range from 0.075 to 0.011. ACE analysis shows SPNN has the lowest ACE, with SPNN2 having a lower median than SPNN1. For interval score and sharpness, SPNN produces the sharpest intervals across all farms. The authors state: Since both QS and IS also measure skill, we can say that SPNN was able to produce the highest quality estimates from all methods.
The paper also compares mean QS to the top GEFCom2014 teams, noting: The winning team in GEFCom2014 was kPower with a mean QS of 0.038. Our method SPNN2 has a close mean QS of 0.042 which would qualify SPNN to be in the top winning teams.
The conclusion states: "This paper proposes a novel approach we call SPNN for probabilistic wind forecasting using a neural network with a smooth approximation to the pinball ball loss function in estimating multiple quantiles. We also introduce non-crossing constraints in the form of a smooth penalty in the loss function... Our results show superior performance across the prediction horizons, which verify the effectiveness of the model for forecasting while preventing estimated quantiles from overlapping."
Future work directions include applying SPNN to solar and ocean wave power forecasting, electricity pricing and load demand forecasting, and very short-term probabilistic forecasting using only past wind power data.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
Improvement: Replace the standard pinball loss (non-differentiable at zero) with the smooth logistic approximation:
S(τ,α)(u) = τu + α·log(1 + exp(-u/α))
This enables direct use of gradient-based optimizers (e.g., Adam) without subgradient methods or heuristic smoothing schedules.
What the improved system can do:
-
Train quantile regression neural networks with standard backpropagation, achieving faster convergence and more stable training compared to methods using the Huber norm or subgradient approaches.
-
Handle large-scale, high-dimensional data (e.g., wind farm data with multiple weather features) without needing specialized solvers.
Improvement: Add a penalty term to the loss function:
p = c · Σ Σ [max(0, ε - (q̂(τ m-1) - q̂(τ m)))]2
This penalizes violations of monotonicity (lower quantiles exceeding higher ones) without hard constraints, making it compatible with gradient-based optimization.
Improvement: Use a feedforward neural network with M output nodes (one per quantile) and a shared hidden layer, trained jointly with the smooth pinball loss and non-crossing penalty. This contrasts with training separate models per quantile.
Improvement: Support both single-hidden-layer (SPNN1) and two-hidden-layer (SPNN2) architectures with specific hyperparameters: 2000 training iterations, batch size 200, hidden nodes 40 (SPNN1) or 20+40 (SPNN2), smoothing rate α=0.01, L2 regularization λ=0.01, penalty coefficient c=1000, and margin ε=0.
-
Probabilistic Wind Power Forecasting: Given hourly NWP wind speed components (U, V at 10m and 100m) and time features, the system outputs 99 non-crossing quantiles for the next 720 hours, with a quantile score of 0.04 (lower is better) and average coverage error below 0.03.
-
Prediction Interval Generation: Automatically produces 10% to 90% prediction intervals that are reliable (empirical coverage within ±3% of nominal) and sharp (mean interval width 0.2–0.3 normalized power), outperforming persistence, climatology, uniform, QR, and SVQR benchmarks.
-
Scalable Multi-Quantile Regression: Handles 10 wind farms × 12 months × 720 hours = 86,400 test samples with a single model, retrained monthly via sliding window, without quantile crossing—a critical requirement for grid operators using intervals for reserve sizing or bidding.
-
Gradient-Based Training Efficiency: Trains in minutes on a standard CPU (Intel i7, 16GB RAM) using Adam optimizer, unlike support vector quantile regression which requires quadratic programming and is computationally prohibitive for large datasets.
-
Transferability to Other Domains: The same architecture can be applied to solar power, load demand, or electricity price forecasting, where nonparametric probabilistic outputs are needed, and where quantile crossing has historically been a barrier to adoption.
Bottom line: The improved AI system provides fast, reliable, non-crossing probabilistic forecasts for volatile renewable energy sources, enabling better grid integration, reserve planning, and market bidding—all without requiring complex feature engineering or bespoke optimization solvers.
Sources
Related papers
- Runtime Assurance Under Measurement Attack: Necessary and Sufficient Observability Conditions for Learned Control in Radio Access Networks
- Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography
- Uncertainty Quantification in Machine Learning for Biosignal Applications -- A Review
- Continuous Orthogonal Mode Decomposition: Haptic Signal Prediction in Tactile Internet
- Generative Models for Modeling and Synthesizing MIMO Channels in Adverse Weather Conditions
- Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements