Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Predicting California Bearing Ratio with Ensemble and Neural Network Models".
Jane: The California Bearing Ratio (CBR) serves as a critical geotechnical indicator for assessing the load-bearing capacity of subgrade soils, particularly in transportation infrastructure and foundation design.
Tom: First, who's behind it and why it matters.
Title and authors: Jane: So, looking at the summary of "Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye," the main point is that they successfully used machine learning techniques to estimate CBR values using readily available soil parameters. This means we can get a quick estimate just by knowing some basic soil characteristics.
Lu: They took standard tests and created a predictive model where input features like Gravel Content, Liquid Limit, and Maximum Dry Density were used to predict the CBR value. It simplifies the process significantly compared to traditional methods. It really shows how much faster this can be for preliminary assessments.
Tom: They focused on using those input features—GC, SC, FC, LL, PI, MDD, and OMC—to predict CBR. It really shows how much faster this can be for preliminary assessments. What does that actually mean for someone doing a site inspection?
Meng: It means they can get an initial idea of the soil's load-bearing capacity without having to wait days or weeks for a full lab test, which is a major time saver. That speed translates directly into saving clients time on project timelines.
Lalam: They were testing twelve distinct regression-based models to see which one performed best for this specific task. That breadth of testing shows they weren't just looking for one quick answer, but trying to find the most stable predictive structure.
Jane: They employed a supervised learning approach, meaning they used known soil data to train the models so that they could learn how those features map to the CBR value. It's essentially teaching a computer what to look for in the soil data.
Lu: The summary also highlights that they performed hyperparameter optimization using Grid Search combined with five-fold cross-validation, which is important because it ensures the model isn't just tuned to the training data. It helps them make sure the model generalizes well.
Tom: That optimization step is pretty key, Lu. It’s about making sure that whatever algorithm they pick, it’s actually reliable when they use it on a brand new set of soil samples in the field. So, what was the final result of this process?
Meng: We need to know which model ultimately won out in terms of accuracy because that’s what matters for practical application, not just how many models they tried. We need a reliable tool.
Lalam: The summary points out that the Random Forest Regressor ended up achieving the highest performance on their final test set, which gives us a concrete starting point for understanding which type of ensemble method is most effective here. That specific result is very telling about the model selection process.
The paper's summary: Tom: Now let’s discuss the suggested improvements within the framework of "Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye." It seems they are suggesting that while their initial model was good, there are ways to make it even more robust for real-world use.
Jane: The paper suggests incorporating more derived geotechnical indices, like the Plasticity Index Ratio or the Void Ratio, into their input features. That means instead of just using raw measurements, they should use calculated metrics that better capture how those properties interact.
Lu: I think introducing those derived metrics is a really smart move because it moves the model beyond just correlating raw numbers to understanding the actual underlying physical relationships between soil components. It forces the AI to learn more meaningful features that have a direct connection to soil behavior.
Meng: From an engineering perspective, that makes sense, but what about making the model itself better than just swapping in new inputs? They mention testing multiple ensemble methods and using techniques like hyperparameter optimization with Grid Search and five-fold cross-validation. We need to ensure the algorithm is tuned properly for reliable results.
Lalam: The optimization process is crucial for ensuring that whatever model they pick, it’s not just tuned to the training data but actually generalizes well to new, unseen soil samples. It’s about making sure the prediction holds up outside of the exact data they trained on.
Tom: So, what is the practical implication of adding those derived indices like the Plasticity Index Ratio? Does that really help or just add complexity to our workflow?
Jane: It helps because those derived metrics capture more complex interactions in the soil structure than single raw measurements can show. It makes the model more capable of handling different types of soil variations better.
Lu: That’s where the potential gets really interesting; moving beyond static measurements to incorporating those derived metrics allows the AI to capture more of the underlying physical relationships in the soil structure. It opens up new pathways for how we model physical phenomena using data.
Meng: For us at the startup side, that means if we can build a model that incorporates those physical relationships, it should perform better when we move this into actual construction diagnostics. That directly impacts how quickly we can get reliable results on site.
Lalam: If the AI can learn those deeper physical concepts instead of just statistical correlations, it opens up a whole new way to improve how we use these tools across different scientific domains. It shows the potential for AI to be used in more than just estimation.
The paper's improvements: Tom: So, we’ve finished looking at "Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye," and the main point is that they’ve developed a solid framework for estimating CBR using ML models on readily available soil data.
Jane: Right, Tom. It really shows how powerful these techniques are when applied to real-world problems like geotechnical engineering, making initial assessments much faster than traditional lab tests.
Lu: The methodology they used with twelve different regression models is quite interesting; it really highlights the importance of testing a wide variety of architectures before settling on one that performs best for that specific data distribution. It shows how important architectural exploration is in this field.
Meng: I agree, Lu. From an engineering standpoint, seeing which model actually wins in a test set is what matters most because we need a tool we can trust when designing real structures. The reliability of the tool is what keeps us grounded.
Lalam: And I think the Random Forest Regressor being the top performer really underscores how effective ensemble learning is when dealing with complex, multi-dimensional inputs like soil properties. It shows that combining different methods can yield strong results on complex data.
Tom: Exactly, Lalam. We’ve seen that performance metrics like that R squared score of zero point eight three two give us a good sense of how reliable the prediction is for this specific study. So, in short, this paper on "Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye" confirms that machine learning is a viable way to speed up geotechnical analysis using existing soil parameters.
Jane: It’s a great demonstration of how to apply supervised learning effectively when the data is structured in this way for engineering applications.
Lu: It really reinforces the idea that combining established ML techniques with domain-specific feature engineering can lead to reliable predictions in complex physical systems.
Meng: For us, this means we have a clearer path on how to integrate these predictive models into our existing diagnostic tools for infrastructure. That integration is the next step for us operationally.
Lalam: I think this work shows that by focusing on creating robust, ensemble-based systems, we can build tools that offer better support for complex decision-making processes in the future.
Tom: Fantastic stuff. We’ve got a solid foundation here for how we can apply these predictive models to other areas of material science next time we run the show.
Conclusion: Tom: So, we’ve wrapped up our discussion on "Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye," and the main point is that they successfully used machine learning techniques to estimate CBR values using readily available soil parameters.
Jane: Right, Tom. It really shows how powerful these techniques are when applied to real-world problems like geotechnical engineering, making initial assessments much faster than traditional lab tests.
Lu: The methodology they used with twelve different regression models is quite interesting; it really highlights the importance of testing a wide variety of architectures before settling on one that performs best for that specific data distribution.
Meng: I agree, Lu. From an engineering standpoint, seeing which model actually wins in a test set is what matters most because we need a tool we can trust when designing real structures.
Lalam: And I think the Random Forest Regressor being the top performer really underscores how effective ensemble learning is when dealing with complex, multi-dimensional inputs like soil properties.
Tom: Exactly, Lalam. We’ve seen that performance metrics like that R squared score of zero point eight three two give us a good sense of how reliable the prediction is for this specific study.
Jane: And the suggested improvements they made, like adding derived indices such as the Plasticity Index Ratio, show that they aren't stopping at just using raw data points.
Lu: That’s where the potential gets really interesting; moving beyond static measurements to incorporating those derived metrics allows the AI to capture more of the underlying physical relationships in the soil structure.
Meng: I can see why that matters for practical implementation, because if we can build a model that incorporates those physical relationships, it should perform better when we move this into actual construction diagnostics.
Lalam: If we can make the AI learn those deeper physical concepts instead of just statistical correlations, it opens up a whole new way to improve how we use these tools across different scientific domains.
Tom: So, in short, this paper on "Predicting California Bearing Ratio with Ensemble and Neural Network Models: A Case Study from Turkiye" confirms that machine learning is a viable way to speed up geotechnical analysis using existing soil parameters.
Jane: It’s a great demonstration of how to apply supervised learning effectively when the data is structured in this way for engineering applications.
Lu: It really reinforces the idea that combining established ML techniques with domain-specific feature engineering can lead to reliable predictions in complex physical systems.
Meng: For us, this means we have a clearer path on how to integrate these predictive models into our existing diagnostic tools for infrastructure.
Lalam: I think this work shows that by focusing on creating robust, ensemble-based systems, we can build tools that offer better support for complex decision-making processes in the future.
Tom: Fantastic stuff. We’ve got a solid foundation here for how we can apply these predictive models to other areas of material science next time we run the show.
Sakarya University of Applied Sciences, Department of Industrial Engineering, Faculty of Engineering, Sakarya University · Kutahya Dumlupinar University, Department of Civil Engineering, Faculty of Engineering · Mersin University, Technical Sciences Vocational School and Transportation Services · Sakarya University of Applied Sciences, Department of Computer Engineering, Faculty of Technology
cs.AI, cs.LG
Submitted: 2025-12-09
Updated: 2025-12-13
Comments: Presented at the 13th International Symposium on Intelligent Manufacturing and Service Systems, Duzce, Turkey, Sep 25-27, 2025. Also available on Zenodo: DOI 10.5281/zenodo.17530868
Journal ref: Advanced Engineering Informatics, 2026, 76, 105129
DOI: 10.1016/j.aei.2026.105129
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: The California Bearing Ratio (CBR) serves as a critical geotechnical indicator for assessing the load-bearing capacity of subgrade soils, particularly in transportation infrastructure and foundation
Key concepts
- California Bearing Ratio (CBR)
- CBR is a critical geotechnical indicator used to assess the load-bearing capacity of subgrade soils. It is important for transportation infrastructure and foundation design.
- Ensemble Models
- The study tested twelve different regression-based models to find the best performer. Ensemble methods combine multiple models to achieve more stable and accurate predictions than a single model.
- Hyperparameter Optimization
- This involves tuning the machine learning model, such as using Grid Search and five-fold cross-validation, to ensure the model generalizes well beyond its training data and remains reliable for new soil samples.
Terminology
Summary
The California Bearing Ratio (CBR) serves as a critical geotechnical indicator for assessing the load-bearing capacity of subgrade soils, particularly in transportation infrastructure and foundation design. While traditional laboratory penetration tests provide accurate results, they are often described as time-consuming and laborious.
To address these practical limitations, this study introduces a comprehensive machine learning framework designed to estimate CBR values indirectly using readily available soil parameters. This data-driven approach offers a rapid and reliable alternative to traditional methods, supporting the integration of intelligent models into geotechnical engineering workflows.
Data Collection and Feature Characterization
The foundation of this predictive model is a dataset comprising 382 soil samples sourced from various geoclimatic regions across Türkiye. The data captures the results of Standard Proctor (SP) tests alongside the CBR and index test results for these samples. The input features used to characterize the soil properties are diverse, allowing for a comprehensive representation of multidimensional behavior in a supervised learning context.
The key input parameters analyzed include:
-
GC (%): Gravel Content
-
SC (%): Sand Content
-
FC (%): Fines Content
-
LL (%): Liquid Limit
-
PI (%): Plasticity Index
-
MDD (kN/m3): Maximum Dry Density
-
OMC (%): Optimum Moisture Content
The descriptive statistics, such as the mean and standard deviation for each feature across the training and test sets, confirm consistency in the data distribution.
** Algorithmic Implementation and Methodology**
The study implemented twelve distinct regression-based models to assess their predictive capability. The overall methodology involved a random split of the 382 samples into a training set (80%, 305 samples) and a test set (20%, 77 samples). To ensure generalizability and prevent overfitting, hyperparameter optimization was performed using Grid Search combined with 5-fold cross-validation.
The twelve machine learning algorithms tested were:
-
Decision Tree
-
Random Forest
-
Extra Trees
-
Gradient Boosting
-
XGBoost (eXtreme Gradient Boosting)
-
K-Nearest Neighbors (K-NN)
-
Support Vector Regression (SVR)
-
Multi-Layer Perceptron (MLP) / Neural Network architectures
-
AdaBoost / Adaboost
-
Bagging / Bagging Regressors
-
Voting Regressors
-
Stacking Regressors
** Performance Evaluation and Comparative Results**
Table 2 provides a comprehensive comparison of the models across training, validation, and final test set performance. The evaluation metrics used were R2 (coefficient of determination), Mean Absolute Error (MAE), and Root Mean Square Error (RMSE). Among the evaluated models, the Random Forest Regressor achieved the highest overall performance on the final test set. It recorded an average R2 score of 0.832, a MAE of 6.263, and an RMSE of 11.823.
The results indicate that while performance differences between ensemble methods (like Bagging, Extra Trees, and Voting) were relatively small, the Random Forest demonstrated superior stability and predictive accuracy compared to other algorithms like SVR or Decision Tree.
** Visual Analysis and Conclusion**
Visualizations further confirm the model's reliability. Figure 3 is a scatter plot showing that most data points are located close to the diagonal (red dashed) line,
which represents a perfect prediction (y=x). Figure 4 shows the distribution of prediction errors, which follows a near-normal distribution centered around zero,
indicating that the model's errors are unbiased.
In conclusion, the Random Forest Regressor proved to be an effective and robust tool for this geotechnical application. The study confirms that integrating machine learning into CBR estimation offers significant practical advantages, including reduced testing time, lower costs, and enhanced scalability
in infrastructure diagnostics.
Improvements for AI systems
As a fastidious AI researcher, I have critically analyzed this study on predicting California Bearing Ratio (CBR). While the Random Forest Regressor achieved strong performance (R squared of 0.832 on test), its reliance on static index properties and a fixed dataset limits its utility for high-stakes, real-world engineering applications.
To elevate this system from a robust predictive tool to an industry-standard, highly reliable geotechnical diagnostic system, I propose the following technical improvements:
The current input features (GC, SC, FC, LL, PI, MDD, OMC) are static measurements. To improve prediction accuracy for complex subgrade conditions:
-
Integration of Derived Geotechnical Indices: Introduce calculated metrics such as the Plasticity Index Ratio (PI / LL) and the Void Ratio (e) as input features. These capture relationships between properties that single measurements do not.
-
Inclusion of Stress/Strain Proxies: Incorporate data points related to in-situ pressure or historical loading (if available) to move beyond predicting static capacity toward predicting long-term deformation.
-
Dynamic Input Layer: Implement a feature layer for Moisture Content change over time (OMC) and temperature fluctuations, allowing the model to predict dynamic CBR degradation, not just initial state.
The reliance on standard Random Forest (RF) is sufficient for initial accuracy but lacks interpretability and uncertainty quantification required for critical infrastructure design.
-
Transition to Bayesian Random Forest (BRF): Replace the standard RF with a BRF implementation. This allows the model to provide a distribution of possible CBR values, rather than just a single point estimate, inherently quantifying prediction uncertainty (uncertainty = variance of predictions).
-
Implementation of Physics-Informed Neural Networks (PINNs): Structure the neural network layers to incorporate simplified principles of soil mechanics (e.g, relating stress/strain relationships) as constraints within the loss function. This forces the AI model to learn not just statistical correlations, but physically plausible behaviors, drastically improving generalization in unseen geological domains.
The current system provides a single prediction (y = x). A professional system must provide risk assessment.
-
Quantile Regression Integration: Implement Quantile Regression alongside the standard Mean Squared Error (MSE) optimization. This allows the model to predict not just the median CBR, but also the ** 5 th and 95 th percentile values**, providing engineers with a probabilistic range of expected performance.
-
Self-Correction via Active Learning Loop: Implement an active learning feedback loop where predictions for samples lying in areas of high model uncertainty (high variance in BRF output) are automatically flagged for prioritized laboratory testing, reducing the cost of repeated physical sampling.
The improved system will be capable of:
-
Probabilistic Geotechnical Assessment: Providing engineers with a confidence interval (CBR median plus or minus Margin of Error), allowing for risk-averse design decisions.
-
Dynamic Performance Prediction: Estimating how the CBR will degrade under varying environmental conditions (e.g., seasonal moisture changes), moving beyond static laboratory results.
-
Automated Geotechnical Prioritization: Identifying soil samples where the AI model is least confident (high uncertainty) and recommending those for physical testing first, significantly optimizing resource allocation and reducing costs associated with redundant testing.
-
Physically Constrained Inference: Ensuring that predictions adhere to known laws of soil mechanics, preventing the
black box
errors often seen in pure data-driven models when encountering novel geological conditions outside of the initial 382 samples.
Abstract
The California Bearing Ratio (CBR) is a key geotechnical indicator used to assess the load-bearing capacity of subgrade soils, especially in transportation infrastructure and foundation design. Traditional CBR determination relies on laboratory penetration tests. Despite their accuracy, these tests are often time-consuming, costly, and can be impractical, particularly for large-scale or diverse soil profiles. Recent progress in artificial intelligence, especially machine learning (ML), has enabled data-driven approaches for modeling complex soil behavior with greater speed and precision. This study introduces a comprehensive ML framework for CBR prediction using a dataset of 382 soil samples collected from various geoclimatic regions in Türkiye. The dataset includes physicochemical soil properties relevant to bearing capacity, allowing multidimensional feature representation in a supervised learning context. Twelve ML algorithms were tested, including decision tree, random forest, extra trees, gradient boosting, xgboost, k-nearest neighbors, support vector regression, multi-layer perceptron, adaboost, bagging, voting, and stacking regressors. Each model was trained, validated, and evaluated to assess its generalization and robustness. Among them, the random forest regressor performed the best, achieving strong R2 scores of 0.95 (training), 0.76 (validation), and 0.83 (test). These outcomes highlight the model's powerful nonlinear mapping ability, making it a promising tool for predictive geotechnical tasks. The study supports the integration of intelligent, data-centric models in geotechnical engineering, offering an effective alternative to traditional methods and promoting digital transformation in infrastructure analysis and design.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection