Early prediction of ready-mixed concrete compressive strength using machine learning-based algorithms: a large-scale industrial dataset study
Construction and Building Materials, cilt.538, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 538
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.conbuildmat.2026.147273
- Dergi Adı: Construction and Building Materials
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC
- Anahtar Kelimeler: Compressive strength prediction, Deep neural network, Ensemble learning, Industrial dataset, LightGBM, Machine learning, MRMR feature selection, Random forest, Ready-mixed concrete, SHAP interpretability, Support vector machine, Voting Regressor
- Çukurova Üniversitesi Adresli: Evet
Özet
The compressive strength of ready-mixed concrete is a critical quality indicator in the construction industry. Conventional 28-day testing procedures delay the identification of substandard batches, causing costly rework and structural risks. This study presents a machine learning (ML)-based framework for the early prediction of ready-mixed concrete compressive strength, using a large-scale dataset of 9007 genuine production records collected from the Akçansa İzmir Aliağa Ready-Mixed Concrete Plant between September 2016 and March 2024. Each record corresponds to an actual concrete batch dispatched from an active industrial plant, and the dataset is substantially larger than those in the existing literature, which typically rely on 100–1000 laboratory-produced specimens. Unlike laboratory datasets generated under controlled conditions, these records capture the full operational variability of real-world RMC production: seasonal raw-material fluctuations, 65 concurrent product families, 30 cement types, and plant-level measurement noise. Seven algorithms (Random Forest, Support Vector Machine, LightGBM, XGBoost, CatBoost, a Deep Neural Network, and a Voting Regressor ensemble), together with a linear-regression baseline, were evaluated under three distinct input scenarios: (i) mixture composition and meteorological features only; (ii) the addition of fresh-state experimental measurements; and (iii) further inclusion of 2-day compressive strength as a predictor. Feature selection was performed using the minimum Redundancy Maximum Relevance (mRMR) algorithm, and hyperparameters were optimised via grid search with 10-fold cross-validation. Model performance was assessed using MAE, MAPE, RMSE, and SMAPE. To prevent data leakage, all scaling and mRMR feature selection were performed strictly within each cross-validation fold, and every metric was computed from out-of-fold predictions. The gradient-boosting models (LightGBM and XGBoost) achieved the best accuracy: in Scenario 1, 28-day MAE = 3.02 MPa (R2 = 0.68); in Scenario 3, MAE = 2.52 MPa (7-day) and 2.50 MPa (28-day, R2 = 0.78), with the Voting Regressor within about 0.1 MPa of the best model. Prediction accuracy improved monotonically as more information was added, while mRMR feature selection provided no benefit and became increasingly counterproductive once 2-day strength was included. Incorporating experimental fresh-state data consistently improves prediction accuracy across all algorithms, and the inclusion of 2-day strength as an additional input variable yields the most accurate predictions across all target days and algorithms evaluated. The proposed framework is provided as a deployable web-based prototype (Python, Streamlit) that produces real-time strength predictions from batching-stage data. To the authors' knowledge, this is one of the largest industrial-scale ML studies on ready-mixed concrete compressive strength prediction reported to date.