In collaboration with Iranian Watershed Management Association

Document Type : Research Paper

Authors

Associate Professor, Soil Conservation and Watershed Management Research Institute, Agricultural Research, Education and Extension Organization (AREEO), Tehran, Iran

Abstract

Introduction
Accurate estimation of suspended sediment in high-sediment rivers, especially in semi-arid regions such as the Golestan River watershed, poses significant challenges in water resource management and sediment control in reservoir dams. The increase in sediment concentration not only affects water quality but also leads to considerable economic and environmental damage by reducing the lifespan of hydraulic structures and altering river morphology. In this context, the Atrak River, as one of the most important sediment sources in northeastern Iran, serves as a prominent example of these challenges. Given the limitations of traditional direct measurement methods and labor-intensive physical models, the development of high-accuracy data-driven models emerges as an efficient solution for monitoring and predicting sediment.
 
Materials and methods
In this study, to estimate suspended sediment in the Atrak River at the Hootan gauging station, a combination of classical and intelligent methods was employed, including: (1) sediment rating curves (using both the midpoint and linear methods), (2) neural networks (MLP and SOM), (3) deep learning models, and (4) ensemble learning methods. The modeling process was carried out in three main stages: first, using the Random Forest algorithm, the key variables affecting sediment (including flow rate, daily precipitation, and their lagged values) were identified. Then, the data were divided into homogeneous groups through clustering, allowing for balanced sampling from each cluster to create homogeneous training (70%), evaluation (15%), and testing (15%) datasets. To enhance the model’s efficiency, various strategies were employed, including optimizing objective functions and managing data skewness. After a comprehensive evaluation of the models using standard criteria, the ensemble learning model XGBoost was selected as the best model, which was ultimately used to reconstruct and complete suspended sediment data for a 40-year period (1982-2021).
 
Results and discussion
The comparison of the results of various models in this study indicated that the ensemble learning model (XGBoost) was recognized as the selected model among other data-driven models for estimating suspended sediment in the Atrak River, having the lowest error rates (MAE of 10,652 tons per day and RMSE of 36,219 tons per day) and the highest performance index (NSE of 0.87). This model not only outperformed other machine learning methods (with the MLP neural network achieving NSE=0.80 and the deep learning model achieving NSE=0.84), but it also demonstrated significant superiority over classical methods such as the midpoint sediment rating curve (with MAE of 20,909 tons per day and RMSE of 45,632 tons per day). A detailed analysis of the results indicates that this superiority is primarily due to the ability of the XGBoost model to manage highly skewed data and identify complex nonlinear relationships between hydrological variables. In this regard, the model’s bias (PBias) decreased significantly by approximately 15.5% from -18.93 in the sediment rating curve model to -3.52 in the XGBoost model, supporting this point. On the other hand, a comparative analysis with similar studies in other high-sediment rivers shows that the combined approach used in this research (utilizing clustering before model training and optimizing objective functions) has led to improved predictions. However, it is important to note that the effectiveness of these models largely depends on the quality and completeness of the input data, and long-term hydrological changes may necessitate periodic recalibration (training) of the models. These limitations provide a foundation for future research aimed at developing models that are more adaptable to environmental changes.
 
Conclusion
This research examined the performance of data-driven models in estimating suspended sediment in the Atrak River at the Hootan gauging station, yielding significant results. The ensemble learning model XGBoost was identified as the best option, demonstrating high accuracy and minimal error in predicting changes in suspended sediment. These findings highlight the importance of utilizing advanced machine learning techniques in analyzing hydrological data and emphasize the need for a connection between data science and water resource management. Key points in this study include the high skewness of suspended sediment data and the specific hydrological conditions of the Atrak River, which arise from gully erosion and other natural phenomena in the region. This skewness and the sharp fluctuations in sediment concentration pose major challenges for accurate modeling. However, the combined approach used in this study provided a significant improvement in prediction accuracy compared to traditional and classical methods. This indicates that employing clustering and data optimization can serve as effective tools in future research. Nonetheless, the results of this study should be interpreted with caution, as the quality of data and environmental changes may impact model performance. Therefore, this research not only aids in a better understanding of sedimentation processes but also lays the groundwork for future studies aimed at enhancing and developing prediction models that are more adaptable to changing environmental conditions. In the end, the results of this research can serve as a model for simulating suspended sediment in other hydrometric stations across the country, providing valuable insights for researchers and experts in this field.
 

Keywords

Adnan, R.M., Petroselli, A., Heddam, S., Santos, C.A.G., Kisi, O., 2022. Suspended sediment modeling using a heuristic regression method hybridized with kmeans clustering. Sustain. 13(9), 4648. https://doi.org/10.3390/su13094648
Bari, S.H., Yokoo, Y., Leong, C., 2024. A brief review of recent global trends in suspended sediment estimation studies. Hy`drolog. Res. Letters 18(2), 51–57. https://doi.org/10.3178/hrl.18.51
Bezak, N., Lebar, K., Bai, Y., Rusjan, S., 2025. Using machine learning to predict suspended sediment transport under climate change. Water Resour. Manage. 39, 3311–3326. https://doi.org/10.1007/s11269-025-03809-0
Chachan, L.J., Bahnam, B.S., 2023. Long short-term memory for predicting monthly suspended sediment load. Tianjin Daxue Xuebao (Ziran Kexue yu Gongcheng Jishu Ban). J. Tianjin Uni. Sci. Technol. 56(05), 36-49. https://doi.org/10.17605/OSF.IO/R9MA2
Chen, T., Guestrin, C., 2016, August. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd. International Conference on Knowledge Discovery and Data Mining, 785-794.
Doroudi, S., Sharafati, A., Mohajeri, S.H., 2021. Estimation of daily suspended sediment load using a novel hybrid support vector regression model incorporated with observer‐teacher‐learner‐based optimization method. Complex. 2021(1), 5540284 (in Persian).
Ezzaouini, M.A., Mahé, G., Kacimi, I., El Bilali, A., Zerouali, A., Nafii, A., 2022. Predicting daily suspended sediment load using machine learning and NARX hydro-climatic inputs in semi-arid environment. Water 14(6), 862. https://doi.org/10.3390/w14060862
Etka, 2024. Data and reports collection and analysis of existing information: site selection and design of a sedimentation pond and water storage reservoir on the Atrak River and pumping station. Etka Organization, Tehran, Iran, 94 pages (in Persian).
Ferguson, R.I., 1987. Accuracy and precision of methods for estimating river loads. Earth Surf. Process. Landform. 12(1), 95-104.
Ghanbari-Adivi, E., 2025. A new machine learning model for predicting suspended sediment load. J. Hydraulic. 19(4), 31-45.
Goodfellow, I., Bengio, Y., Courville, A., 2016. Deep learning. MIT Press, 708 pages.
Hornik, K., Stinchcombe, M., White, H., 1989. Multilayer feedforward networks are universal approximators. Neural Net. 2(5), 359-366.
Hosseiny, H., Masteller, C.C., Dale, J.E., Phillips, C.B., 2023. Development of a machine learning model for river bed load. Earth Surf. Dynamic. 11, 681–693. https://doi.org/10.5194/esurf-11-681-2023
Humphreys, K.M., Mays, D.C., Water quality modeling and machine learning applied to suspended sediment and land management data from the South Fork Clearwater River Basin, Idaho County, Idaho, USA. Available at SSRN: https://ssrn.com/abstract=4887564 or http://dx.doi.org/10.2139/ssrn.4887564
Jansson, M.B., 1996. Estimating a sediment rating curve of the Reventazon river at Palomo using logged mean loads within discharge classes. J. Hydrol. 183(3), 227-241.
Jarbais, G., Harshavardhanan, P., 2025. Comparative analysis of machine learning models for daily suspended sediment concentration prediction in environmental monitoring. Sadhana 50(63). https://doi.org/10.1007/s12046-025-02705-1
Jones, K.R., Berney, O., Carr, D.P., Barrett E.C., 1981. Arid zone hydrology for agricultural development. FAO Irrigation and Drainage Paper, 37, 271 pages.
Kaufman, L., Rousseeuw, P.J., 2009. Finding groups in data: An introduction to cluster analysis (Vol. 344). John Wiley and Sons, New York, USA, 342 pages.
Kaveh, K., Kaveh, H., Bui, M.D., Rutschmann, P., 2021. Long short-term memory for predicting daily suspended sediment concentration. Engin. Comput. 37, 2013-2027.
Khosravi, K., Golkarian, A., Saco, P.M., Booij, M.J.,  Melesse, A.M., 2022. Model identification and accuracy for estimation of suspended sediment load. Geocarto Int. 37(27), 18520-18545. doi:10.1080/10106049.2022.2142964
Khosravi, M., Ghoochani, S., Shabanian, H., 2024. Deep learning-based modeling of daily suspended sediment concentration and discharge in esopus. In 2024 International Conference on Machine Learning and Applications (ICMLA), 882-887). IEEE.
Kohonen, T., 2002. The self-organizing map. Proceedings of the IEEE, 78(9), 1464-1480.
Kumar, D., Bhattacharya, R.K., Singh, A., Banerjee, D., 2022. Modelling of suspended sediment concentration using conventional and machine learning approaches, in the Upper Godavari basin, India. ISH J. Hydraulic Engin. 28(1), 81-92. https://doi.org/10.1080/09715010.2020.1723122
Kwon, S., Noh, H., Seo, I.W., Park, Y.S., 2023. Effects of spectral variability due to sediment and bottom characteristics on remote sensing for suspended sediment in shallow rivers. Sci. Total Environ. 878, 163125. https://doi.org/10.1016/j.scitotenv.2023.163125
Lund, J.W., Groten, J.T., Karwan, D.L., Babcock, C., 2022. Using machine learning to improve predictions and provide insight into fluvial sediment transport. Hydrolog. Process. 36(8), e14648. https://doi.org/10.1002/hyp.14648
Mir, A.A., Patel, M., 2024. A comprehensive review on sediment transport, flow dynamics, and hazards in steep channels. J. Water Manage. Model. 32, C517. https://doi.org/10.14796/JWMM.C517
Mohamadi, S., 2019. The suspended sediment load modeling by artificial neural networks, neural-fuzzy and rating curve in Hlilrood watershed. Watershed Engin. Manage. 11(2), 452-466 (in Persian).
Mohammadi-Raigani, Z., Gholami, H., Mohamadi, M., 2025. Evaluating the performance of machine learning models for predicting suspended sediment load, case study: Taleghan watershed, Iran. E.E.R. 15(1), 83-104 (in Persian).
Mousavi, S.M., Elmizadeh, H., Sakiani, M.A., Zoratipour, A., 2024. Estimation of river-suspended sediments using ANNs and ANFIS methods with modeling in MATLAB, case study: Dez River in Khuzestan. Oceanograph. Fisheries Open Access J. 17(2), 555958. https://doi.org/10.19080/OFOAJ.2024.17.555958
Noh, H., Son, G., Kim, D., Park, Y.S., 2023a. A novel efficient method of estimating suspended‐to‐total sediment load fraction in natural rivers. Water Resour. Res. 59(10), e2022WR034401. https://doi.org/10.1029/2022WR034401
Noh, H., Son, G., Kim, D., Park, Y.S., 2023b. A SVR based-pseudo modified einstein procedure incorporating H-ADCP model for real-time total sediment discharge monitoring. KSCE J. Civil Environ. Engin. Res. 43(3), 321–335. https://doi.org/10.12652/Ksce.2023.43.3.0321
Piraei, R., Afzali, S.H., Niazkar, M., 2023. Assessment of XGBoost to estimate total sediment loads in rivers. Water Resour. Manage. 37, 5289–5306. https://doi.org/10.1007/s11269-023-03603-y
Rahgoshay, M., Feiznia, S., Arian, M., Hashemi, S.A.A., 2019. Simulation of daily suspended sediment load using an improved model of support vector machine and genetic algorithms and particle swarm. Arab. J. Geosci. 12(9), 277.
Sahoo, B.B., Sankalp, S., Kisi, O., 2023. A novel smoothing-based deep learning time-series approach for daily suspended sediment load prediction. Water Resour. Manage. 37(11), 4271-4292.
Shadkani, S., Hemmatzadeh, Y., Pak, A., Abolfathi, S., 2025. Prediction of suspended sediment concentration in fluvial flows using novel hybrid deep learning model. Int. J. Sediment Res. 40(4), 573-587. https://doi.org/10.1016/j.ijsrc.2025.02.004
Shaukat, N., Hashmi, A., Abid, M., Aslam, M.N., Hassan, S., Sarwar, M.K., Masood, A., Shahid, M.L.U.R., Zainab, A., Tariq, M.A.U.R., 2022. Sediment load forecasting of gobindsagar reservoir using machine learning techniques. Front. Earth Sci. 10, 1047290. https://doi.org/10.3389/feart.2022.1047290
Sit, M., Demiray, B.Z., Xiang, Z., Ewing, G.J., Sermet, Y., Demir, I., 2020. A comprehensive review of deep learning applications in hydrology and water resources. Water Sci. Technol. 82(12), 2635-2670.
Swami, S., Underwood, K.L., Wshah, S., Davis, W.D., Rizzo, D.M., 2023. Advancing deep learning techniques for turbidity forecasting in rivers and reservoirs. Graduate College Dissertations and Theses. 1787. https://scholarworks.uvm.edu/graddis/1787
Tabatabaei, M., Salehpour Jam, A., Mosaffaie, J., 2023. Suspended sediment simulation using machine learning algorithms and CHIRPS satellite precipitation data with emphasis on data clustering and gamma test, case study: Ramyan Watershed, Golestan Province. Watershed Engin. Manage. 15(3), 328-350.
Taşar, B., Üneş, F., Demirci, M., Güzel, H., Varçin, H., 2024. Suspended sediment estimation using machine learning methods. 2024 "Air and Water – Components of the Environment" Conference Proceedings, Cluj-Napoca, Romania, 105-114. https://doi.org/10.24193/AWC2024_10.
Tayfur, G., 2012. Soft computing in water resources engineering: Artificial neural networks, fuzzy logic and genetic algorithms. WIT Press, Dorset, UK, 288 pages.
Ulke, A., Tayfur, G., Ozku, S., 2009. Predicting suspended sediment loads and missing datafor gediz river, Turkey. J. Hydrol. Engin. 14(9), 954-965.
Üneş, F., Tasar, B., Varçin, H., 2025. Forecasting of suspended sediment in river using artificial intelligence methods. Air and Water – Components of the Environment Conference Proceedings, Cluj-Napoca, Romania, 101-107. https://doi.org/10.24193/AWC2025_09
Zhang, C., Liu, Y., Chen, X., Gao, Y., 2022. Estimation of suspended sediment concentration in the yangtze main stream based on sentinel-2 MSI data. Remote Sens. 14(18), 4446.