Document Type : Research Paper
Authors
Associate Professor, Soil Conservation and Watershed Management Research Institute, Agricultural Research, Education and Extension Organization (AREEO), Tehran, Iran
Abstract
Introduction
Accurate estimation of suspended sediment in high-sediment rivers, especially in semi-arid regions such as the Golestan River watershed, poses significant challenges in water resource management and sediment control in reservoir dams. The increase in sediment concentration not only affects water quality but also leads to considerable economic and environmental damage by reducing the lifespan of hydraulic structures and altering river morphology. In this context, the Atrak River, as one of the most important sediment sources in northeastern Iran, serves as a prominent example of these challenges. Given the limitations of traditional direct measurement methods and labor-intensive physical models, the development of high-accuracy data-driven models emerges as an efficient solution for monitoring and predicting sediment.
Materials and methods
In this study, to estimate suspended sediment in the Atrak River at the Hootan gauging station, a combination of classical and intelligent methods was employed, including: (1) sediment rating curves (using both the midpoint and linear methods), (2) neural networks (MLP and SOM), (3) deep learning models, and (4) ensemble learning methods. The modeling process was carried out in three main stages: first, using the Random Forest algorithm, the key variables affecting sediment (including flow rate, daily precipitation, and their lagged values) were identified. Then, the data were divided into homogeneous groups through clustering, allowing for balanced sampling from each cluster to create homogeneous training (70%), evaluation (15%), and testing (15%) datasets. To enhance the model’s efficiency, various strategies were employed, including optimizing objective functions and managing data skewness. After a comprehensive evaluation of the models using standard criteria, the ensemble learning model XGBoost was selected as the best model, which was ultimately used to reconstruct and complete suspended sediment data for a 40-year period (1982-2021).
Results and discussion
The comparison of the results of various models in this study indicated that the ensemble learning model (XGBoost) was recognized as the selected model among other data-driven models for estimating suspended sediment in the Atrak River, having the lowest error rates (MAE of 10,652 tons per day and RMSE of 36,219 tons per day) and the highest performance index (NSE of 0.87). This model not only outperformed other machine learning methods (with the MLP neural network achieving NSE=0.80 and the deep learning model achieving NSE=0.84), but it also demonstrated significant superiority over classical methods such as the midpoint sediment rating curve (with MAE of 20,909 tons per day and RMSE of 45,632 tons per day). A detailed analysis of the results indicates that this superiority is primarily due to the ability of the XGBoost model to manage highly skewed data and identify complex nonlinear relationships between hydrological variables. In this regard, the model’s bias (PBias) decreased significantly by approximately 15.5% from -18.93 in the sediment rating curve model to -3.52 in the XGBoost model, supporting this point. On the other hand, a comparative analysis with similar studies in other high-sediment rivers shows that the combined approach used in this research (utilizing clustering before model training and optimizing objective functions) has led to improved predictions. However, it is important to note that the effectiveness of these models largely depends on the quality and completeness of the input data, and long-term hydrological changes may necessitate periodic recalibration (training) of the models. These limitations provide a foundation for future research aimed at developing models that are more adaptable to environmental changes.
Conclusion
This research examined the performance of data-driven models in estimating suspended sediment in the Atrak River at the Hootan gauging station, yielding significant results. The ensemble learning model XGBoost was identified as the best option, demonstrating high accuracy and minimal error in predicting changes in suspended sediment. These findings highlight the importance of utilizing advanced machine learning techniques in analyzing hydrological data and emphasize the need for a connection between data science and water resource management. Key points in this study include the high skewness of suspended sediment data and the specific hydrological conditions of the Atrak River, which arise from gully erosion and other natural phenomena in the region. This skewness and the sharp fluctuations in sediment concentration pose major challenges for accurate modeling. However, the combined approach used in this study provided a significant improvement in prediction accuracy compared to traditional and classical methods. This indicates that employing clustering and data optimization can serve as effective tools in future research. Nonetheless, the results of this study should be interpreted with caution, as the quality of data and environmental changes may impact model performance. Therefore, this research not only aids in a better understanding of sedimentation processes but also lays the groundwork for future studies aimed at enhancing and developing prediction models that are more adaptable to changing environmental conditions. In the end, the results of this research can serve as a model for simulating suspended sediment in other hydrometric stations across the country, providing valuable insights for researchers and experts in this field.
Keywords