نوع مقاله : مقاله پژوهشی
نویسندگان
دانشیار پژوهشی پژوهشکده حفاظت خاک و آبخیزداری، سازمان تحقیقات، آموزش و ترویج کشاورزی، تهران، ایران
چکیده
مقدمه
برآورد دقیق رسوب معلق در رودخانههای پررسوب، بهویژه در مناطق نیمهخشک مانند حوزه گرگانرود، از چالشهای اساسی در مدیریت منابع آب و کنترل رسوبات مخازن سدها محسوب میشود. افزایش غلظت رسوبات نهتنها بر کیفیت آب تأثیر میگذارد، بلکه با کاهش عمر مفید سازههای هیدرولیکی و تغییر مورفولوژی رودخانه، خسارات اقتصادی و زیستمحیطی قابلتوجهی بهدنبال دارد. در این میان، رودخانه اترک بهعنوان یکی از مهمترین منابع رسوبی در شمال شرق ایران، نمونهای بارز از این چالشهاست. با توجه به محدودیتهای روشهای سنتی اندازهگیری مستقیم و مدلهای فیزیکی پرزحمت، توسعه مدلهای دادهمبنا با دقت بالا بهعنوان راهکاری کارآمد برای پایش و پیشبینی رسوب مطرح میشود.
مواد و روشها
در این پژوهش، برای برآورد رسوب معلق رودخانه اترک در ایستگاه هوتن، از ترکیبی از روشهای کلاسیک و هوشمند شامل: (1) منحنی سنجه رسوب (با دو روش حد وسط دستهها و یک خطی)، (2) شبکههای عصبی (MLP و SOM)، (3) مدلهای یادگیری عمیق و (4) روشهای یادگیری جمعی استفاده شد. فرایند مدلسازی در سه مرحله اصلی انجام پذیرفت. ابتدا با بهرهگیری از الگوریتم جنگل تصادفی، متغیرهای کلیدی مؤثر بر رسوب (شامل دبی جریان، بارش روزانه و مقادیر تأخیری آنها) شناسایی شدند. سپس دادهها با روش خوشهبندی به گروههای همگن تقسیم شدند که این امر امکان نمونهبرداری متوازن از هر خوشه برای ایجاد مجموعههای همگن آموزش (70 درصد)، ارزیابی (15 درصد) و آزمون (15 درصد) را فراهم نمود. بهمنظور افزایش کارائی مدلها، از راهکارهای مختلفی شامل بهینهسازی توابع هدف و مدیریت چولگی دادهها استفاده شد. پس از ارزیابی جامع مدلها با معیارهای استاندارد، مدل یادگیری جمعی XGBoost بهعنوان مدل برتر انتخاب شد که در نهایت برای بازسازی و تکمیل دادههای رسوب معلق در دوره 40 ساله (1401-1360) بهکار گرفته شد.
نتایج و بحث
مقایسه نتایج مدلهای مختلف در این پژوهش نشان داد که مدل یادگیری جمعی (XGBoost) با داشتن کمترین میزان خطا MAE برابر با 10652 تن در روز و RMSE برابر 36219 تن در روز) و بیشرین شاخص کارایی (NSE برابر با 0.87) بهعنوان مدل منتخب از بین سایر مدلهای داده مبنا برای برآورد رسوب معلق در رودخانه اترک شناخته شد. این مدل نه تنها در مقایسه با سایر روشهای یادگیری ماشین مانند شبکه عصبی MLP با NSE=0.80)و مدل یادگیری عمیق با NSE=0.84 عملکرد بهتری داشت، بلکه نسبت به روشهای کلاسیک مانند منحنی سنجه حد وسط دستهها با MAE برابر با 20909 تن در روز و RMSE برابر 45632 تن در روز، نیز برتری قابل توجهی نشان داد. تحلیل دقیق نتایج حاکی از آن است که این برتری عمدتاً ناشی از توانایی مدل XGBoost در مدیریت دادههای با چولگی بالا و شناسایی رابطههای غیرخطی پیچیده بین متغیرهای هیدرولوژیکی است. در این رابطه مقدار بایاس (PBias) مدل با کاهش تقریباً 15.5 درصدی از 18.93- در مدل منحنی سنجه رسوب به 3.52- در مدل XGBoost مؤید این نکته است. از سوی دیگر، مقایسه تطبیقی با مطالعات مشابه نشان میدهد که رویکرد ترکیبی بهکار گرفته شده در این پژوهش (استفاده از خوشهبندی پیش از آموزش مدل و بهینهسازی توابع هدف) منجر به بهبود پیشبینی شده است. با این حال، باید توجه داشت که کارایی این مدلها تا حد زیادی به کیفیت و کامل بودن دادههای ورودی وابسته است و تغییرات هیدرولوژیکی بلندمدت ممکن است نیاز به بازواسنجی (آموزش) دورهای مدلها را ایجاد کند. این محدودیتها زمینهای برای تحقیقات آتی در جهت توسعه مدلهای سازگارتر با تغییرات محیطی فراهم میسازد.
نتیجهگیری
این پژوهش به بررسی کارایی مدلهای دادهمبنا در برآورد رسوب معلق در رودخانه اترک، ایستگاه آبسنجی هوتن پرداخته و به نتایج قابل توجهی دست یافته است. مدل یادگیری جمعی XGBoost بهعنوان بهترین گزینه شناسایی شد که با دقت بالا و کمترین خطا، توانایی پیشبینی تغییرات رسوب معلق را به نمایش گذاشت. این یافتهها نشاندهنده اهمیت استفاده از تکنیکهای پیشرفته یادگیری ماشین در تحلیل دادههای هیدرولوژیکی است و بر لزوم پیوند بین علم داده و مدیریت منابع آب تأکید میکند. از نکات کلیدی در این تحقیق، چولگی بالای دادههای رسوب معلق و شرایط خاص هیدرولوژیکی رودخانه اترک است که ناشی از فرسایشهای خندقی و پدیدههای طبیعی دیگر در این منطقه است. این چولگی و تغییرات شدید در غلظت رسوب، چالشهای عمدهای را برای مدلسازی دقیق به وجود میآورد. با این حال، رویکرد ترکیبی بهکاررفته در این مطالعه، بهبود قابل توجهی در دقت پیشبینی نسبت به روشهای سنتی و کلاسیک فراهم آورد. این امر حاکی از آن است که استفاده از خوشهبندی و بهینهسازی دادهها میتواند بهعنوان ابزاری مؤثر در تحقیقات آینده مورد توجه قرار گیرد. با این حال، نتایج این تحقیق باید با احتیاط تفسیر شوند. چرا که کیفیت دادهها و تغییرات محیطی ممکن است بر عملکرد مدلها تأثیر بگذارد. بنابراین، این مطالعه نهتنها به درک بهتر فرایندهای رسوبگذاری کمک میکند، بلکه زمینهساز تحقیقات آتی در جهت بهبود و توسعه مدلهای پیشبینی سازگارتر با شرایط متغیر محیطی خواهد بود. در نهایت نتایج این تحقیق میتواند بهعنوان یک کار الگویی در شبیهسازی رسوب معلق در دیگر ایستگاههای هیدرومتری کشور مورد استفاده محققین و کارشناسان این حوزه قرار گیرد.
کلیدواژهها
عنوان مقاله [English]
Assessment of the performance of data-driven models in estimating suspended sediment in high-sediment rivers: a case study of the Hootan gauging station, Atrak River, Maraveh Tappeh study unit
نویسندگان [English]
- Mahmoudreza Tabatabaei
- Mohammadreza Gharib Reza
Associate Professor, Soil Conservation and Watershed Management Research Institute, Agricultural Research, Education and Extension Organization (AREEO), Tehran, Iran
چکیده [English]
Introduction
Accurate estimation of suspended sediment in high-sediment rivers, especially in semi-arid regions such as the Golestan River watershed, poses significant challenges in water resource management and sediment control in reservoir dams. The increase in sediment concentration not only affects water quality but also leads to considerable economic and environmental damage by reducing the lifespan of hydraulic structures and altering river morphology. In this context, the Atrak River, as one of the most important sediment sources in northeastern Iran, serves as a prominent example of these challenges. Given the limitations of traditional direct measurement methods and labor-intensive physical models, the development of high-accuracy data-driven models emerges as an efficient solution for monitoring and predicting sediment.
Materials and methods
In this study, to estimate suspended sediment in the Atrak River at the Hootan gauging station, a combination of classical and intelligent methods was employed, including: (1) sediment rating curves (using both the midpoint and linear methods), (2) neural networks (MLP and SOM), (3) deep learning models, and (4) ensemble learning methods. The modeling process was carried out in three main stages: first, using the Random Forest algorithm, the key variables affecting sediment (including flow rate, daily precipitation, and their lagged values) were identified. Then, the data were divided into homogeneous groups through clustering, allowing for balanced sampling from each cluster to create homogeneous training (70%), evaluation (15%), and testing (15%) datasets. To enhance the model’s efficiency, various strategies were employed, including optimizing objective functions and managing data skewness. After a comprehensive evaluation of the models using standard criteria, the ensemble learning model XGBoost was selected as the best model, which was ultimately used to reconstruct and complete suspended sediment data for a 40-year period (1982-2021).
Results and discussion
The comparison of the results of various models in this study indicated that the ensemble learning model (XGBoost) was recognized as the selected model among other data-driven models for estimating suspended sediment in the Atrak River, having the lowest error rates (MAE of 10,652 tons per day and RMSE of 36,219 tons per day) and the highest performance index (NSE of 0.87). This model not only outperformed other machine learning methods (with the MLP neural network achieving NSE=0.80 and the deep learning model achieving NSE=0.84), but it also demonstrated significant superiority over classical methods such as the midpoint sediment rating curve (with MAE of 20,909 tons per day and RMSE of 45,632 tons per day). A detailed analysis of the results indicates that this superiority is primarily due to the ability of the XGBoost model to manage highly skewed data and identify complex nonlinear relationships between hydrological variables. In this regard, the model’s bias (PBias) decreased significantly by approximately 15.5% from -18.93 in the sediment rating curve model to -3.52 in the XGBoost model, supporting this point. On the other hand, a comparative analysis with similar studies in other high-sediment rivers shows that the combined approach used in this research (utilizing clustering before model training and optimizing objective functions) has led to improved predictions. However, it is important to note that the effectiveness of these models largely depends on the quality and completeness of the input data, and long-term hydrological changes may necessitate periodic recalibration (training) of the models. These limitations provide a foundation for future research aimed at developing models that are more adaptable to environmental changes.
Conclusion
This research examined the performance of data-driven models in estimating suspended sediment in the Atrak River at the Hootan gauging station, yielding significant results. The ensemble learning model XGBoost was identified as the best option, demonstrating high accuracy and minimal error in predicting changes in suspended sediment. These findings highlight the importance of utilizing advanced machine learning techniques in analyzing hydrological data and emphasize the need for a connection between data science and water resource management. Key points in this study include the high skewness of suspended sediment data and the specific hydrological conditions of the Atrak River, which arise from gully erosion and other natural phenomena in the region. This skewness and the sharp fluctuations in sediment concentration pose major challenges for accurate modeling. However, the combined approach used in this study provided a significant improvement in prediction accuracy compared to traditional and classical methods. This indicates that employing clustering and data optimization can serve as effective tools in future research. Nonetheless, the results of this study should be interpreted with caution, as the quality of data and environmental changes may impact model performance. Therefore, this research not only aids in a better understanding of sedimentation processes but also lays the groundwork for future studies aimed at enhancing and developing prediction models that are more adaptable to changing environmental conditions. In the end, the results of this research can serve as a model for simulating suspended sediment in other hydrometric stations across the country, providing valuable insights for researchers and experts in this field.
کلیدواژهها [English]
- Data-driven models
- Ensemble learning
- Gradient boosting
- Machine learning methods
- Prediction
- XGBoost