نشریه علمی پژوهشی مهندسی آبیاری و آب ایران

نشریه علمی پژوهشی مهندسی آبیاری و آب ایران

تخمین داده های مفقود جریان رودخانه با استفاده از الگوریتم‌های یادگیری جمعی و ماشینی (مطالعه موردی: رودخانه کرخه)

نوع مقاله : مقاله پژوهشی

نویسندگان
1 دانشجوی دکترای تخصصی، گروه مهندسی آب و سازه های هیدرولیکی، دانشکده مهندسی عمران، دانشگاه سمنان، ایران
2 دانشیار، گروه مهندسی آب و سازه های هیدرولیکی، دانشکده مهندسی عمران، دانشگاه سمنان، ایران
3 استاد، گروه مهندسی آب و سازه های هیدرولیکی، دانشکده مهندسی عمران، دانشگاه سمنان، ایران
10.22125/iwe.2025.491838.1842
چکیده
در این پژوهش، از 9 الگوریتم یادگیری جمعی و ماشینی شامل الگوریتم‌های Xgboost، Catboost، Extra Trees، Random Forest، M5، MLP، K-NN، Decision Tree وSVR برای تخمین داده‌های مفقود جریان روزانه رودخانه کرخه استفاده شد. جهت برآورد داده‌های مفقود ایستگاه عبدالخان و پای پل، داده‌های جریان روزانه ایستگاه هیدرومتری حمیدیه به عنوان ایستگاه همسایه در دوره آماری 40 ساله مورد بررسی قرار گرفت. بهینه‌سازی فراپارامترهای الگوریتم‌های مذکور، به روش Optuna انجام شد. مقایسه عملکرد مدل‌ها نشان داد که الگوریتم Xgboost با یادگیری روابط غیرخطی پیچیده،دقت بیشتری در تخمین داده‌های مفقود دارد. الگوریتم مذکور، در ایستگاه‌های عبدالخان و پای پل، با داشتن بیشترین مقدار ضریب تعیین (R2) به ترتیب برابر با 95/0 و 78/0 و کمترین مقدار میانگین خطای مطلق (MAE) بترتیب برابر با 76/18 و 45/36 بهترین عملکرد را دارد. همچنین، کمترین مقدار ریشه میانگین مربع خطاها (RMSE) برابر با 75/43 و 87/108 به‌دست آمد.علاوه براین، الگوریتم Xgboost کمترین مقدار مجذور میانگین مربعات خطای نسبی (RRMSE) برابر با 20/0 و 46/0 ثبت کرد.بنابراین، الگوریتم Xgboost بیشترین کارایی را در تخمین داده‌های مفقود نسبت به بقیه مدل‌ها در هر دو ایستگاه دارد. همچنین، می‌تواند بر چالش‌های مکانی و داده‌های محدود غلبه کند. نتایج نمودار تیلور نیز حاکی از برتری مدل Xgboost در هر دو ایستگاه مذکور است. مدل Catboost نیز در ایستگاه‌های عبدالخان و پای پل به ترتیب 11% و 5% دقت کمتر از مدل Xgboost داشت و دومین جایگاه را میان مدل‌های بررسی‌شده کسب کرد .نتایج این پژوهش می‌تواند جهت تخمین جریان رودخانه در سایر ایستگاه‌های فاقد آمار مفید واقع شود.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Estimation of Missing Streamflow Data in River Using Ensemble and Machine Learning Algorithms (Case Study: Karkheh River)

نویسندگان English

Mahsa Boustani 1
Saeed Farzin 2
Sayed-Farhad Mousavi 3
1 PhD Candidate, Department of Water Engineering and Hydraulic Structures, Faculty of Civil Engineering, Semnan University, Semnan, Iran,
2 Associated Professor, Department of Water Engineering and Hydraulic Structures, Faculty of Civil Engineering, Semnan University, Semnan, Iran
3 Professor, Department of Water Engineering and Hydraulic Structures, Faculty of Civil Engineering, Semnan University, Semnan
چکیده English

In this study, nine ensemble and machine learning algorithms, including Xgboost, Catboost, Extra Trees, Random Forest, M5, MLP, K-NN, Decision Tree, and SVR, were employed to estimate missing daily streamflow data for the Karkheh River in southwestern Iran. To estimate the missing data at Abdolkhan and Paye-Pol stations, the daily flow data from the Hamidiyeh hydrometric station, as a neighboring station, was analyzed over a 40-year period. Hyperparameter optimization for these algorithms was carried out using the Optuna method. A thorough comparison of model performance showed that the Xgboost algorithm, by learning complex nonlinear relationships, provided the highest estimation accuracy. The results revealed that at Abdolkhan and Paye-Pol stations, Xgboost achieved the highest efficiency, with the highest coefficient of determination (R²) values of 0.95 and 0.78, the lowest mean absolute error (MAE) values of 18.76 and 36.45, the lowest root mean square error (RMSE) values of 43.75 and 108.87, and the lowest relative root mean square error (RRMSE) values of 0.20 and 0.46, respectively.Furthermore, Taylor diagrams confirmed the superiority of the Xgboost model at both stations. These findings highlight the ability of Xgboost to overcome spatial challenges and handle limited data effectively.The Catboost model achieved second place among the models evaluated, with 11% and 5% lower accuracy compared to the Xgboost model at the Abdolkhan and Pay-Pol stations, respectively. The results of this study can be valuable for estimating missing flow data at other stations of this river and play a significant role in the effective management of water resources.

کلیدواژه‌ها English

Estimation of missing data
Extreme Gradient Boosting
Ensemble learning
Machine learning. Optuna
آهنی، ع. و شوریان، م. 1396. پیش‌بینی جریان ماهیانه رودخانه با استفاده از مدل‌های داده مبنا. تحقیقات منابع آب ایران. سال سیزده، شماره دو، ص 207-214.
بوستانی، م.، کرمی، ح.، موسوی، س. ف.، و فرزین، س. 1398. بررسی ارتباط بین شاخص‌های نظریه آشوب در رفتارنگاری جریان رودخانه‌ای در مقیاس‌های زمانی کوتاه‌مدت. نشریه علمی پژوهشی مهندسی آبیاری و آب ایران، سال نهم، شماره چهار، ص 98-116.
دارابی، ف.، نجفی نژاد، ع.، پورقاسمی، ح. ر. و  سعدالدین، ا. 1403. پیش‌بینی اثر اقدامات بیولوژیک بر سیل‌خیزی حوزه آبخیز بهشت‌آباد با استفاده از روش‌های یادگیری ماشین. مدیریت جامع حوزه­های آبخیز. 10.22034/iwm.2024.2032264.1159
سیدیان، س. م.، سلیمانی، م. و کاشانی، م. 1393. پیش‌بینی دبی جریان رودخانه با استفاده از داده‌کاوی و سری زمانی. اکوهیدرولوژی. سال یک، شماره سه، ص 167-179.
غفاری، غ. ع. و وفاخواه، م. 1392. شبیه‌سازی فرآیند بارش- رواناب با استفاده از شبکه عصبی‌ مصنوعی و سیستم فازی-عصبی تطبیقی (مطالعه موردی: حوزه آبخیز حاجی‌قوشان). پ‍‍ژوهشنامه مدیریت حوزه آبخیز. سال چهار، شماره هشت، ص 136-120.
محمدی، س.، حسن پور، ف.، شریف آذری، س. و فروغی، ف. 1400. ارزیابی روش‌های رگرسیونی نوین جهت تخمین بار رسوبی معلق در رودخانه سیستان. نشریه علمی پژوهشی مهندسی آبیاری و آب ایران. سال دوازده، شماره دو، ص 1-15.
میرنورالهی، ع.، کرمی، ح.، فرزین، س. و عامری، م. 1401. بررسی عملکرد ماشین‌های یادگیری در تخمین ضریب دبی آبگذری آبگیرهای کفی با روزنه دایره‌ای. نشریه علمی پژوهشی مهندسی آبیاری و آب ایران، سال دوازده، شماره چهار، ص 21-41.
یوسفی، ح.، یونسی، ح، ا.، داودی مقدم، د.، ارشیا، آ. و شمسی، ز. 1401. تعیین پتانسیل سیل با استفاده از مدل‌های یادگیری ماشین CART، GLM و GAM مطالعه موردی: حوضه کشکان. نشریه علمی پژوهشی مهندسی آبیاری و آب ایران، سال دوازده، شماره چهار، ص 105-84.
 
Ahadiyan, J. (2016). Application of ANFIS adaptive system to estimate the potential consolidation of clay soils. Journal of Modeling in Engineering, 14(45), 17-31.
Ahadiyan, J., Kiani, S., Asiaban, P., Azizi Nadian, H., & Omidvarinia, M. (2023). Optimizing the dimensions of the agricultural water transfer system from the Karun 3 dam to the northeastern cities of Khuzestan province. Journal of New Approaches in Water Engineering and Environment, 1(2), 112-126.
Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. 2019. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining (pp. 2623-2631).
Bagherian Marzouni, M., Akhoundali, A. M., Moazed, H., Jaafarzadeh, N., Ahadian, J., & Hasoonizadeh, H. (2014). Evaluation of Karun river water quality scenarios using simulation model results. International Journal of Advanced Biological and Biomedical Research, 2(2), 339-318.
Boustani, M., Farzin, S., & Mousavi, S. F. (2025). Estimation of R-Vine copula parameters in multivariate flood frequency analysis using arithmetic optimization algorithm and comparing the performance with genetic algorithm. Water Resources Management, 1-19.
Boser, B. E., Guyon, I. M., & Vapnik, V. N. 1992. A training algorithm for optimal margin classifiers. In Proceedings of the fifth annual workshop on computational learning theory (pp. 144-152).
Breiman, L. 2001. Random forests. Machine Learning, 45, 5-32.
Breiman, L., Friedman, J., & Olshen, R. A. 2017. Classification and regression trees. Routledge.
Chemura, A., Rwasoka, D., Mutanga, O., Dube, T., & Mushore, T. 2020. The impact of land-use/land cover changes on water balance of the heterogeneous Buzi sub-catchment, Zimbabwe. Remote Sensing Applications: Society and Environment, 18, 100292.
Chen, T., & Guestrin, C. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining (pp. 785-794).
Daryaei, M., Kashefipour, S. M., Ahadian, J., & Ghobadian, R. (2010). Modeling the compression index of fine soils using artificial neural network and comparison with the other empirical equations. Journal of Water and Soil, 24(4), 659-667.
Dariane, A. B., & Borhan, M. I. 2024. Comparison of classical and machine learning methods in estimation of missing streamflow data. Water Resources Management, 38(4), 1453-1478.
Fan, J., Wang, X., Wu, L., Zhou, H., Zhang, F., Yu, X., ... & Xiang, Y. 2018. Comparison of Support Vector Machine and Extreme Gradient Boosting for predicting daily global solar radiation using temperature and precipitation in humid subtropical climates: A case study in China. Energy Conversion and Management, 164, 102-111.
Friedman, J. H. 2001. Greedy function approximation: A gradient boosting machine. Annals of Statistics, 1189-1232.
Geurts, P., Ernst, D., & Wehenkel, L. 2006. Extremely randomized trees. Machine Learning, 63, 3-42.
Jabeur, S. B., Gharib, C., Mefteh-Wali, S., & Arfi, W. B. 2021. Catboost model and artificial intelligence techniques for corporate failure prediction. Technological Forecasting and Social Change, 166, 120658.
Katipoğlu, O. M., & Sarıgöl, M. 2023. Prediction of flood routing results in the Central Anatolian region of Türkiye with various machine learning models. Stochastic Environmental Research and Risk Assessment, 37(6), 2205-2224.
Khampuengson, T., & Wang, W. 2023. Novel methods for imputing missing values in water level monitoring data. Water Resources Management, 37(2), 851-878.
Khoramipoor, Z., Valikhan Anaraki, M., & Farzin, S. 2024. A new approach in flood routing based on the integration of bayes theory, support vector machine and meta-heuristic optimization algorithm. Iranian Journal of Irrigation & Drainage, 18(3), 409-420.
Krysanova, V., & White, M. 2015. Advances in water resources assessment with SWAT-an overview. Hydrological Sciences Journal, 60(5), 771-783.
Latifoğlu, L., & Canpolat, Ü. (2022). Prediction of daily streamflow data using ensemble learning models. The European Journal of Research and Development, 2(4), 356-371.
Li, Y., Liang, Z., Hu, Y., Li, B., Xu, B., & Wang, D. 2020. A multi-model integration method for monthly streamflow prediction: Modified stacking ensemble strategy. Journal of Hydroinformatics, 22(2), 310-326.
Minns, A. W., & Hall, M. J. 1996. Artificial neural networks as rainfall-runoff models. Hydrological Sciences Journal, 41(3), 399-417.
Mohammadi, M., Vagharfard, H., Mahdavi Najafabadi, R., Daneshkar Arasteh, P., & Nazemosadat, M. J. 2021. Rainfall-runoff modelling of coastal watersheds near Hormuz Strait using data mining. Iranian Journal of Soil and Water Research, 52(2), 313-327.
Molnar, C. 2020. Interpretable machine learning. Lulu. com.
Ni, L., Wang, D., Wu, J., Wang, Y., Tao, Y., Zhang, J., & Liu, J. 2020. Streamflow forecasting using extreme gradient boosting model coupled with Gaussian mixture model. Journal of Hydrology, 586, 124901.
Quinlan, J. R. 1992. Learning with continuous classes. In 5th Australian joint conference on artificial intelligence (Vol. 92, pp. 343-348).
Razavi, T., & Coulibaly, P. 2013. Streamflow prediction in ungauged basins: Review of regionalization methods. Journal of Hydrologic Engineering, 18(8), 958-975.
Samadi, M., Bahremand, A., & Fathabadi, A. 2019. The Boustan Dam monthly inflow forecasting using data-driven and ensemble models in the Golestan Province. Watershed Engineering and Management, 11(4), 1044-1058.
Schratz, P., Muenchow, J., Iturritxa, E., Richter, J., & Brenning, A. 2019. Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data. Ecological Modelling, 406, 109-120.
Sharma, V., & Yuden, K. 2021. Imputing missing data in hydrology using machine learning models. International Journal of Engineering Research and Technology, 10, 78-82.
Sharififard, E., Azizipour, M., Ahadiyan, J., & Haghighi, A. (2024). Determination of creep function coefficients of viscoelastic pipes using a transient-guided machine learning model. AQUA—Water Infrastructure, Ecosystems and Society, 73(11), 2132-2149.
Sumayli, A. 2023. Development of advanced machine learning models for optimization of methyl ester biofuel production from papaya oil: Gaussian process regression (GPR), multilayer perceptron (MLP), and K-nearest neighbor (KNN) regression models. Arabian Journal of Chemistry, 16(7), 104833.
Taylor, K. E. 2001. Summarizing multiple aspects of model performance in a single diagram. Journal of Geophysical Research: Atmospheres, 106(D7), 7183-7192.
Terzi, Ö., Küçüksille, E. U., Baykal, T., & Taylan, E. D. 2023. Deep and machine learning for daily streamflow estimation: A focus on LSTM, RFR and Xgboost. Water Practice & Technology, 18(10), 2401-2414.
Varga, M., Balogh, S., & Csukas, B. 2016. GIS based generation of dynamic hydrological and land patch simulation models for rural watershed areas. Information Processing in Agriculture, 3(1), 1-16.
Venkatesan, E., & Mahindrakar, A. B. 2019. Forecasting floods using extreme gradient boosting–a new approach. International Journal of Civil Engineering and Technology, 10(2), 1336-1346.
Xie, T., Chen, L., Yi, B., Li, S., Leng, Z., Gan, X., & Mei, Z. 2024. Application of the improved k-nearest neighbor-based multi-model ensemble method for runoff prediction. Water, 16(1), 69.
Yohaness, Y. 1999. Classification and regression tree: An introduction. Research Institute of Washington, DC.
Zhang, Y., Zhao, Z., & Zheng, J. 2020. Catboost: A new approach for estimating daily reference crop evapotranspiration in arid and semi-arid regions of Northern China. Journal of Hydrology, 588, 125087.