Document Type : Research Article
Authors
Department of Geomorphology & Meteorology, Faculty of Geography and Environmental Sciences, Hakim Sabzevari University, Sabzevar, Iran
Abstract
Floods, as one of the most destructive natural hazards, cause irreparable damage to infrastructure and human communities on an annual basis. The present study aims to produce an accurate flood hazard zonation map for the Kashafrud basin by employing two advanced machine learning algorithms, namely Logistic Model Tree (LMT) and Random Forest (RF). In this research, thirteen influential parameters including elevation, curvature, precipitation, drainage density, aspect, soil and geological characteristics, slope, distance from streams, land use, normalized difference vegetation index (NDVI), topographic wetness index (TWI), and stream power index (SPI) were analyzed as input variables. Data corresponding to 145 recorded flood locations were identified and divided into training (70%) and validation (30%) datasets. Following the weighting of the thematic layers, flood susceptibility maps were generated and classified into five hazard classes using the natural breaks classification method. Model performance was evaluated using the Receiver Operating Characteristic (ROC) curve, which indicated that the LMT model, with an AUC value of 0.897, exhibited higher predictive accuracy than the RF model, which achieved an AUC of 0.811. The flood hazard zonation results reveal that high-risk areas are predominantly concentrated in regions characterized by gentle slopes, impermeable formations, low elevations, and floodplains. Areas classified as very high and high risk are mainly located in the central and outlet sections of the basin. These findings not only confirm the effectiveness of machine learning algorithms in flood hazard zonation but also underscore the necessity of strategic management planning and targeted basin management practices in this flood-prone region.
Introduction
Floods are among the most significant and destructive natural hazards worldwide, annually affecting millions of people and causing severe damage to human life, critical infrastructure, agriculture, industry, and both urban and rural areas (Organization, 1999; Das, 2019; Daneshparvar, 2021). In recent years, climate change, rapid population growth, and accelerated urban expansion have contributed to increases in both the frequency and severity of flood events, underscoring the urgent need for effective flood risk management and mitigation strategies. The Kashafrud Basin, located in northeastern Iran, is particularly vulnerable to severe flooding due to its geomorphological characteristics, diverse geological structures, variable precipitation patterns, and intensive human activities. Effective flood risk management in this basin requires accurate identification of flood-prone areas, comprehensive analysis of influential factors, and the application of advanced modeling techniques. Accordingly, this study aims to identify and classify flood-prone areas within the Kashafrud Basin by applying machine learning algorithms specifically the Logistic Model Tree (LMT) and Random Forest (RF) to evaluate the relative contributions of both natural and anthropogenic factors within a data-driven scientific framework.
Material and Methods
In this study, 145 flood and non-flood locations were identified using data obtained from the Watershed Management Organization, complemented by satellite imagery interpretation. The dataset was randomly divided into two subsets, comprising 70% for model training and 30% for validation. Key explanatory variables, including land use, elevation classes, drainage density, lithology, slope, aspect, distance from waterways, rainfall, soil types, and vegetation and moisture indices (TWI, SPI, and NDVI), were derived from topographic maps, geological maps, and 2024 satellite imagery. All spatial datasets were processed and analyzed using ArcGIS 10.4 and ENVI 5.4 software, generating both quantitative and qualitative indicators for each variable.
To address potential multicollinearity among the independent variables, the Variance Inflation Factor (VIF) and Tolerance (TOL) indices were calculated. Variables exhibiting high interdependence were excluded to enhance model reliability and predictive performance. Flood susceptibility modeling was conducted using two advanced machine learning algorithms: the LMT model, which integrates decision tree structures with logistic regression, and the RF model, which employs an ensemble of decision trees. Model performance was evaluated using sensitivity, specificity, the Kappa coefficient, the Receiver Operating Characteristic (ROC) curve, and the Area Under the Curve (AUC). Flood hazard zonation maps were produced using the natural breaks classification method, and the model outputs were converted into thematic flood risk maps.
Results and Discussion
Multicollinearity analysis confirmed that none of the selected variables exceeded the threshold values (VIF > 10 or TOL < 0.1), indicating a robust and stable modeling framework. The results identified thirteen key factors including slope, elevation, curvature, distance from rivers, drainage density, rainfall, soil type, land use, and vegetation and moisture indices as significant determinants of flood risk. Areas characterized by low to moderate slopes and elevations below 1000 m were found to exhibit increased surface runoff velocity, thereby intensifying flood risk and severity, particularly in the central and outlet sections of the basin. Terrain curvature and dense drainage networks significantly influenced runoff concentration and flow pathways, highlighting the critical role of topography and drainage structure in flood dynamics.
Resistant geological formations and impermeable clay-rich soils substantially reduced infiltration capacity, leading to increased surface runoff and the occurrence of sudden and hazardous flood events. In addition, weak vegetation cover and inappropriate land use changes particularly urban and industrial development indirectly exacerbated both runoff volume and flow velocity. Low NDVI values reflected degraded vegetation conditions and a heightened susceptibility to flooding.
Flood hazard zonation maps generated using the LMT model indicated that approximately 81.23% of the basin falls within moderate to very high flood risk classes. Similarly, the RF model estimated that 68.44% of the study area is exposed to comparable risk levels, with high-risk zones predominantly concentrated in the central and outlet regions of the basin. Statistical validation demonstrated that the LMT model achieved higher predictive accuracy (AUC = 0.897) compared to the RF model (AUC = 0.811), although both models exhibited satisfactory and reliable performance for flood susceptibility assessment. High sensitivity, specificity, and Kappa coefficient values further confirmed the robustness and operational applicability of both modeling approaches.
The findings indicate that geomorphological and hydrological factors such as slope, elevation, drainage density, soil characteristics, and impermeable geological formations play a more dominant role in flood occurrence than anthropogenic factors, including land use and human-induced disturbances. Nevertheless, the influence of human activities on flood intensity and spatial distribution remains significant, emphasizing the necessity for integrated and sustainable watershed management strategies. These results are consistent with recent international studies, which highlight the primary contribution of natural variables to flood risk while acknowledging the amplifying effects of human interventions.
Conclusion
By integrating remote sensing data, GIS-based spatial analysis, and advanced machine learning algorithms, this study presents a comprehensive and innovative framework for flood risk assessment, analysis, and spatial zonation in one of the most critical watersheds in northeastern Iran. The results demonstrate that the LMT model outperforms the RF model in accurately identifying flood-prone areas within the Kashafrud basin. The resulting flood hazard maps provide valuable decision-support tools for prioritizing crisis management actions, land use planning, and the development of resilient water infrastructure.
While accurate consideration of natural and geomorphological factors is fundamental to effective flood risk management, regulating human activities such as land use planning, construction of flood resilient infrastructure, restoration of vegetation cover, and implementation of early warning systems is equally essential for minimizing flood-related damages. Emphasizing data-driven and machine learning based approaches, this research contributes to bridging the gap between applied scientific research and practical flood management needs in Iranian catchments. Future studies incorporating advanced algorithms, uncertainty analysis, and higher-resolution satellite data are strongly recommended to further enhance the resilience of aquatic and urban ecosystems.
Keywords
Subjects