Document Type : Research Article
Authors
1
Department of Geography, Faculty of Literature and Humanities, Razi University, Kermanshah, Iran
2
Surveying of Department, Yasouj University, Yasouj, Iran
3
Researcher at Sejong University, Seoul, South Korea
Abstract
In this study, a drought susceptibility map for Khuzestan Province, Iran, was produced using the Random Forest (RF) algorithm combined with the Shapley Additive Explanations (SHAP) interpretability approach. Drought occurrence data were derived from the Standardized Precipitation Index (SPI) for the period 2018–2022. A total of 18 environmental and climatic factors-including relative humidity, wind speed, evapotranspiration, minimum and maximum temperatures, elevation, slope, slope aspect, topographic wetness index (TWI), land cover, normalized difference vegetation index (NDVI), soil water content, river density, water level, sandy soil, soil bulk density, clay content, and soil texture-were used as input variables for modeling drought susceptibility. Model performance was evaluated using the Receiver Operating Characteristic (ROC) curve, yielding an Area Under the Curve (AUC) value of 0.987, which demonstrates the excellent predictive performance of the Random Forest algorithm. According to the drought susceptibility map, 31.03% of the study area falls in the “very low” class, 8.18% in “low”, 8.30% in “moderate”, 10.88% in “high”, and 41.58% in the “very high” susceptibility class. The highest spatial frequency was observed in the “very low” and “very high” categories. Based on SHAP results, elevation, maximum temperature, topographic wetness index, wind speed, and slope were identified as the most influential factors contributing to drought occurrence. This research highlights the advantages of interpretable machine learning approaches in drought susceptibility assessment, providing valuable insights for environmental planners and decision-makers to mitigate the adverse impacts of drought.
Extended Abstract
Introduction
The frequent occurrence of droughts in recent decades has posed significant challenges to the sustainability of ecosystems, agriculture, and socio-economic development. Drought is defined as a prolonged deficiency of precipitation over a specific period and represents a complex phenomenon characterized by reductions in soil moisture, declines in surface water flows, decreases in groundwater levels, and simultaneous increases in temperature. Droughts are typically classified into meteorological, agricultural, hydrological, socio-economic, and environmental types, and they exert both direct and indirect impacts, including global warming, deforestation, and urbanization. Machine learning algorithms, including SVM, BRT, RF, XGBoost, KNN, and CNN, exhibit strong capabilities in drought prediction, particularly the Random Forest (RF) algorithm, which offers high accuracy and computational efficiency. However, the “black-box” nature of these models limits physical interpretability, making the application of explainable artificial intelligence (XAI) techniques and the Shapley method essential for analyzing both the positive and negative effects of features and their interactions. Recent studies have demonstrated that integrating the Random Forest (RF) algorithm with the Shapley method enhances both the accuracy of drought prediction and the analysis of climatic, soil, topographic, and socio-economic factors. This study aims to improve the interpretability of the RF algorithm and to develop a drought sensitivity map for Khuzestan Province, with its novelty lying in the application of a spatially interpretable approach to analyze the factors influencing drought occurrence.
Material and Methods
Khuzestan Province, located in southwestern Iran, covers an area of approximately 64,236 km² with elevations ranging from 0 to 3,740 m, and experiences a climate spectrum from arid to humid. Despite its major rivers and extensive water resources, the region has been increasingly affected by environmental crises due to recurrent droughts and overexploitation of water. The development of water-intensive industries, expansion of large-scale agricultural lands, and cultivation of high-water-demand crops have intensified pressure on water resources. Additionally, declining precipitation in neighboring provinces and increased groundwater extraction have exacerbated drought conditions, impacting approximately 98.7% of the province between 2012 and 2021.
To model drought susceptibility, a total of 19 climatic, topographic, soil, and hydrological variables were considered, including precipitation, relative humidity, wind speed, evapotranspiration, minimum and maximum temperatures, elevation, slope and aspect, topographic wetness index, land cover, vegetation index, soil water content, river density, water table level, sand and clay fractions, bulk density, and overall soil texture. Climatic data were derived from the annual averages of 22 meteorological stations across Khuzestan Province (2018–2022), while soil data were obtained from the USDA database via Google Earth Engine (GEE). Elevation, slope, slope direction, and watershed power index layers were generated from the SRTM digital elevation model with a resolution of 30 x 30 meters, and the vegetation index was generated from NDVI based on Landsat 8 images from 2019 to 2022. River density was calculated using the Line Density tool in ArcMap, and the water surface layer was extracted from well water table depth data (piezometric data from the National Water Resources Management Company). The land cover layer was generated by integrating Sentinel-1 and Sentinel-2 imagery based on the 13 classes defined by Ghorban et al. (2020). Subsequently, all layers were transformed into raster format with a spatial resolution of 250 m using the Inverse Distance Weighting (IDW) method to facilitate drought modeling.
The Standardized Precipitation Index (SPI) was employed to produce the drought map, as it quantifies precipitation deficits across various time scales and is recognized as a reliable indicator for drought monitoring. Drought susceptibility was modeled using the Random Forest algorithm, which is capable of handling large, multivariate datasets while providing robust predictive performance. To enhance model interpretability, the Shapley Additive Explanations (SHAP) method was applied to assess the contribution of each input variable to drought prediction. Model accuracy was assessed using the Root Mean Square Error (RMSE) and Mean Absolute Error (MAE) indices, while the coefficient of determination (R²) was calculated as a dimensionless measure of model fit. The performance of the drought susceptibility map was further evaluated using the Receiver Operating Characteristic (ROC) curve and the Area Under the Curve (AUC) index, where values approaching 1 indicate a strong capability of the model to accurately predict vulnerable areas.
Results and Discussion
In this study, the Random Forest model was employed to predict drought occurrences in Khuzestan Province based on environmental and climatic variables. Model performance was assessed using a 10-fold cross-validation approach, with 70% of the data allocated for training and the remaining 30% for testing. The results demonstrated that the model exhibited excellent performance, with an R² of 0.925, MAE of 0.035, and RMSE of 0.137, indicating a high level of accuracy in drought prediction. Interpretability analysis using the Shapley method revealed that precipitation, elevation, topographic wetness index, relative humidity, and wind speed were the most influential factors, with minimum precipitation exerting a particularly significant effect. The model output was generalized in ArcMap software and using the natural failure method, the drought sensitivity map was classified into five classes: "very low, low, medium, high, and very high." The results showed that 45.26% of the area was in the "very low" class, 4.25% in "low," 8.5% in "medium," 25.8% in "high," and 36.4% in "very high." An AUC value of 0.972 further confirmed the model’s outstanding performance. These results highlight the Random Forest model as a powerful tool for drought identification, capable of capturing complex and nonlinear relationships while processing both numerical and categorical data, thereby providing effective support for environmental management and decision-making.
Conclusion
Recent drought events have inflicted substantial impacts on both the environment and human activities. In this study, a drought susceptibility map for Khuzestan Province was developed using the Random Forest model in combination with the Shapley Additive Explanations (SHAP) method for interpretability. The model demonstrated excellent predictive performance, with an AUC of 0.972. Analysis revealed that precipitation, elevation, relative humidity, the topographic wetness index, and wind speed were the most influential factors driving drought occurrence, whereas soil texture and land cover exhibited comparatively lower impacts. The most prevalent categories in the drought susceptibility map were “very low” and “very high.” Areas including Andimeshk and Dezful were classified as “very low,” while parts of Mahshahr and Ahvaz fell into the “very high” category. Drought susceptibility mapping serves as an effective tool for identifying vulnerable regions and informing management strategies, thereby enhancing the resilience of both ecosystems and human communities.
Keywords
Subjects