Document Type : Research Article
Author
Geography Department, Faculty of Social Science and Humanities, University of Mazandaran, Babolsar, Iran
Abstract
Landslides are among the most significant geomorphological hazards, posing serious threats to transportation networks and infrastructure. This study aims to compare the performance of a reference statistical method, Logistic Regression (LR), and a machine learning algorithm, Random Forest (RF), for spatial landslide susceptibility mapping along the Haraz Road corridor in Mazandaran Province, northern Iran. In this study, a balanced dataset consisting of 10,000 pixels representing landslide and non-landslide locations, together with a set of environmental conditioning factors associated with landslide occurrence, was prepared. To improve model stability, ten independent training datasets and ten independent testing datasets were generated and used in the modelling process for both approaches. Model performance was evaluated using overall accuracy, the Kappa coefficient, and the area under the receiver operating characteristic curve (AUC). Finally, the best-performing models from both approaches were selected based on the AUC values obtained from the testing datasets, and landslide susceptibility maps were produced accordingly. The results indicate that both methods exhibit considerable capability in predicting landslide occurrences. The LR model achieved AUC values of 0.90 for the training datasets and 0.89 for the testing datasets, whereas the RF model yielded corresponding values of 0.93 and 0.89. Analysis of factor importance revealed that lithology and land use are the most influential factors controlling landslide occurrence along the Haraz Road. Moreover, the spatial distribution of susceptible zones in the produced landslide susceptibility maps shows a notable degree of consistency, suggesting that the quality and characteristics of input data may play a more decisive role than the complexity of the modelling algorithm. The findings of this study can contribute to landslide hazard management and to improving safety conditions along the Haraz Road.
Introduction
Landslides are among the most significant natural disasters affecting susceptible mountainous regions, causing substantial damage to infrastructure, environment, and human life. Reliable identification of landslide-prone areas is therefore essential for hazard and risk management.
The Haraz Road in northern Iran represents one of the most critical transportation routes connecting capital to Mazandaran province, and it is frequently affected by landslides due to susceptible geological structures and lithology, steep slopes, climatic variability, and human interferences. These conditions highlight the necessity of developing accurate and scientifically robust landslide susceptibility models for effective risk reduction and planning.
In recent decades, landslide susceptibility modelling has increasingly relied on quantitative modelling approaches, particularly statistical approaches. Logistic Regression (LR) has been widely used as a reference statistical technique due to its simplicity, interpretability, and solid probabilistic foundation. In recent years, machine learning algorithms such as Random Forest (RF) have gained attention because of their ability to model nonlinear relationships, handle complex interactions among variables, and achieve high predictive accuracy without strict statistical assumptions. Despite the growing popularity of machine learning approaches, there remains ongoing debate regarding whether advanced algorithms significantly outperform traditional statistical approaches when datasets are properly prepared and validated.
Therefore, the main objective of this research is to implement a comparative evaluation of Logistic Regression and Random Forest models for spatial landslide susceptibility modelling along the Haraz road. The specific aims include: (1) evaluating the predictive performance and accuracy of both models using the same multiple training and testing datasets, (2) identifying the relative importance of environmental conditioning factors affecting landslide occurrence, and (3) generating landslide susceptibility maps suitable for hazard management and infrastructure planning.
Material and Methods
In order to model spatial landslide susceptibility using LR and RF, a comprehensive landslide inventory was prepared from field studies and existing geological data sources. Based on this inventory, an equal and balanced dataset consisting of 10,000 pixels was prepared, including 5,000 pixels with landslides and 5,000 non-landslide pixels. This balanced dataset was used to reduce classification bias and improve model reliability. Also, a set of environmental factors related to the 10000 pixels was used including lithology, land use, normalized difference vegetation index (NDVI), slope gradient, topographic position index (TPI), topographic wetness index (TWI), stream power index (SPI), plan and profile curvature, aspect-derived indices (northness and eastness), and distance-based factors such as proximity to roads, rivers, and faults.
To enhance model robustness and reduce uncertainty associated with random sampling, ten independent training datasets and ten corresponding testing datasets were generated using repeated random partitioning from the balanced dataset. Then, both LR and RF were implemented using identical datasets to ensure fair comparison. Logistic Regression modelling was performed using a binomial generalized linear modelling framework, while Random Forest modelling involved ensemble classification through multiple decision trees. For the Random Forest algorithm, hyperparameter tuning was performed to specify the optimal number of trees and the number of predictor variables per node of trees, thereby improving predictive performance and minimizing overfitting.
At the end, model performance was evaluated using multiple statistical metrics, including overall accuracy, Cohen’s kappa coefficient, and the area under the receiver operating characteristic curve (AUC). These metrics were computed for each of the ten training and testing datasets, and the mean and standard deviation values were reported to assess model stability and reliability. The best-performing model for each approach was selected based on the highest AUC values obtained from validation datasets and based on that, landslide susceptibility maps were generated.
Results and Discussion
The results indicate that both LR and RF approaches achieved high predictive performance with strong classification capability. The mean AUC values of approximately 0.90 for 10 training datasets and 0.89 for 10 testing datasets in LR, demonstrate very good performance and accuracy. The RF approach achieved slightly higher training AUC values (approximately 0.93), while testing AUC values remained similar to LR (approximately 0.89), indicating comparable generalization performance between the two Approaches. Accuracy and kappa statistics also showed consistent results across the 10 training and 10 testing datasets for both models, confirming the robustness and reliability of the employed approaches.
Results related to the variable importance revealed that lithology and land use were the most influential factors in both approaches affecting landslide occurrence in the Haraz road. The dominance of lithology highlights the critical role of weak geological formations and weathered materials such as Shemshak formation and Scree deposits in controlling landslide susceptibility along the Haraz road.
The landslide susceptibility maps generated from the best-performing models in LR and RF showed similar spatial patterns, with high and very high susceptibility zones concentrated along steep slopes within regions characterized by weak lithological formations. The similarity between maps generated by LR and RF suggests that both models successfully captured the dominant controlling factors and their spatial relationships with landslides distribution along the Haraz road.
The relative similar predictive performance obtained by LR and RF in this study suggests that data quality, balanced sampling design, and rigorous validation procedures may play a more critical role than algorithm complexity alone in achieving reliable landslide susceptibility prediction. This finding supports previous research findings indicating that advanced machine learning methods do not always significantly outperform well-established statistical models when datasets are properly prepared. Therefore, LR remains a valuable reference approach, while RF provides complementary advantages in handling nonlinear relationships and variable interactions.
Conclusion
This study demonstrates that both LR and RF models provide reliable and robust outcomes for landslide susceptibility modelling. Although RF exhibited slightly higher performance during training phase, both models yielded similar predictive accuracy on independent testing datasets, indicating similar practical applicability for landslide susceptibility assessment. The selection of lithology and land use as dominant landslide influencing factors emphasizes the importance of geological formation and human land-use practices in influencing slope instability along mountainous roads.
From an applied point of view, the susceptibility maps generated in this research offer valuable tools and information for infrastructure management, land use planning, and disaster risk reduction strategies. The decision makers and planners can use these maps to prioritize landslides monitoring programs, engineering interventions, and land use regulation in high and very high risk zones, thereby improving road safety and reducing socio-economic and environmental losses.
Future research could consider dynamic triggering factors such as rainfall intensity, temporal land-use changes, and detailed topographic changes through radar remote sensing imageries to improve predictive capability and support early warning landslide systems. Additionally, the integration of remote sensing time-series data, advanced hybrid models, and deep learning approaches may further improve the performance and accuracy of landslide susceptibility models and provide a more reliable foundation for landslide hazard and risk assessments in complex mountainous environments.
Keywords
Subjects