1. Introduction
Landslides are major natural hazards that cause substantial loss of life worldwide [
1]. Mountainous regions of southwestern China are particularly susceptible to slope instability because of their steep relief, tectonic activity, and strongly dissected terrain [
2,
3].
The Wumeng Mountainous Region is characterized by deeply dissected mountainous terrain and complex geological conditions. In Zhenxiong County, Wu et al. [
4] described the source area of the catastrophic landslide that occurred on 22 January 2024 as exhibiting a “boot-shaped terrain” configuration. This regional observation motivates the terrain concept investigated in the present study. Rather than treating this morphology as a deterministic precursor of slope instability, we abstract it as a steep-upper–gentle-lower longitudinal terrain pattern and formalize it as an operational geomorphological screening criterion.
Zhenxiong County in northeastern Yunnan Province was selected as the study area. The catastrophic landslide of 22 January 2024 in Liangshui Village resulted in severe casualties and highlighted the need for improved screening of landslide-prone terrain [
4]. The ability to identify terrain warranting further investigation before slope failure is important for proactive disaster-risk reduction and land-use planning [
5].
Conventional approaches to landslide assessment include inventory- and geomorphology-based investigation [
6,
7] and data-driven susceptibility modelling using conditioning factors such as slope, aspect, lithology, land use, and precipitation [
8,
9]. Statistical and machine-learning methods, including logistic regression [
10], support vector machines [
11], and random forests [
12], have been widely used for landslide susceptibility mapping, generally producing spatial susceptibility estimates from predefined conditioning variables.
The emergence of deep learning, and semantic segmentation in particular, has opened new possibilities for pixel-level landslide detection and mapping from remote-sensing imagery [
13,
14]. Architectures such as U-Net [
15], DeepLabV3+ [
16], and Transformer-based segmentation models [
17,
18] provide widely used methodological foundations for this task. Chen et al. [
19] demonstrated automated landslide detection from multi-temporal remote-sensing imagery in mountainous urban settings, while Ji et al. [
20] introduced the open Bijie landslide dataset containing optical imagery, DEM information, and interpreted landslide boundaries. Recent studies have further explored hybrid CNN–Transformer architectures [
21], multi-scale feature fusion [
22], attention-based segmentation [
23], and learning strategies for class-imbalanced landslide mapping [
24].
Despite these advances, three limitations remain relevant to the present task. First, many landslide-detection studies rely strongly on optical remote-sensing imagery, whose effectiveness may be reduced by cloud cover, vegetation occlusion, and limited visibility of underlying terrain [
18,
25]. Second, supervised deep-learning methods depend on reliable labelled inventories, whereas landslide inventories can be incomplete, spatially biased, and inconsistent among mapping sources [
6,
26]. Third, purely data-driven workflows do not necessarily encode explicit geomorphological constraints, and biases in the spatial distribution or collection of landslide labels can affect both model predictions and their physical interpretation [
27]. These considerations motivate the use of topographic data and explicit geomorphological screening in the present study.
Airborne LiDAR-derived DEMs provide a complementary source of information because they describe bare-earth topography beneath vegetation and can reveal terrain morphology that may be difficult to interpret directly from optical imagery [
28]. Recent deep-learning studies have increasingly incorporated DEM-derived terrain information into landslide detection. For example, Li et al. [
25] proposed the multimodal DemDet framework combining LiDAR-derived terrain information with optical imagery for forested landslide detection, while Liu et al. [
29] developed a feature-fusion segmentation network integrating high-resolution remote-sensing imagery with DEM-derived terrain features. Hybrid approaches have also combined satellite imagery with DEM-derived topographic channels [
21]. These studies demonstrate the value of topographic information as a complement to spectral imagery, while motivating further investigation of whether DEM-derived terrain variables alone can support approximation of a specifically defined geomorphological screening pattern.
The second limitation, namely the scarcity and inconsistency of labeled landslide data, motivates the use of rule-derived candidate masks as an alternative source of supervision. In many mountainous regions, comprehensive and spatially consistent landslide inventories are difficult to obtain, and available inventories are often biased toward post-event mapping or visually evident surface failures. This limitation is particularly relevant when the research target is not a complete landslide inventory, but a specific terrain morphology that can be described by expert geomorphological knowledge. In this context, geomorphologically interpretable terrain-screening rules provide a practical way to generate candidate labels from DEM-derived topographic information, thereby reducing the dependence on large manually interpreted landslide inventories.
The third limitation, namely the lack of explicit geomorphological priors in many deep learning workflows, motivates the use of slope units as the basic terrain analysis units. Slope units are terrain compartments bounded by ridge and valley lines that represent natural geomorphological units over which gravitational mass movements may initiate and propagate [
30,
31]. Compared with regular grid cells, slope-unit-based approaches provide physically meaningful boundaries, reduce within-unit topographic heterogeneity, and better capture whole-slope processes [
32]. Recent studies have begun to combine slope units with deep learning; for example, Wang et al. [
33] applied graph convolutional networks to slope-unit-based susceptibility mapping by treating each slope unit as a node in a geomorphological graph. However, few studies have explicitly used slope-unit-level geomorphological rules to generate rule-derived candidate masks and then trained deep learning models to approximate such rule-based terrain-screening outputs.
This study formalizes the boot-shaped terrain description reported for the Zhenxiong landslide [
4] as a quantitative geomorphological screening criterion. The targeted pattern is represented as a steep-upper–gentle-lower transition along a representative longitudinal slope-unit profile. Rather than assuming that this morphology constitutes an independently validated precursor of landslide occurrence, the proposed rule defines a reproducible terrain-candidate class for subsequent analysis. By coupling this explicit screening procedure with semantic segmentation, the framework uses the resulting rule-derived masks as structured supervisory targets and evaluates whether their spatial patterns can be approximated directly from DEM-derived raster inputs. The segmentation model is therefore intended as a raster-based surrogate of the upstream screening procedure rather than as an independently supervised landslide detector.
The specific objectives of this study are as follows:
- (1)
To formalize boot-shaped terrain morphology as a parameterized screening criterion based on representative longitudinal profiles of slope units, and to quantify the sensitivity of the resulting rule-derived candidate definition to the screening parameters;
- (2)
To compare four representative semantic segmentation architectures (U-Net, U-Net++, DeepLabV3+, and SegFormer-B0) for approximating the rule-derived boot-shaped terrain masks from DEM-derived raster inputs;
- (3)
To evaluate the contribution of DEM-derived terrain variables through controlled input-feature comparisons and to identify a parsimonious raster-input configuration for subsequent surrogate-model evaluation;
- (4)
To evaluate model–rule agreement under complementary spatial settings, including pixel- and patch-level evaluation on the controlled patch dataset, cross-subregion transfer, and whole-area evaluation under the natural prevalence of rule-derived candidates.
The remainder of this paper is organized as follows.
Section 2 describes the study area and the LiDAR DEM data.
Section 3 presents the methodology, including the boot-shaped terrain screening procedure, dataset construction, and deep learning segmentation architecture.
Section 4 presents the qualitative and quantitative results, including architecture, input-feature, and loss-function comparisons, sensitivity analysis of the rule-derived candidate definition, cross-subregion transfer, and whole-area natural-prevalence evaluation.
Section 5 discusses the findings, limitations, and future research directions.
Section 6 summarizes the main conclusions.
4. Results
The results are presented from qualitative observations to controlled factor-wise comparisons and increasingly stringent spatial evaluations. Unless otherwise stated, the segmentation metrics are calculated against the rule-derived reference masks and therefore quantify model–rule agreement rather than accuracy against independently verified landslide ground truth. The qualitative overview is presented first to provide physical and spatial context for the subsequent quantitative analyses, followed by evaluations of model architecture, terrain input composition, loss function, screening-parameter sensitivity, cross-subregion transfer, and whole-area performance under natural candidate prevalence.
4.1. Qualitative Overview of Terrain Context and Model Predictions
Field and aerial photographs from representative locations within the investigated area provide visual context for the mountainous terrain from which the LiDAR data were acquired (
Figure 5a,b). The investigated landscape is characterized by strongly dissected hillslopes, locally steep rock faces, and strong topographic contrasts between hillslopes and lower-gradient slope-foot or valley-floor terrain. A ground-based photograph from Tangfang Township and an aerial photograph from Pingshang Township illustrate the regional geomorphological setting in which the boot-shaped terrain screening task was investigated. These photographs are included to document the physical landscape and facilitate interpretation of the DEM-derived terrain morphology; they are not used as independent validation of the rule-derived reference masks or model predictions.
Figure 5c,d show two illustrative positive examples from the held-out test partition of the final DeepLabV3+ configuration, using DEM and slope as input features and Dice + Focal as the loss function. For each example, the DEM, rule-derived reference mask, predicted positive-class probability, and thresholded model–rule agreement are presented side by side. The examples were selected to provide visually clear cases of model–reference spatial correspondence; they are intended for qualitative illustration only, whereas quantitative performance is evaluated using the complete test partitions in the subsequent analyses.
The probability maps visualize the continuous reference-positive scores produced by the model, whereas the agreement maps show the spatial relationship between the thresholded predictions and the rule-derived reference masks. In the agreement maps, true-positive (TP) regions denote pixels identified as positive by both the model and the rule-derived reference, false-positive (FP) regions denote model-positive pixels outside the reference mask, and false-negative (FN) regions denote reference-positive pixels not reproduced by the model. The examples illustrate that the model can reproduce the principal spatial pattern of the rule-derived targets while disagreement remains along parts of the candidate boundaries and in adjacent terrain.
Such disagreement should not be interpreted directly as geological misclassification. A model-positive region outside the rule-derived reference does not establish the presence of an independently verified landslide or potential hazard, and a reference-positive region missed by the model does not necessarily indicate failure to detect a confirmed landslide. Because independent field-verified boundaries are unavailable for the individual rule-derived candidates, the qualitative comparison demonstrates model–rule spatial agreement rather than independently verified landslide-detection accuracy.
The qualitative overview therefore links the raster-based segmentation task to the physical terrain setting of the study area and illustrates the spatial form of agreement and disagreement between the learned model and the screening-derived reference. Quantitative differences among the candidate segmentation architectures are evaluated next under the common experimental configuration.
4.2. Architecture Comparison Under the Common Experimental Configuration
The four segmentation architectures were compared under the common experimental configuration defined in
Section 3.3, with DEM + slope-gradient rasters as the input and Dice + Focal as the loss function. Each architecture was evaluated using the same three spatial folds and three random seeds, yielding nine fold–seed evaluations per model. The mean and standard deviation across these nine evaluations is reported descriptively in
Table 3, whereas statistical inference was conducted using the three spatial folds as the inferential blocks, as described in
Section 3.3.5.
DeepLabV3+ achieved the highest mean segmentation performance among the four evaluated architectures, with a pixel-level F1-score of 0.351 ± 0.036, positive-class IoU() of 0.214 ± 0.026, and mIoU of 0.557 ± 0.014. U-Net ranked second in mean F1 (0.259 ± 0.054), followed by U-Net++ (0.226 ± 0.057) and SegFormer-B0 (0.118 ± 0.046). DeepLabV3+ also yielded the highest mean precision (0.328 ± 0.032) and recall (0.394 ± 0.088) among the evaluated architectures.
The fold-blocked permutation analysis indicated an overall architecture effect on pixel-level F1 (Q = 9.00, permutation p = 0.0017, Kendall’s W = 1.00). DeepLabV3+ achieved a higher fold-level mean F1 than U-Net, U-Net++, and SegFormer-B0 in all three spatial folds, with mean paired differences of +0.093, +0.125, and +0.234, respectively. However, because only three independent spatial folds were available, the exact two-sided pairwise sign-flip tests had limited resolution and did not reach statistical significance after multiplicity correction. The architecture ranking is therefore interpreted from the omnibus test together with effect magnitude and fold-wise consistency rather than from pairwise significance alone.
Patch-level scores were higher than the corresponding pixel-level scores for all four architectures. DeepLabV3+ obtained the highest mean Patch-F1 (0.814 ± 0.145), compared with 0.764 ± 0.194 for U-Net, 0.752 ± 0.185 for U-Net++, and 0.624 ± 0.165 for SegFormer-B0. As defined in
Section 3.3.4, these patch-level values are auxiliary metrics obtained from the controlled, approximately balanced patch dataset and should not be interpreted as estimates of whole-area operational performance.
Taken together, the controlled comparison indicates that DeepLabV3+ provided the highest mean model–rule agreement among the four evaluated architectures under the matched experimental setting. Together with the significant omnibus architecture effect and the consistent direction of the fold-level differences, these results support retaining DeepLabV3+ as the final segmentation architecture for the subsequent cross-subregion and whole-area evaluations. However, the comparison evaluates complete model architectures rather than individual architectural components; therefore, the observed performance differences cannot be attributed specifically to ASPP, the ResNet-34 encoder, or any other single design element without additional component-level ablation experiments.
4.3. Contribution of Terrain Input Variables
The contribution of the terrain variables was evaluated using DeepLabV3+ with Dice + Focal loss while varying only the raster-input composition. Six configurations were considered: DEM, Slope, DEM + Slope, DEM + Aspect (sin + cos), Slope + Aspect (sin + cos), and DEM + Slope + Aspect (sin + cos), where Slope denotes the spatially continuous slope-gradient raster described in
Section 3.2.3. Aspect was represented by the complete sine–cosine pair in all configurations in which directional information was included, consistent with its circular encoding. Each configuration was evaluated using the same three spatial folds and three random seeds, and the descriptive results are summarized in
Table 4.
Among the single-variable inputs, Slope yielded higher mean segmentation performance than DEM, with F1-scores of 0.344 ± 0.051 and 0.225 ± 0.055, respectively. Combining DEM with Slope resulted in an F1-score of 0.351 ± 0.036, positive-class IoU () of 0.214 ± 0.026, and mIoU of 0.557 ± 0.014. The highest mean F1 among the six configurations was obtained by Slope + Aspect (sin + cos), at 0.358 ± 0.039, although its numerical advantage over DEM + Slope was small. The full four-channel configuration, DEM + Slope + Aspect (sin + cos), yielded an F1-score of 0.328 ± 0.040, lower than the corresponding DEM + Slope configuration. These results indicate that the contribution of terrain aspect depends on the accompanying terrain variables rather than following a uniform positive or negative pattern.
The fold-blocked permutation analysis indicated an overall effect of input composition on pixel-level F1 (Q = 12.33, permutation p = 0.0053, Kendall’s W = 0.82). However, the limited number of independent spatial folds constrained the resolution of pairwise inference. At the fold level, DEM + Slope exceeded DEM in all three folds, whereas its difference from Slope was small and changed direction across the folds. Similarly, DEM + Slope exceeded Slope + Aspect (sin + cos) in two folds but was lower in the remaining fold. Accordingly, the omnibus result supports an overall input-composition effect, while individual configuration differences are interpreted together with their effect magnitudes and fold-wise consistency rather than as statistically significant pairwise superiority.
The effect of aspect was examined further by comparing each base input with the corresponding configuration containing the complete sine–cosine aspect representation. Adding aspect to DEM reduced the mean F1 by 0.046 and produced lower fold-level F1 in all three spatial folds. Adding aspect to slope produced a modest mean increase of 0.014, with improvement in two of the three folds. In contrast, adding aspect to DEM + Slope reduced mean F1 by 0.024, with lower F1 in two of the three folds. Thus, the experiments do not support describing aspect as uniformly beneficial, uniformly detrimental, or irrelevant. Instead, its incremental contribution was configuration-dependent.
Although Slope + Aspect (sin + cos) produced the numerically highest mean F1, its advantage over DEM + Slope was only approximately 0.007, and the direction of this difference was not consistent across the three spatial folds. DEM + Slope was therefore retained as the final parsimonious two-channel input configuration for the subsequent cross-subregion and whole-area evaluations. It provided competitive pixel-level agreement with fewer input channels, exhibited relatively low descriptive variability, and avoided reliance on an incremental aspect contribution that was not consistent across the evaluated base-input configurations. This selection should therefore be interpreted as a balance among predictive performance, input dimensionality, and fold-wise consistency rather than as evidence that DEM + Slope was statistically or universally superior to the alternative input configurations.
4.4. Loss-Function Evaluation
The effect of the optimization objective was evaluated using DeepLabV3+ with the DEM + slope-gradient input while varying only the loss function. Six loss formulations were compared: Cross-Entropy, Weighted Cross-Entropy, Focal loss, Dice loss, Dice + Cross-Entropy, and Dice + Focal. Each loss function was evaluated using the same three spatial folds and three random seeds, resulting in nine fold–seed evaluations per configuration. The resulting descriptive performance is summarized in
Table 5. The nine-run mean ± SD values are reported descriptively, whereas statistical inference was conducted using the three spatial folds as the inferential blocks, as described in
Section 3.3.5.
Dice + Focal achieved the highest mean pixel-level F1-score among the six evaluated loss functions (0.351 ± 0.036), together with the highest positive-class IoU (0.214 ± 0.026) and mIoU (0.557 ± 0.014). Weighted Cross-Entropy produced the second-highest mean F1 (0.333 ± 0.058) and the highest mean recall (0.496 ± 0.134), although its mean precision was lower (0.262 ± 0.040). Dice + Cross-Entropy yielded an intermediate F1 of 0.309 ± 0.071, while Dice and standard Cross-Entropy produced mean F1-scores of 0.297 ± 0.070 and 0.275 ± 0.100, respectively. Standalone Focal loss showed the lowest mean F1 (0.198 ± 0.065) and positive-class IoU (0.111 ± 0.041) under the present training configuration.
The relative ordering differed across individual metrics. Dice + Focal achieved the highest mean precision, F1, positive-class IoU, and mIoU, whereas Weighted Cross-Entropy achieved the highest recall. Dice + Cross-Entropy produced the highest mean Patch-F1 (0.831 ± 0.051), slightly exceeding Dice + Focal (0.814 ± 0.145) and Weighted Cross-Entropy (0.813 ± 0.130). This difference between pixel- and patch-level rankings further indicates that the loss functions influence boundary-level agreement and patch-level candidate presence differently.
Despite these descriptive differences, the fold-blocked omnibus analysis did not provide evidence of a statistically significant overall loss-function effect on pixel-level F1 (Q = 7.76, permutation
p = 0.168, Kendall’s W = 0.52). The statistical analysis treated the three spatial folds as the independent blocks after averaging the three random seeds within each fold, consistent with the procedure described in
Section 3.3.5. Because only three independent spatial blocks were available, the corresponding pairwise exact sign-flip tests had limited inferential resolution.
At the fold level, Dice + Focal achieved a higher mean F1 than Weighted Cross-Entropy in all three spatial folds, with an average paired difference of approximately +0.018. It also exceeded standalone Focal and Dice in all three folds, whereas its differences from Cross-Entropy and Dice + Cross-Entropy were not directionally consistent across every fold. None of the planned pairwise comparisons reached statistical significance after multiplicity correction. Therefore, the loss-function results do not support claiming statistical superiority of Dice + Focal over every alternative objective.
Although the omnibus comparison did not establish a statistically significant overall loss-function effect, Dice + Focal yielded the highest descriptive mean pixel-level F1, , and mIoU among the six evaluated objectives. It was therefore retained as the final loss formulation for the subsequent cross-subregion and whole-area evaluations. This selection was based on its overall pixel-level descriptive performance under the present dataset and experimental protocol and should not be interpreted as evidence that Dice + Focal is statistically or universally superior to the alternative loss formulations. The higher Patch-F1 obtained by Dice + Cross-Entropy further indicates that the preferred loss formulation depends in part on the evaluation scale and metric of interest.
4.5. Sensitivity of the Rule-Derived Candidate Definition
The sensitivity analysis examined how the three boot-shaped profile parameters—the upper-segment angle threshold
, lower-segment angle threshold
, and profile partition ratio r affected both the rule-derived candidate definition and its subsequent learnability by the segmentation model. Nine predefined parameter combinations were evaluated, with T2 (45°, 30°, 0.6) retained as the reference configuration. Because each parameter combination generates a different set of reference masks, the resulting F1 and
values represent model agreement with different rule-derived targets and therefore cannot, by themselves, be interpreted as evidence that one parameter setting is physically more valid than another. The parameter configurations, candidate abundance, downstream model–rule agreement, and spatial-overlap statistics are summarized in
Table 6.
As summarized in
Table 6, the nine parameter combinations produced marked differences in candidate abundance. The number of reference-positive patches ranged from 42 under (50°, 30°, 0.6) to 664 under (40°, 30°, 0.6), compared with 207 for the reference configuration. Correspondingly, the total spatial extent of the rule-derived candidate masks varied substantially. Relative to the reference configuration, the candidate-area ratio ranged from approximately 0.17 to 3.39, indicating that parameter changes affected not only the number of sampled patches but also the geographic extent of the candidate definition.
Spatial-overlap analysis further demonstrated that the alternative rule configurations did not simply produce minor boundary perturbations around an otherwise stable candidate set. Global Jaccard overlap with the reference configuration ranged from 0.045 to 0.683 among the eight alternative settings. Configuration (45°, 35°, 0.6) showed the greatest spatial agreement with the reference among the alternatives (J = 0.683) and retained all reference-positive pixels while expanding the total candidate area by a factor of approximately 1.46. In contrast, the more restrictive configurations (50°, 30°, 0.6) and (50°, 35°, 0.7) retained only approximately 17.5% and 5.1% of the reference candidate pixels, respectively. The relaxed configuration (40°, 30°, 0.6) retained the complete reference candidate set but expanded the mapped candidate area to approximately 3.39 times that of the reference. These results demonstrate substantial sensitivity of the rule-derived spatial target to the profile parameters.
The downstream segmentation results also varied across parameter settings (
Table 6). The reference configuration produced an F1-score of 0.377 ± 0.018 and
of 0.232 ± 0.013 across the three spatial folds. The numerically highest F1 was obtained under T5 (45°, 35°, 0.6), at 0.391 ± 0.024. However, this setting also expanded the candidate area relative to T2 and therefore represents a different rule-derived target rather than an independently demonstrated improvement in physical validity. When
was increased to 50°, the number of positive patches decreased sharply and segmentation agreement was low: F1 declined to 0.051 ± 0.024 for (50°, 30°, 0.6) and 0.068 ± 0.047 for (50°, 35°, 0.7). Conversely, more permissive configurations generated substantially larger candidate sets but did not necessarily yield higher model–rule agreement; for example, (40°, 30°, 0.6) produced 664 positive patches but an F1 of 0.247 ± 0.121.
The contrasting responses of downstream F1 and the spatial candidate definition across the nine parameter configurations are visualized in
Figure 6.
Changing the partition ratio also materially altered both downstream model–rule agreement and the spatial candidate definition. With = 45° and = 30°, changing r from 0.5 to 0.6 and 0.7 resulted in F1-scores of 0.315 ± 0.037, 0.377 ± 0.018, and 0.332 ± 0.098, respectively. The corresponding spatial Jaccard overlaps of the r = 0.5 and r = 0.7 variants with the reference were only 0.451 and 0.446, showing that even apparently moderate changes in the profile partition alter the spatial candidate definition appreciably.
Taken together, these results indicate that the screening parameters affect two related but distinct properties: the spatial definition of the rule-derived candidate set and the ability of the segmentation model to reproduce that definition. The predefined reference setting T2 (45°, 30°, 0.6) was retained as the configuration used throughout the main experiments because it corresponds to the original expert-informed operational rule, not because the sensitivity analysis demonstrated statistical or physical optimality. The observed variation instead shows that the Stage-1 screening criterion is an operational geomorphological definition whose parameterization materially influences candidate abundance, spatial extent, and downstream model–rule agreement.
4.6. Cross-Subregion Transfer Validation
Cross-subregion transfer was evaluated using the leave-one-subregion-out (LOSO) protocol described in
Section 3.3.3. In each experiment, one complete LiDAR-covered subregion was excluded from model fitting, validation, early stopping, and normalization-statistic estimation, and the trained DeepLabV3+ model was subsequently evaluated on patches from that held-out subregion. The same DEM + slope-gradient input and Dice + Focal loss were used in all four LOSO settings. Each held-out setting was repeated with three random seeds, and the resulting region-level descriptive mean ± SD values are summarized in
Table 7.
Pixel-level model–rule agreement varied across the four held-out subregions. The mean F1-score ranged from 0.303 ± 0.051 in xp04 to 0.342 ± 0.004 in xp01, while positive-class IoU(
) ranged from 0.179 ± 0.035 to 0.206 ± 0.003. The macro-average across the four held-out subregions was 0.321 ± 0.016 for F1, 0.192 ± 0.011 for (
), and 0.534 ± 0.026 for mIoU. For context, the common experimental configuration achieved an F1 of 0.351 ± 0.036 under the spatially grouped three-fold evaluation described in
Section 4.2. The lower LOSO macro-average indicates that reproducing the rule-derived target in a completely withheld subregion was more challenging than evaluation with spatial folds drawn from the broader multi-subregion patch dataset.
The precision–recall balance differed substantially among subregions. xp01 exhibited the highest precision (0.500 ± 0.043) but the lowest recall (0.261 ± 0.015), indicating relatively conservative positive predictions relative to the rule-derived reference. In contrast, xp04 showed the highest recall (0.614 ± 0.106) but the lowest precision (0.205 ± 0.048), indicating a more expansive prediction pattern. xp02 and xp05 occupied intermediate positions, with F1-scores of 0.323 ± 0.015 and 0.319 ± 0.020, respectively. These region-specific differences show that the observed precision–recall balance varied substantially among the four held-out subregions.
Patch-F1 remained numerically high in the LOSO patch evaluation, ranging from 0.840 ± 0.061 to 0.886 ± 0.034, with a macro-average of 0.860 ± 0.022. However, this result should be interpreted in the context of the controlled patch-sampling procedure described in
Section 3.2.2 and
Section 3.3.4. The LOSO test patches retain the constructed patch-level class distribution and the 100-pixel patch-presence criterion; therefore, the high Patch-F1 does not demonstrate equivalent performance under the natural class prevalence of a complete subregion. The corresponding whole-area evaluation is presented separately in
Section 4.7.
The held-out subsets were also strongly unequal in size: xp01 contained 143 test patches and xp02 contained 214, whereas xp04 and xp05 contained only 29 and 20 patches, respectively. Region-specific performance estimates for xp04 and xp05 should therefore be interpreted with particular caution because they were evaluated on substantially smaller held-out patch sets. Moreover, the three seeds within each held-out setting represent repeated stochastic fits to the same spatial test region rather than independent geographic replicates.
Overall, the LOSO experiment provides evidence that the learned segmentation surrogate retains measurable agreement with the rule-derived candidate masks when transferred among the four LiDAR-covered subregions of Zhenxiong County. The experiment should therefore be interpreted as an assessment of within-county cross-subregion transfer, not as evidence of external geographic generalization. Evaluation in independent counties or geologically distinct regions would be required to establish broader transferability.
4.7. Whole-Area Evaluation Under Natural Candidate Prevalence
The patch-based experiments in
Section 4.2,
Section 4.3,
Section 4.4,
Section 4.5 and
Section 4.6 were conducted on intentionally sampled datasets and therefore do not reproduce the natural prevalence of rule-derived candidate pixels across the complete study regions. To evaluate the final segmentation surrogate under complete-region natural candidate prevalence, the LOSO models were applied to the complete LiDAR-covered extent of each held-out subregion. Whole-area inference used overlapping 256 × 256-pixel windows with a stride of 128 pixels, with overlapping reference-positive probabilities fused by arithmetic averaging as described in
Section 3.3.6. The resulting predictions were evaluated against the complete rule-derived reference mask of each held-out subregion.
The natural reference-positive prevalence was low in all four subregions, ranging from 1.02% in xp04 to 1.85% in xp01, with a macro-average of 1.41% (
Table 8). Under the fixed decision threshold of t = 0.5, whole-area model–rule agreement was substantially lower than that observed in the controlled patch experiments. Macro-averaged precision, recall, F1, and positive-class IoU (
) were 0.153, 0.094, 0.072, and 0.038, respectively. Region-level F1 ranged from 0.021 ± 0.006 in xp01 to 0.117 ± 0.040 in xp05. These results show that performance measured on the approximately balanced patch dataset does not directly translate to complete-region evaluation under natural candidate prevalence.
The four subregions exhibited markedly different operating characteristics at the fixed threshold. xp01 produced relatively high precision (0.316 ± 0.086) but extremely low recall (0.011 ± 0.003), indicating that only a small fraction of the complete rule-derived candidate area exceeded the nominal 0.5 decision threshold. In contrast, xp02 and xp04 showed higher recall (0.143 ± 0.026 and 0.130 ± 0.124, respectively) but substantially lower precision. xp05 showed an intermediate precision–recall balance, with an F1 of 0.117 ± 0.040. These regional differences show that the same nominal decision threshold yielded substantially different precision–recall behavior across the four held-out subregions.
To examine whether part of this whole-area performance gap was attributable to operating-threshold transfer, a second evaluation used thresholds selected exclusively from each model’s original LOSO validation split. The mean validation-derived thresholds varied substantially among regions, from 0.327 ± 0.055 for xp01 to 0.643 ± 0.168 for xp02 (
Table 8). Applying these locked thresholds to the complete held-out regions increased the macro-average F1 from 0.072 to 0.095 and
from 0.038 to 0.051, corresponding to an absolute F1 increase of 0.023. The macro-average F1 increased by 0.023, but the absolute whole-area agreement remained limited.
The effect of validation-derived operating-threshold selection was region-dependent. Whole-area F1 increased from 0.021 to 0.109 in xp01, from 0.043 to 0.057 in xp04, and from 0.117 to 0.131 in xp05, whereas it decreased from 0.109 to 0.085 in xp02. The selected threshold for xp05 also showed substantial seed-to-seed variability (0.377 ± 0.323), indicating instability of the validation-derived operating point in this held-out setting. Thus, operating-threshold adjustment partially improved whole-area model–rule agreement in some subregions but did not restore performance to the levels observed in the controlled patch experiments.
This interpretation is further supported by the false-positive area relative to the rule-derived reference masks. Across the four held-out subregions, the macro-average false-positive area was approximately 0.498 km2 at t = 0.5 and 0.512 km2 under the validation-derived operating thresholds. Operating-threshold adjustment therefore changed the precision–recall trade-off rather than uniformly reducing disagreement with the reference masks. Together with the substantial regional variation, these results indicate that the whole-area performance gap cannot be attributed to decision-threshold mismatch alone.
Threshold-independent precision–recall analysis provided complementary information about probability-ranking behavior under natural class imbalance. The whole-area precision–recall curves for the four held-out subregions are shown in
Figure 7. Mean AUPRC values were 0.078, 0.060, 0.034, and 0.114 for xp01, xp02, xp04, and xp05, respectively. Because the corresponding reference-positive prevalences were only 1.85%, 1.20%, 1.02%, and 1.56%, the regional AUPRC values were approximately 4.22, 4.96, 3.37, and 7.34 times their respective prevalence baselines. The macro-average AUPRC across the four held-out subregions was 0.072, while the mean of the four region-specific AUPRC-to-prevalence ratios was approximately 4.97. These results indicate that the model retained some ability to rank rule-derived reference-positive pixels above background under natural class imbalance, although this ranking ability did not translate into high threshold-dependent whole-area F1.
Taken together, the whole-area evaluation reveals a substantial gap between controlled patch-based evaluation and complete-region performance. Validation-derived thresholds partially reduced this gap in three of the four subregions, but the remaining degradation indicates that threshold mismatch alone does not explain the observed performance loss. The residual gap is consistent with additional effects associated with the shift to natural candidate prevalence, heterogeneous background terrain, and cross-subregion variation in model-score distributions. Because the reference masks remain rule-derived, these results characterize the regional transfer behavior of the learned screening surrogate rather than independently verified landslide-detection accuracy.
5. Discussion
5.1. Geomorphological Meaning and Parameter Sensitivity of the Rule-Derived Screening Criterion
The boot-shaped screening criterion translates the qualitative steep-upper–gentle-lower terrain morphology reported for the Zhenxiong landslide [
4] into an explicit slope-unit screening rule. By using representative longitudinal profiles and predefined inclination thresholds, the procedure provides a reproducible way to identify terrain units conforming to this morphology. However, the resulting candidates represent conformity with an operational geomorphological criterion rather than confirmed landslides or independently validated unstable slopes. Accordingly, the rule-derived masks should be interpreted as structured geomorphological supervision for the subsequent segmentation model rather than as independent landslide ground truth.
The sensitivity analysis shows that this candidate definition is materially dependent on its parameterization. Across the tested configurations, candidate abundance, spatial overlap, and downstream model–rule agreement varied substantially, indicating that the screening rule is not parameter-invariant. The reference configuration (45°, 30°, 0.6) was retained because it represents the predefined expert-informed operational setting used in the main experiments, not because the sensitivity analysis demonstrated that it was physically optimal. In particular, a higher model F1 under an alternative threshold combination does not establish a more valid geomorphological criterion because each parameter combination generates a different reference target. Independent field or engineering-geological evidence would be required to determine physically preferable screening thresholds.
5.2. Configuration-Dependent Contribution of Terrain Variables
The input-feature experiments show that local terrain gradient information is particularly informative for approximating the rule-derived boot-shaped terrain pattern. Slope alone produced substantially higher model–rule agreement than DEM alone, while combining DEM and slope yielded competitive performance with relatively low variability. Terrain aspect, however, did not show a consistent incremental effect. Adding the complete sine–cosine aspect representation reduced mean F1 when combined with DEM or DEM + slope, but slightly improved performance when added to slope alone; indeed, Slope + Aspect produced the numerically highest mean F1 among the evaluated input configurations. These results therefore do not support treating aspects as uniformly beneficial or detrimental for this task.
DEM + Slope was retained as the final two-channel input because it provided a parsimonious representation of elevation context and local terrain gradient while maintaining performance close to the numerically best configuration. The additional contribution of aspect appeared to depend on the accompanying terrain variables, suggesting that conventional terrain factors used in landslide susceptibility analysis [
8,
9] should not be assumed to contribute in the same way when the learning target is a specific rule-derived geomorphological pattern. Accordingly, the selected DEM + Slope configuration represents a balance between model–rule agreement, input dimensionality, and fold-wise consistency rather than a statistically or universally optimal feature combination.
5.3. Effects of Segmentation Architecture and Optimization Objective
The architecture comparison showed a clear overall effect on model–rule agreement, with DeepLabV3+ achieving the highest mean F1 among the four evaluated models and maintaining the same favorable ranking across the three spatial folds. Its multi-scale contextual aggregation through ASPP is conceptually compatible with terrain patterns expressed over different spatial extents, but the present experiment compares complete architectures rather than individual components. Therefore, the observed advantage of DeepLabV3+ cannot be attributed specifically to ASPP, the ResNet-34 encoder, or any other single architectural element without additional component-level ablation. Likewise, the lower performance of SegFormer-B0 should not be interpreted as evidence that Transformer-based segmentation is inherently unsuitable for this task, because only one lightweight Transformer configuration was evaluated under the present dataset and training protocol.
The loss-function comparison showed weaker evidence of systematic differences. Dice + Focal produced the highest descriptive mean pixel-level F1, , and mIoU, whereas Weighted Cross-Entropy achieved the highest recall and Dice + Cross-Entropy the highest Patch-F1. The omnibus loss-function test was not statistically significant, indicating that no single objective can be regarded as uniformly superior under the present experimental design. Dice + Focal was therefore retained for the final cross-subregion and whole-area evaluations because of its overall descriptive pixel-level performance, rather than because of demonstrated statistical superiority. This result also suggests that the preferred optimization objective depends partly on the evaluation scale and the balance between pixel-level overlap and patch-level candidate presence.
5.4. Cross-Subregion Transfer and Whole-Area Deployment Gap
The LOSO experiments indicate that the learned segmentation surrogate retained measurable model–rule agreement when transferred among the four LiDAR-covered subregions within Zhenxiong County. The macro-averaged pixel-level F1 decreased from 0.351 under spatially grouped cross-validation to 0.321 under LOSO, while the precision–recall balance varied substantially among held-out subregions. These results indicate measurable cross-subregion transfer of the learned rule approximation within the investigated county, although the evidence remains limited to the four LiDAR-covered subregions of Zhenxiong County. Because all four regions share the same broader geographic setting and data-processing framework, the present results should not be interpreted as demonstrating generalization to independent counties or geologically distinct areas.
A substantially larger performance gap emerged when the LOSO models were applied to complete held-out regions under the natural prevalence of rule-derived candidates. At the fixed t = 0.5 operating threshold, the macro-average whole-area F1 was only 0.072, compared with the substantially higher values obtained on the intentionally sampled patch datasets. Validation-derived thresholds increased the macro-average F1 to 0.095, but the improvement was region-dependent and did not restore patch-level performance; performance even decreased in xp02, while the selected threshold for xp05 was highly variable across seeds. This indicates that operating-threshold mismatch explains only part of the whole-area deployment gap. Threshold adjustment alone therefore cannot reconcile the controlled patch results with complete-region deployment. The remaining degradation is consistent with additional effects associated with the shift from an approximately balanced sampled-patch distribution to complete terrain containing only about 1–2% reference-positive pixels, heterogeneous background conditions, and regional variation in model-score distributions.
Despite the low threshold-dependent whole-area F1, AUPRC remained above the corresponding prevalence baseline in all four held-out regions, indicating that the model retained some ability to rank rule-derived reference-positive pixels above background. This distinction is important for operational interpretation: the surrogate contains useful spatial ranking information, but the current model and sampling strategy do not yet support strong binary whole-area segmentation at a single transferable operating threshold. Future deployment should therefore emphasize evaluation under natural prevalence and independent geographic validation rather than relying primarily on performance obtained from controlled patch samples.
5.5. Practical Implications, Limitations, and Future Directions
The practical value of the proposed segmentation surrogate lies primarily in simplifying the operational screening workflow after model training. The original Stage-1 procedure requires hydrology-based slope-unit delineation, limited boundary refinement, representative-profile construction, segment-inclination calculation, and rule-based screening before candidate masks can be produced. In contrast, the trained model operates directly on co-registered DEM and slope-gradient rasters without reconstructing slope units or longitudinal profiles during inference. In the engineering benchmark, whole-area inference over approximately 147.23 km2 of valid LiDAR coverage required 40.602 s on an NVIDIA Tesla T4 GPU, with a peak resident memory of approximately 2.17 GiB. These measurements characterize the computational cost of the neural inference stage, but a reproducible end-to-end speedup relative to the historical rule-based workflow cannot be claimed because complete runtime, memory, and manual-refinement records for that workflow were not retained.
Several limitations remain. First, the segmentation targets are derived from an expert-informed geomorphological rule rather than from independently verified landslide ground truth, and the sensitivity analysis shows that the resulting candidate definition depends materially on the screening parameters. Second, all four study subregions are located within Zhenxiong County, so the LOSO experiment evaluates within-county transfer rather than external geographic generalization. Third, the small number of independent spatial folds limits statistical resolution, while the whole-area results reveal a substantial gap between controlled patch evaluation and natural-prevalence deployment. In addition, the same validation subsets were used for checkpoint selection and operating-threshold selection, and limited remote-sensing imagery was used for upstream slope-unit boundary quality control, although the final segmentation model itself requires only DEM and slope-gradient inputs. Future work should therefore prioritize independent geographic validation, independently interpreted or field-supported reference data, sampling strategies closer to natural prevalence, and external validation or calibration of operating thresholds.