Next Article in Journal
High-Resolution Typhoon Risk Assessment Based on Geospatial Big Data: A Case Study of Haikou, China
Previous Article in Journal
Surface Deformation Monitoring and Subsidence Risk Zonation Along the Middle Route of the South-to-North Water Diversion Project Coupling Time-Series InSAR with AHP-FCE
Previous Article in Special Issue
Adaptive Spatial–Frequency Information Fusion for SAR Ship Detection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Multimodal Remote Sensing Framework Based on an Improved YOLO Instance Segmentation Model for Automatic Glacial Lake Extraction in Southeastern Tibet

1
College of Water Conservancy and Civil Engineering, Xizang Agriculture and Animal Husbandry University, Linzhi 860000, China
2
Yarlung Zangbo Grand Canyon Water Cycle Monitoring and Research Station, Xizang Autonomous Region, Linzhi 860000, China
3
College of Water Sciences, Beijing Normal University, Beijing 100875, China
4
State Key Laboratory of Hydroscience and Engineering, Tsinghua University, Beijing 100084, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(16), 2769; https://doi.org/10.3390/rs18162769
Submission received: 1 July 2026 / Revised: 12 August 2026 / Accepted: 13 August 2026 / Published: 16 August 2026

Highlights

What are the main findings?
  • An improved YOLO11-seg model integrating SPDConv, MSCA, C3k2_Mamba, C2PSA_MSCA, and CARAFE achieved an F1-Score of 0.9205 and an mAP50(M) of 0.9331.
  • Using remote sensing imagery acquired in 2024, the model extracted 4766 glacial lakes in southeastern Tibet, covering 428.13 km2.
What is the implication of the main finding?
  • Southeastern Tibet’s steep terrain and numerous small lakes make it important for GLOF research; the 2024 inventory provides a regional basis for hazard assessment.
  • The multimodal deep-learning framework for glacial lake extraction serves as a methodological reference for compiling glacial lake inventories in environmentally similar high-mountain regions.

Abstract

Glacial lakes are sensitive indicators of climate-driven cryospheric change, and their accurate mapping provides fundamental spatial information for water-resource assessment and glacial lake outburst flood (GLOF) hazard assessment. In southeastern Tibet, automatic extraction remains difficult because glacial lakes are small and easily confused with snow, mountain shadows, dark bedrock, riverine wetlands, and non-glacial water bodies. In this study, we integrate optical bands and water indices from Sentinel-2, topographic information derived from a digital elevation model, and radar backscatter from Sentinel-1 into a nine-channel multimodal dataset, and develop an improved YOLO11-seg model that combines spatial-to-depth downsampling, multi-scale attention, long-range context modeling, and content-aware upsampling to enhance small-lake detection, background suppression, and boundary delineation. Compared with U-Net, DeepLabV3+, YOLOv8-seg, YOLO11-seg, YOLO12-seg, and YOLO26-seg, the proposed model achieved the highest F1-Score of 0.9205 and an mAP50(M) of 0.9331, while its F1-Score and IoU on the independent test set reached 0.9209 and 0.8533, respectively. Using remote sensing imagery acquired in 2024, the model extracted 4766 glacial lakes in southeastern Tibet, covering 428.13 km2; 80.84% of these lakes were smaller than 0.10 km2. The results demonstrate an effective and reproducible framework for automatic glacial lake mapping in complex alpine environments.

1. Introduction

Glacial lakes are sensitive indicators of cryospheric change and important components of high-mountain hydrological systems. Their number, area, and spatial distribution record glacier retreat and meltwater redistribution, while glacial lake inventories provide essential information for water-resource assessment, lake-evolution analysis, and GLOF hazard assessment [1,2,3,4,5]. A recent global semi-automated inventory mapped 117,352 glacial lakes larger than 0.01 km2 and found that lakes between 0.01 and 0.10 km2 accounted for 77.24% of the total number [4]. Multi-temporal observations in the Everest region also revealed sustained increases and pronounced seasonal variability in supraglacial ponds [6]. In the China–Pakistan Economic Corridor, 10 m Sentinel-2 imagery identified substantially more small lakes than 30 m Landsat imagery [5]. These findings demonstrate the need for accurate, repeatable, and scalable methods that can resolve small lakes and support regular inventory updates.
Southeastern Tibet is an important but particularly challenging region for glacial lake mapping. The region contains extensively distributed maritime glaciers, deeply incised valleys, strong topographic relief, abundant seasonal snow, and numerous small glacial lakes [7,8,9,10,11,12]. The coexistence of active maritime glaciers, expanding lakes, and steep downstream valleys creates a complex GLOF hazard setting and increases the potential for cascading downstream impacts in southeastern Tibet [10,12]. Persistent monsoonal cloud cover, seasonal snow, mountain shadows, and dark bare rock severely limit the availability of consistently clear optical imagery [5,6,7,8,9,12]. Small and irregular lakes are also easily confused with snow margins, shadowed terrain, turbid water, riverine wetlands, and non-glacial water bodies [8,9,12,13,14,15]. The 2023 Tibetan Plateau inventory identified 29,089 glacial lakes larger than 0.005 km2, with the main concentrations occurring in southeastern mountain regions and small lakes dominating by number [11]. A recent annual inventory for southeastern Tibet further reported persistent lake expansion and continuing difficulties in detecting micro glacial lakes under complex backgrounds [12]. Reliable extraction in this region therefore requires methods that can compensate for limited optical information and remain robust across heterogeneous alpine environments.
Existing glacial lake mapping methods can be classified into visual interpretation, semi-automatic extraction, and automatic extraction. Visual interpretation relies on manual delineation using single-date or multi-temporal satellite imagery [1,5]. It allows interpreters to examine ambiguous lake boundaries directly but is labor-intensive, operator-dependent, and difficult to apply consistently over large regions. Semi-automatic methods combine spectral indices, threshold segmentation, object-based classification, decision trees, active contour models, terrain constraints, or historical lake boundaries with manual inspection and correction [4,5,6,16,17,18,19,20,21,22,23,24]. These methods reduce the workload of manual interpretation and have supported regional and global inventories. However, index- and threshold-based approaches remain sensitive to empirical threshold selection and image quality and may confuse water with mountain shadows, dark bare rock, snow, ice, or other low-reflectance surfaces. Their regional applicability also depends on locally designed rules and the consistency of manual quality control.
Automatic glacial lake extraction has increasingly shifted from conventional machine-learning classification to deep-learning-based semantic and instance segmentation. Semantic segmentation approaches based on U-Net, DeepLab, and related encoder–decoder architectures have been used for pixel-level delineation of glacial lake surfaces [25,26,27,28,29,30,31,32,33,34,35,36,37]. Instance-oriented methods based on Mask R-CNN, YOLO, and related architectures further support object-level lake delineation and the representation of small targets [38,39,40,41]. Optical–SAR fusion networks have also been developed to improve glacial lake extraction under cloud, shadow, and other complex alpine conditions [42,43,44,45,46,47]. In addition, recent studies have explored prior geographical knowledge, weak supervision, cloud-based processing, transfer learning, and automated inventory construction [11,12,36,37,48,49,50,51,52,53]. Nevertheless, omission of small lakes, false detections caused by snow, shadows, wetlands, and dark bedrock, and incomplete boundaries of irregular lakes remain important challenges in complex alpine environments [4,8,9,12,35,36,37,39,41,42,45].
Against this background, this study addresses three scientific questions. First, can the joint use of optical, water-index, topographic, and SAR inputs improve lake discrimination when clear optical observations are limited? Second, can an instance-segmentation framework designed for small objects and complex boundaries reduce omissions, false detections, and boundary fragmentation? Third, can the resulting model support consistent regional glacial lake mapping across southeastern Tibet? To answer these questions, we construct a multimodal remote sensing dataset, evaluate an enhanced instance-segmentation framework against mainstream models, and conduct module and input-channel ablation experiments. The selected model is then applied to remote sensing imagery acquired in 2024 to compile a regional glacial lake inventory. The objectives are to improve automatic glacial lake extraction in complex alpine environments and provide updated spatial information for subsequent regional GLOF hazard assessment.

2. Materials and Methods

2.1. Study Area and Data Sources

The study area covers Nyingchi and Qamdo in southeastern Tibet and spans the eastern Nyainqentanglha Mountains, the eastern Himalayas, and the western Hengduan Mountains. The region is characterized by pronounced topographic relief, deeply incised river valleys, and extensively glacierized high mountains, with an average elevation of approximately 4000 m [8,9,12]. Located at the intersection of the westerlies and the South Asian monsoon system, the region receives warm and humid airflow transported through the Yarlung Tsangpo valley. Precipitation is strongly concentrated in summer and varies markedly from the humid south to the relatively dry north, creating diverse glacier, snow, vegetation, wetland, and river environments over short horizontal distances [8,9,12,45].
This distinctive alpine-humid environment supports one of the principal concentrations of maritime glaciers on the Tibetan Plateau. These glaciers are highly sensitive to climatic warming and have experienced substantial mass loss and terminus retreat, promoting the formation and expansion of numerous glacial lakes [8,10,12,45]. Many glacial lakes are distributed within steep tributary valleys connected to the Yarlung Tsangpo and other downstream river systems. Consequently, a GLOF may propagate through confined valleys and interact with debris-flow and river processes, increasing the potential for cascading downstream impacts on settlements, transportation corridors, hydropower facilities, and other infrastructure [9,10,12]. An accurate and up-to-date glacial lake inventory is therefore particularly important as a spatial basis for identifying lake locations, sizes, elevation distributions, and glacier–lake relationships required for subsequent regional GLOF hazard assessment.
The same environmental characteristics make reliable glacial lake mapping particularly difficult. Persistent monsoonal cloud cover restricts the availability of consistently clear optical observations, while seasonal snow, terrain shadows, supraglacial debris, turbid water, wetlands, river segments, dark bedrock, and seasonal meltwater ponds can produce spectral or backscatter responses similar to those of glacial lakes [8,9,12,45]. Small and irregular glacial lakes are especially susceptible to omission because of mixed-pixel effects and weak water–land boundaries. Moreover, the steep terrain that generates extensive optical shadows can also cause layover and radar-shadow distortions in SAR imagery [45]. These combined difficulties have contributed to incomplete or temporally inconsistent glacial lake inventories in southeastern Tibet [9,12]. The region therefore provides both an important target for updated glacial lake inventorying and a stringent test area for evaluating whether optical, water-index, topographic, and SAR information can jointly improve extraction under complex high-mountain conditions.
The regional glacial lake inventory was produced for southeastern Tibet, whereas the samples used for model development and evaluation were collected from ten geographically separated regions across Tibet. The sample regions were purposively selected to represent variations in lake size, morphology, elevation, glacier-contact condition, surface state, and surrounding background. Collectively, they include very small and clustered lakes, compact and elongated lakes, irregular or fragmented boundaries, supraglacial and glacier-contact settings, open-water and partially frozen surfaces, and backgrounds affected by snow, mountain shadows, dark bare rock, debris-covered glaciers, rivers, or wetlands. The ten regions were subsequently assigned at the region level to the training, validation, and test subsets, as shown in Figure 1.
The data sources used in this study include Sentinel-2 MSI, Sentinel-1 SAR, DEM data, glacier inventory data, and semi-automatically generated glacial lake labels. Sentinel-2 MSI was used to construct visible, near-infrared, shortwave-infrared, and water-index features; Sentinel-1 SAR was used to provide radar backscatter information; and Copernicus DEM was used to derive slope information. Glacier boundaries were obtained from the Second Chinese Glacier Inventory and RGI 7.0 [54,55]. All raster datasets were resampled, co-registered, and spatially aligned to a 10 m grid. The main data sources and their roles are summarized in Table 1.
Sentinel-2 and Sentinel-1 images used for each sample region were retrieved using the same region-specific temporal window listed in Table 2. Within each temporal window, Sentinel-2 images satisfying the corresponding cloud-cover thresholds were selected. For each selected image, the QA60 quality band was used to identify pixels flagged as opaque cloud or cirrus, and these pixels were masked and excluded from subsequent compositing. For each spectral band and spatial pixel location, the median reflectance was then calculated from all valid, unmasked observations available across the selected images, producing a single Sentinel-2 optical composite for the corresponding region. Sentinel-1 images acquired during the same period were filtered to retain IW-mode observations with VV polarization. A pixel-wise median VV-backscatter composite was subsequently calculated from the selected Sentinel-1 images. The resulting Sentinel-2 optical composite and Sentinel-1 SAR composite were co-registered and incorporated into the corresponding nine-channel input patches.
The temporal windows were adjusted among regions because the availability of usable Sentinel-2 observations varied spatially owing to monsoonal cloud cover, seasonal snow, and terrain shadows. This region-specific selection provided sufficiently clear optical observations while maintaining a common observation period for the Sentinel-2 and Sentinel-1 inputs within each region. After the Sentinel-2 composite was generated, NDWI and MNDWI were calculated using the same band definitions and preprocessing procedure for all sample regions. The cloud-cover thresholds, shared Sentinel-1/Sentinel-2 temporal windows, patch counts, and dataset split assignments are summarized in Table 2.
To reduce spatial information leakage among the dataset subsets, the dataset was partitioned at the sample-region level rather than by randomly assigning individual image patches. Regions 1–6 were assigned to the training subset, Regions 7–8 to the validation subset, and Regions 9–10 to the test subset, as shown in Figure 1 and Table 2. The resulting training, validation, and test subsets contained 15,400 (79.5%), 1390 (7.2%), and 2583 (13.3%) image patches, respectively. All patches originating from the same sample region were assigned to the same subset; consequently, no sample region was shared among the three subsets. Spatially overlapping patches and directly neighboring patches generated within a given sample region were retained in the same subset and were not distributed across different subsets. This region-level partitioning was used to reduce spatial dependence between the training data and the validation and test data and to provide a more rigorous assessment of model generalization to geographically independent regions.
Based on the study-area characteristics, data sources, and dataset partitioning described above, the overall workflow comprises multimodal data preprocessing and channel construction, sample construction and annotation, enhanced model development, regional glacial lake extraction and post-processing, and inventory validation and spatial analysis (Figure 2).

2.2. Multimodal Sample Construction and Annotation

The model input constructed in this study consists of nine channels, including RGB, NIR, SWIR, NDWI, MNDWI, slope, and SAR:
X = [ R , G , B , N I R , S W I R , N D W I , M N D W I , S l o p e , S A R ] ,
where X denotes the multimodal remote sensing feature tensor input to the model; RGB represents the visible bands of Sentinel-2; NIR and SWIR denote the near-infrared and shortwave-infrared bands, respectively; NDWI and MNDWI are water indices; slope is derived from the DEM; and SAR represents the Sentinel-1 VV-polarized backscatter feature.
Here, R, G, B, NIR, and SWIR correspond to Sentinel-2 bands B4, B3, B2, B8, and B11, respectively. NDWI and MNDWI were calculated using Sentinel-2 band numbers as follows:
N D W I = B 3 B 8 B 3 + B 8
MNDWI = B 3 B 11 B 3 + B 11
where B3 is the green band, B8 is the near-infrared band, and B11 is the shortwave-infrared band of Sentinel-2. NDWI enhances water responses by using the reflectance difference between green and near-infrared bands, whereas MNDWI introduces the shortwave-infrared band to further improve the separability between water bodies and bare land or dark background features.
Each input channel was selected according to its physical relevance to glacial lake extraction. RGB bands provide visible morphology, texture, and color information for identifying open-water surfaces and surrounding land-cover patterns. NIR and SWIR bands provide additional spectral contrast between open water and surrounding alpine backgrounds. Open water generally exhibits strong absorption and low reflectance in the near-infrared and shortwave-infrared wavelengths, whereas vegetation, bare rock, snow, ice, and moist surfaces exhibit different spectral responses across the visible, NIR, and SWIR bands. NDWI and MNDWI further enhance water-related responses and improve the separability between glacial lakes and non-water backgrounds, particularly bare land and dark background features. The slope layer provides topographic constraints that help reduce terrain-related false detections in steep mountain areas, while Sentinel-1 SAR backscatter provides illumination-independent information that may assist lake identification in shadowed or partially cloud-affected scenes.
To reduce numerical-scale differences among the heterogeneous input variables, each channel was independently clipped to its predefined physically valid range and linearly rescaled to 1–255 before channel concatenation, as expressed in Equation (4). Invalid pixels, including NoData, NaN, Inf, and other abnormal values, were assigned a value of 0. The nine rescaled channels were then stacked into a 512 × 512 × 9 tensor. During model preprocessing, this tensor was converted to floating-point format and divided by 255, yielding a common numerical range of approximately 0–1 for all valid input channels. Therefore, reflectance, water-index, slope, and SAR backscatter values were not directly concatenated in their original physical units or numerical ranges. This channel-wise preprocessing reduces the risk of numerical dominance caused solely by differences in the original value scales, although it does not make the statistical distributions or physical meanings of the different modalities identical.
x n o r m = c l i p ( x , x m i n , x m a x ) x m i n x m a x x m i n × 254 + 1 , x Ω v , 0 , x Ω i .
where x denotes the original pixel value of a single channel; x m i n and x m a x represent the lower and upper bounds of the predefined valid physical range for that channel, respectively; c l i p ( · ) restricts pixel values to the valid range; Ω v denotes the set of valid pixels; and Ω i denotes the set of invalid pixels, including NoData, NaN, positive or negative infinity, and other abnormal values.
Glacial lake labels were produced through a semi-automatic annotation workflow. Candidate glacial lakes in the sample patches were first segmented using the Segment Anything Model (SAM) [56] within X-AnyLabeling version 3.3.10 [57] and were then corrected through manual visual interpretation to obtain the final glacial lake instance labels. This workflow balances annotation efficiency and boundary accuracy and is applicable to open-water, frozen or partially frozen, small, and complex-background glacial lake samples. To illustrate the diversity of the annotated samples, Figure 3 presents representative supraglacial lakes, glacier-contact lakes, and non-glacier-contact lakes according to their spatial relationships with glaciers. This grouping is used for representative visualization, whereas all annotated instances are treated as a single glacial lake class during model training and evaluation.

2.3. Enhanced YOLO11 Glacial Lake Segmentation Model

The proposed model is based on YOLO11-seg, with 512 × 512 × 9 multimodal image patches as input and glacial lake instance masks as output. Compared with previous YOLO-series glacial lake segmentation studies [39], this study extends the model input to 9-channel multimodal data and improves feature extraction, attention enhancement, contextual modeling, and boundary recovery.
The enhanced architecture incorporates four targeted modifications for glacial lake segmentation in complex alpine environments. SPDConv and a P2 high-resolution prediction layer were used to preserve fine-scale information for small glacial lakes [39,58]. MSCA and C2PSA_MSCA were incorporated to enhance multi-scale glacial lake features and suppress heterogeneous background responses [59,60]. C3k2_Mamba was introduced to capture longer-range spatial dependencies [61,62], whereas CARAFE was used to improve glacial lake boundary-feature reconstruction during upsampling [63]. Their positions within the network are shown in Figure 4.
At the input stage, the independently rescaled channels were concatenated into a single tensor and jointly processed by the network. This strategy is therefore referred to as input-level early fusion rather than modality-specific feature fusion. Optical bands, water indices, slope, and SAR backscatter were not processed through separate modality-specific branches. Instead, the first convolutional layer applied learnable kernels across all nine channels, transforming the normalized multimodal inputs into joint feature representations. These representations were subsequently refined through the shared backbone. In the neck, deep semantic features were upsampled and concatenated with shallower spatial features, followed by further processing through the attention-enhanced modules. This architecture enables spectral, index-based, topographic, and radar information to be jointly learned for small-object representation, background discrimination, and glacial lake mask prediction.
(1)
SPDConv Module
SPDConv was introduced to reduce the loss of fine-scale glacial lake features during early downsampling [58]. Instead of discarding spatial samples through strided convolution or pooling, it rearranges four interleaved sub-feature maps along the channel dimension and subsequently applies Conv-BN-SiLU. This operation preserves spatial information that is important for very small glacial lakes:
X spd   =   Concat c ( X 00 , X 10 , X 01 , X 11 )
Y spd = SiLU ( BN ( Conv k × k ( X spd ) ) )
where X 00 ,   X 10 ,   X 01 , and X 11 denote the four sub-feature maps obtained from the input feature X by interval sampling along odd and even rows and columns, and Concat c ( · ) denotes concatenation along the channel dimension. X spd transfers local spatial information into the channel dimension, and Y spd is the output feature after convolution, batch normalization, and SiLU activation. The structure of the SPDConv module is shown in Figure 5.
(2)
MSCA and C2PSA_MSCA Modules
MSCA was introduced to enhance glacial lake-related responses and suppress heterogeneous background interference through multi-scale spatial attention [59]. It combines depthwise and strip convolutions to aggregate local and directional contextual information at multiple scales [59,60]. In the proposed network, MSCA was embedded within C2PSA to form C2PSA_MSCA, allowing high-level features to be adaptively reweighted before branch fusion and supporting the discrimination of glacial lakes from terrain shadows, dark bedrock, and wetlands:
A m = σ C o n v 1 × 1 A 0 + A 7 + A 11 + A 21
F m = F A m
A s = σ C o n v 7 × 7 C o n c a t c A v g P o o l c F m , M a x P o o l c F m
F o u t = F m A s
A , B = S p l i t C o n v 1 × 1 ( 1 ) X
Y c 2 p s a = C o n v 1 × 1 ( 2 ) C o n c a t c A , P n P 2 P 1 B
where F is the input feature; A 0 is the base feature obtained by the 5 × 5 depthwise convolution; A 7 , A 11 , and A 21 are the outputs of strip convolution branches at different scales; A m denotes the multi-scale convolutional attention weight; F m denotes the multi-scale attention-enhanced feature; A s denotes the spatial attention weight; and F o u t is the final output of MSCA. In C2PSA_MSCA, A and B are the two branches obtained after the first 1 × 1 convolution, P i ( ) denotes the i-th PSABlock_MSCA, and the second 1 × 1 convolution is used to fuse the two branches. The structures of the MSCA and C2PSA_MSCA modules are shown in Figure 6.
(3)
C3k2_Mamba Module
C3k2_Mamba embeds a MambaBlock within the C3k2 feature-aggregation structure at the P4/16 level to capture longer-range spatial dependencies across the feature map [61,62]. This design helps associate spatially separated features belonging to fragmented, elongated, or partially obscured glacial lakes that may not be fully connected through local convolution alone. The feature map is reshaped into a sequence, processed through LayerNorm and the Mamba operator, restored to its spatial arrangement, and combined through a residual connection:
Q = R e s h a p e ( X )
Q m = M a m b a ( L N ( Q ) )
Y m a m b a = X + R e s h a p e 1 ( Q m )
where X is the input two-dimensional feature map; R e s h a p e ( · ) denotes the operation that flattens the feature map into a sequence; Q is the serialized feature; L N ( · ) denotes LayerNorm; Q m is the sequence feature modeled by the Mamba operator; R e s h a p e 1 ( · ) denotes the operation that restores the sequence to a spatial feature map; and Y m a m b a is the MambaBlock output with a residual connection. The structure of the C3k2_Mamba module is shown in Figure 7.
(4)
CARAFE Upsampling Module
CARAFE was used to replace fixed interpolation-based upsampling and improve the reconstruction of elongated or irregular glacial lake boundaries [63]. It predicts content-aware reassembly kernels from local features and adaptively aggregates neighboring information during upsampling, supporting the recovery of boundary features at weak water–land transitions and along irregular glacial lake shorelines:
Y p = n N ( p , k u p ) W p , n X n , p = p s
W p = S o f t m a x ψ ( X ) p , n N ( p , k u p ) W p , n = 1
where p denotes a spatial position in the output feature map, p denotes the corresponding input feature position, and s denotes the upsampling factor, with p = p s . N p , k u p denotes a local neighborhood of size k u p × k u p centered at p , X n denotes the input feature at neighboring position n, and W p , n denotes the content-aware reassembly weight. ψ ( · ) denotes the prediction function used to generate the reassembly kernel from local content, S o f t m a x ( · ) is used to normalize neighboring weights, and Y p denotes the reassembled output feature. The structure of the CARAFE module is shown in Figure 8.

2.4. Post-Processing of Prediction Results

After model inference, the patch-level mask predictions were mosaicked and converted into glacial lake vector polygons. Regularized post-processing was subsequently applied to reduce isolated noise and topographically implausible detections. NoData and other invalid pixels were first corrected during inference. A 10 km buffer was then constructed from the merged glacier boundaries of the Second Chinese Glacier Inventory and RGI 7.0 [54,55], and candidate polygons intersecting this buffer were retained. The 10 km distance was selected because it has been widely used as a practical upper limit for defining the potential distribution of glacial lakes in large-scale inventories of High-Mountain Asia and southeastern Tibet [1,42]. Predicted pixels located on terrain with slopes greater than 30° were subsequently excluded. This conservative criterion follows previous remote sensing mapping of supraglacial ponds, in which a 30° surface-slope threshold was used to eliminate steep avalanche fans and icefalls [6]. In the present regional workflow, it was applied as a topographic plausibility filter to reduce detections on steep shadowed rock faces while limiting excessive removal caused by DEM resolution and terrain uncertainty. Finally, isolated polygons smaller than 1000 m2 (0.001 km2) were removed. At the 10 m mapping resolution, this minimum mapping unit corresponds to the nominal area of ten pixels and was used to suppress isolated segmentation noise while retaining the very small glacial lakes that constitute an important component of the regional inventory.
To reduce the staircase-like artifacts introduced during raster mask-to-vector conversion, the polygon boundaries were processed sequentially using long-edge densification with a maximum segment length of 25 m, one iteration of Chaikin corner cutting, cubic B-spline interpolation with 200 target points and a smoothing factor of 0.01, paired +9 m and −9 m buffering with a resolution of 10, topology-preserving simplification with a tolerance of 1 m, and final geometry repair. Equal outward and inward buffer distances were used to limit systematic expansion or contraction of the polygons. To quantify the possible influence of this procedure on lake-area estimates, the polygon areas before and after smoothing were compared at both the regional and individual-object levels. Spatially corresponding polygons were identified according to their maximum spatial overlap, and the aggregate area change and distribution of absolute object-level area changes were calculated.

2.5. Experimental Settings and Evaluation Metrics

All experiments were conducted on a workstation equipped with an Intel Core i9-14900K CPU (Intel Corporation, Santa Clara, CA, USA) and an NVIDIA GeForce RTX 4090 GPU with 24 GB of memory (NVIDIA Corporation, Santa Clara, CA, USA). The models were implemented in Python 3.10.18 using PyTorch 2.4.1 with CUDA 12.4. All models were initialized without pretrained weights and trained from scratch for 500 epochs. The input image size was set to 512 × 512 pixels, with a batch size of 32. MuSGD was employed as the optimizer, with an initial learning rate of 0.01, a momentum coefficient of 0.937, and a weight decay of 5 × 10−4. The first 15 epochs were used as a warm-up stage, during which the optimizer momentum increased from 0.50 to 0.937. After the warm-up stage, the learning rate was adjusted using a cosine schedule and gradually reduced to 10% of its initial value. Automatic mixed-precision training was enabled, and the random seed was fixed at 42.
To improve model generalization, the training data were augmented through random horizontal flipping with a probability of 0.50, rotations within ±2°, translations of up to 8%, scaling of up to 5%, Mosaic augmentation with a probability of 0.20, and copy–paste augmentation with a probability of 0.05. Mosaic augmentation was disabled during the final 30 epochs to promote stable convergence toward the end of training.
This study used Precision, Recall, F1-Score, IoU, and mAP50(M) to evaluate model performance. Precision reflects the ability to control false positives, Recall reflects the ability to reduce omissions, and F1-Score provides a combined assessment of Precision and Recall. IoU measures the spatial overlap between predicted masks and ground-truth masks, whereas mAP50(M) represents the mean average precision of masks at an IoU threshold of 0.50.
Precision = TP TP + FP
R e c a l l = T P T P + F N
F 1 s c o r e = 2 × P r e c i s i o n × R e c a l l P r e c i s i o n + R e c a l l
I o U = T P T P + F P + F N
m A P 50 ( M ) = 1 C c = 1 C A P c M ( I o U = 0.50 )
where TP denotes glacial-lake pixels correctly classified as foreground, FP denotes background pixels incorrectly classified as glacial-lake foreground, and FN denotes glacial-lake pixels missed by the model. Here, mAP50(M) is the instance-mask mean average precision, where 50 indicates an IoU threshold of 0.50 and M denotes that the evaluation target is the segmentation mask; this metric was reported only for the YOLO-series instance segmentation models.

2.6. Channel-Attribution Analysis

To quantify the relative contribution of each input channel within the trained 9-channel model, a channel-wise SHAP (SHapley Additive exPlanations) analysis was conducted using 1000 glacial-lake-containing patches randomly selected from the validation set with a fixed random seed of 42. For each channel, the reference value was defined as the channel-specific median estimated from 500 training patches. Channel attribution was calculated with respect to the confidence-weighted post-processed mask score. Global channel importance was quantified as the mean absolute SHAP value across the 1000 samples, and the relative contribution of each channel was obtained by normalizing its mean absolute SHAP value against the total attribution of all nine channels. For the grouped analysis, the mean absolute SHAP values of the constituent channels were summed within five predefined input groups: RGB, NIR–SWIR, water indices, slope, and SAR. This analysis quantifies the relative reliance of the trained model on individual channels and multimodal input groups.

2.7. Regional Product Validation

The regional mapping product was evaluated using two independently produced glacial lake inventories. The first was the Landsat-derived 2018 High-Mountain Asia inventory published by Wang et al. (2020) [1], which was generated from 30 m Landsat imagery through manual vectorization and interactive quality control. The second was the 2023 Tibetan Plateau glacial lake inventory of Tian et al. (2026) [11], which was generated from 10 m Sentinel-2 imagery using Attention DeepLabV3+ and subsequently refined through visual inspection and boundary editing with high-resolution Google Earth imagery. Both inventories were clipped to the study area and the same 10 km glacier-buffer extent. They were evaluated separately rather than merged because they differ in observation period, spatial resolution, glacial lake definition, and inventory-generation procedure.
All three inventories were projected to EPSG:6933 before spatial comparison. Glacier distances were calculated separately using a local azimuthal equidistant projection centered on the study area. For each reference lake, all mapped polygons with a positive-area intersection were identified, and the intersecting polygon with the highest Jaccard index was retained as its best-matching object. The Jaccard index was calculated as J(A,B) = |A∩B|/|A∪B|, where A and B denote the reference and mapped lake polygons, respectively. A reference lake was classified as matched when it had at least one positive-area intersection with the present inventory. No additional common area threshold was imposed for the validation; all polygons contained in each clipped reference Shapefile were retained. The evaluation included the reference-object match rate, matched reference-area proportion, and the distribution of maximum Jaccard indices. Match rates were further stratified by reference-lake area, elevation, and distance to the nearest glacier. Glacier distance was calculated for both reference inventories using the same merged glacier boundaries from the Second Chinese Glacier Inventory and RGI 7.0 [54,55].

3. Results

3.1. Model Accuracy and Ablation Analysis

(1)
Quantitative Comparison with Mainstream Segmentation Models
To ensure a controlled comparison of model architectures, U-Net, DeepLabV3+, YOLOv8-seg, YOLO11-seg, YOLO12-seg, YOLO26-seg, and the proposed model were all trained from scratch using the same predefined training subset and identical training settings. The same dataset partitions, nine-channel input configuration, channel-wise preprocessing procedure, input size, optimization settings, and evaluation protocol were used for all models. The comparison in Table 3 evaluates differences among model architectures under a fixed multimodal input configuration.
The comparison included both semantic and instance segmentation models. U-Net and DeepLabV3+ were retained as representative semantic segmentation baselines because they provide established references for glacial lake foreground-mask delineation [25,26], whereas the YOLO-series models provide direct comparisons under the same instance segmentation task [39]. Although semantic segmentation models do not explicitly distinguish individual lake instances, both model categories generate spatial masks of glacial lake regions. Precision, Recall, and F1-Score were therefore used as common measures of glacial lake segmentation performance, while mAP50(M) was reported only for the YOLO-series instance segmentation models.
As shown in Table 3, the proposed model achieved the highest Precision, Recall, and F1-Score among all compared models, with values of 0.9714, 0.8746, and 0.9205, respectively. Among the YOLO-series instance segmentation models, the proposed model also achieved the highest mAP50(M), reaching 0.9331. Compared with YOLO11-seg under the same complete 9-channel input configuration, the proposed model improved the F1-Score by 0.0587 and mAP50(M) by 0.0433. These results support the effectiveness of the architectural improvements when the input data configuration is held constant.
(2)
Module Ablation Experiment
To evaluate the contribution of the improved model design, two complementary ablation strategies were adopted. In the progressive analysis, YOLO11-seg was used as the baseline, and SPDConv, MSCA, C3k2_Mamba, C2PSA_MSCA, and CARAFE were added successively to examine the performance changes during construction of the complete model. A leave-one-out analysis was then conducted by removing each module individually from the complete configuration to assess its contribution within the integrated model. The corresponding module-ablation configurations and results are summarized in Table 4.
The final model achieved the best overall results, with Precision, Recall, F1-Score, and mAP50(M) values of 0.9714, 0.8746, 0.9205, and 0.9331, respectively. Compared with the baseline, these four metrics increased by 0.0544, 0.0617, 0.0587, and 0.0433, respectively. The progressive analysis therefore shows the overall performance improvement obtained during construction of the complete configuration. In the leave-one-out analysis, every reduced configuration yielded lower F1-Score and mAP50(M) values than the complete model. Together, the two analyses indicate that the complete module combination achieved the best overall performance among the tested configurations and suggest that each module retained a measurable contribution within the final configuration.
(3)
Input-Channel Combination Ablation Experiment
To assess the contribution of multimodal and multi-channel inputs to glacial lake segmentation, five input-channel combinations were designed, ranging from RGB three-channel input to the complete 9-channel input. The corresponding input combinations and results are summarized in Table 5.
As summarized in Table 5, the complete 9-channel input setting (S5) achieved the highest Precision (0.9714), F1-Score (0.9205), and mAP50(M) (0.9331) among the five input configurations. Compared with the RGB-only setting (S1), S5 improved the F1-Score and mAP50(M) by 0.0454 and 0.0333, respectively.
To further examine the contribution of each input within the complete S5 configuration identified in Table 5, a channel-wise SHAP analysis was performed, as shown in Figure 9. All nine channels exhibited measurable mean absolute SHAP attribution. R and NDWI provided the largest individual contributions, accounting for 29.26% and 28.00% of the total attribution, respectively. At the input-group level, RGB, water indices, NIR–SWIR, slope, and SAR accounted for 41.47%, 31.14%, 14.72%, 6.99%, and 5.67%, respectively. The attribution magnitudes differed among the channels and input groups, with RGB and water indices accounting for the largest proportions and NIR–SWIR, slope, and SAR showing smaller but measurable attribution. Together with the channel-combination results in Table 5, these findings show that the complete model drew measurable information from all predefined input groups.
Complementing the quantitative ablation and attribution results presented in Table 5 and Figure 9, Figure 10 visualizes the spatial patterns represented by the nine input components in a representative glacial lake scene. RGB imagery preserves visible shoreline morphology, texture, and surrounding land-cover information. NIR, SWIR, NDWI, and MNDWI provide additional spectral and water-index contrasts between lake surfaces and alpine backgrounds. The slope layer represents local terrain conditions, whereas Sentinel-1 SAR provides radar-backscatter information that is independent of solar illumination. These channel-specific patterns provide distinct input representations for the complete multimodal configuration evaluated in Table 5.
The segmentation performance of the proposed model and the main baseline models was further compared on the test set, as shown in Table 6.
On the test set, the proposed model achieved the highest Recall, F1-Score, and IoU, with values of 0.9283, 0.9209, and 0.8533, respectively. These results indicate that the proposed model maintained strong glacial lake recognition completeness and spatial overlap accuracy on independent test samples.

3.2. Regional-Scale Glacial Lake Mapping in Southeastern Tibet

After model accuracy evaluation, the optimal model was applied to remote sensing imagery acquired in 2024 to extract and map glacial lakes across southeastern Tibet. The regional mapping used complete 9-channel multimodal remote sensing data as input. The post-processing workflow included patch-level prediction mosaicking, NoData correction, small-area noise removal, raster mask-to-vector polygon conversion, 10 km buffer filtering based on merged glacier inventory data, slope filtering, area calculation, elevation extraction, and boundary refinement to produce the final glacial lake vector dataset for southeastern Tibet.
The area comparison conducted before and after boundary smoothing showed that the aggregate mapped lake area changed by only 0.64%. At the individual-object level, the median absolute area change was 0.59%, with an interquartile range of 0.26–1.20%, and 95% of the spatially corresponding lake polygons showed an absolute area change of no more than 2.58%. These results indicate that the smoothing procedure primarily regularized pixel-induced boundary irregularities and had only a limited influence on the regional lake-area statistics. The final spatial distribution of the extracted glacial lakes, glacier boundaries, and the 10 km glacier buffer is shown in Figure 11.
(1)
Elevation-Zone Distribution
Based on the regional-scale glacial lake vector results and DEM data, the number and area of glacial lakes were summarized by elevation zone (Table 7).
In terms of elevation zones, the 4000–4500 m zone contained the largest number of glacial lakes, with 1549 lakes, accounting for 32.50% of the total number. This elevation zone also had the largest glacial lake area, with 152.1352 km2, accounting for 35.54% of the total area. The 4500–5000 m and 5000–5500 m elevation zones contained 1487 and 1180 glacial lakes, respectively, indicating that glacial lakes in the study area were mainly concentrated between 4000 and 5500 m.
(2)
Area-Class Distribution
To further characterize the size structure of glacial lakes, the number and area contribution of glacial lakes were summarized by area class (Table 8).
In terms of area classes, the 0.01–0.05 km2 class contained the largest number of glacial lakes, with 2199 lakes, accounting for 46.14% of the total number. Lakes smaller than 0.10 km2 totaled 3853, accounting for 80.84% of all glacial lakes. In contrast, the area contribution was dominated by larger lakes: the 0.10–0.50 km2 and >0.50 km2 classes together contributed 313.8717 km2, accounting for 73.31% of the total glacial lake area.

3.3. Regional Inventory Validation

Object-level comparison showed similar overall agreement with the two independent reference inventories. Of the 4063 lakes in the 2018 inventory of Wang et al. (2020) [1], 2623 spatially intersected the present inventory, corresponding to a reference-object match rate of 64.6%; these matched objects represented 88.5% of the total reference-lake area. Of the 4486 lakes in the 2023 inventory of Tian et al. (2026) [11], 2911 were matched, giving an object match rate of 64.9% and a matched reference-area proportion of 89.0%. The median maximum Jaccard indices were 0.746 (interquartile range, 0.652–0.821) and 0.717 (0.615–0.800), respectively. Among the matched objects, 92.7% in the Wang inventory and 91.1% in the Tian inventory achieved a maximum Jaccard index of at least 0.50.
The comparison revealed a strong dependence on lake size (Figure 12). For the Wang inventory, object match rates increased from 30.0% for lakes smaller than 0.01 km2 to 100.0% for lakes larger than 0.50 km2; the corresponding rates for the Tian inventory increased from 38.9% to 98.8%. Median maximum Jaccard indices also increased with lake size, from 0.546 and 0.540 in the smallest class to 0.910 and 0.900 in the largest class for the Wang and Tian inventories, respectively. Thus, most object-level discrepancies occurred among very small lakes, whereas medium and large lakes showed substantially higher detection and boundary agreement.
Agreement also varied with elevation and glacier proximity (Figure 12). Match rates were highest at 4000–4500 m for both the Wang (91.9%) and Tian (85.6%) inventories and generally decreased at higher elevations. Using the same merged glacier boundaries, the match rates for the glacier-contact, 0–1 km, 1–5 km, and 5–10 km classes were 53.7%, 56.3%, 71.9%, and 72.6%, respectively, for the Wang inventory, and 44.4%, 59.0%, 72.4%, and 75.0%, respectively, for the Tian inventory. These stratified results show that the lower overall object match rates were concentrated mainly among very small lakes and lakes close to glacier margins.

4. Discussion

4.1. Significance of Multimodal Inputs and Model Improvements

Glacial lakes in southeastern Tibet commonly occur against complex backgrounds comprising seasonal snow, mountain shadows, dark bare rock, riverine wetlands, and glacier-terminal landforms. These conditions make lake boundaries difficult to identify consistently from visible imagery alone. Previous studies have applied Sentinel-2-based YOLO segmentation to glacial lake mapping [39] and have demonstrated the value of combining optical and SAR observations in shadow-affected alpine terrain [42,44]. In contrast to optical-only approaches, the present study jointly incorporated visible and infrared bands, water indices, slope, and SAR backscatter through input-level early fusion. The early-fusion architecture does not contain an explicit arbitration rule for locally inconsistent signals among the input modalities. Instead, the relative influence of the optical, water-index, slope, and SAR channels is learned through cross-channel convolution and subsequent shared feature transformations under the segmentation objective. Accordingly, SAR speckle, optical edges, water-index responses, and terrain information are combined within learned joint representations rather than resolved according to a predefined modality priority. The slope channel used in the network should also be distinguished from the separate 30° slope criterion applied after inference; the latter acts as a deterministic topographic plausibility filter and does not participate in network-level fusion.
The complete 9-channel configuration increased the F1-Score and mAP50(M) by 0.0454 and 0.0333, respectively, compared with the RGB-only configuration. More specifically, adding SAR to the eight-channel S4 configuration increased Precision from 0.9040 to 0.9714, F1-Score from 0.8966 to 0.9205, and mAP50(M) from 0.9299 to 0.9331, although Recall decreased from 0.8893 to 0.8746. This result indicates that the addition of SAR was associated mainly with improved false-positive control rather than a uniform increase in all evaluation metrics. The SHAP analysis further identified measurable but unequal attribution from RGB, water indices, NIR–SWIR, slope, and SAR. Figure 13 provides corresponding qualitative examples in which the complete configuration recovered lake regions that were missed or confused by the RGB-only configuration in several cloud- and shadow-affected scenes. Taken together, these quantitative and qualitative results characterize the contribution of the complete input configuration without assuming that each channel performs a single predetermined function.
Under the same complete 9-channel input configuration, the proposed model improved the F1-Score and mAP50(M) by 0.0587 and 0.0433, respectively, compared with the baseline YOLO11-seg model. It also achieved the highest Precision, Recall, and F1-Score among the compared semantic and instance segmentation models. The architectural modifications address different aspects of the extraction problem: SPDConv and the P2 detection layer strengthen small-target representation; MSCA and C2PSA_MSCA enhance feature selection under heterogeneous backgrounds; C3k2_Mamba captures longer-range context for fragmented or elongated lake components; and CARAFE supports the reconstruction of irregular or weakly defined boundaries. The progressive ablation and representative-scene comparisons therefore support the combined effectiveness of these modifications in reducing small-lake omissions, suppressing complex-background interference, and improving boundary continuity.

4.2. Comparison with Existing Glacial Lake Inventories and Regional Implications

Using remote sensing imagery acquired in 2024, the regional mapping identified 4766 glacial lakes in southeastern Tibet, with a total area of 428.13 km2. After filtering using the 10 km glacier buffer, glacial lakes showed a distinctly uneven spatial distribution across the study area, with clustered patterns in the western part of the region, the central alpine-gorge belt, and the glacier-developed areas in the southeast. Further statistics by elevation zone and area class reveal two notable patterns. First, both the number and area of glacial lakes were mainly concentrated between 4000 and 5500 m, particularly in the 4000–4500 m elevation zone. Second, small glacial lakes dominated in terms of number, whereas the total regional lake area was mainly contributed by a relatively small number of larger lakes (Figure 14).
At the aggregate level, the three inventories showed broadly consistent size- and elevation-distribution patterns (Figure 15). The main differences occurred in the smallest area classes and in total mapped area. Compared with the 2018 Landsat inventory of Wang et al. (2020) [1], the present 10 m inventory contained more small objects and a larger total lake area. The 2023 Sentinel-2 inventory of Tian et al. (2026) [11] provides a closer comparison in spatial resolution and observation period, although differences remain because of the distinct lake definitions, extraction procedures, minimum mapping units adopted during inventory production, and boundary-refinement methods.
The object-level results provide a more direct assessment than aggregate counts alone (Figure 12). The similar reference-object match rates of 64.6% and 64.9%, together with matched reference-area proportions of 88.5% and 89.0%, indicate that the present inventory captured most of the lake area represented by both external inventories. Agreement increased consistently with lake size, whereas mismatches were concentrated among very small lakes and glacier-contact or near-glacier lakes. These patterns are consistent with the greater sensitivity of small lake polygons to pixel size and inventory-production rules, as well as the susceptibility of glacier-margin lakes to differences in acquisition date, water level, ice cover, glacier retreat, and boundary interpretation. Agreement with both an earlier Landsat inventory and a near-contemporaneous Sentinel-2 inventory therefore supports the regional consistency of the present product, while identifying the small-lake and glacier-margin settings in which additional high-resolution or manual verification would be most valuable. Nevertheless, these external inventory comparisons should be interpreted as cross-product consistency assessments rather than as a substitute for validation against a manually checked, independently sampled, and temporally matched high-resolution reference dataset. Therefore, the accuracy of the complete 2024 regional inventory, particularly for very small and glacier-margin lakes, remains to be assessed through targeted independent validation.

4.3. Recognition Performance and Error Sources in Representative Scenarios

To examine model behavior against the principal extraction difficulties identified in the Introduction, representative scenes were selected to cover three major categories of potential failure: omission of small or clustered targets, confusion caused by spectrally complex backgrounds, and incomplete delineation of irregular boundaries. Accordingly, Figure 16 includes small and clustered glacial lakes; scenes affected by snow, shadows, partially frozen surfaces, and exposed bare rock; and lakes with elongated, irregular, or fragmented boundaries. Compared with U-Net, DeepLabV3+, and the baseline YOLO11-seg model, the proposed model produced fewer visible omissions and false detections in the selected scenes and maintained more continuous boundaries for several irregular lake objects. These comparisons provide qualitative evidence of model behavior under the specific difficulties targeted by the multimodal input configuration and architectural modifications.
The higher Precision of the proposed model relative to baseline YOLO11-seg (0.9714 versus 0.9170; Table 3) is consistent with the improved false-positive control observed in the selected representative scenes. Remaining errors mainly involved omissions of very small lakes or lakes with fragmented or weakly defined boundaries, together with boundary offsets on partially frozen or turbid surfaces.

4.4. Uncertainty and Future Improvements

Although this study achieved favorable results in glacial lake extraction and regional-scale mapping in southeastern Tibet, several sources of uncertainty remain. The first is temporal: glacial lake boundaries are affected by image acquisition time, seasonal snow cover, lake ice, and cloud and mountain shadows, while lake water levels, lake-ice conditions, and surrounding snow vary across years and seasons. Although the Sentinel-1/2, DEM, and glacier inventory data provide an effective basis for regional mapping, results derived from a single period or a limited temporal window may not fully represent long-term stable boundaries. A second source arises from post-processing: glacier-buffer filtering, slope filtering, small-patch removal, and boundary smoothing reduce false detections of non-glacial water bodies, steep-slope errors, vector noise, and jagged boundaries, but may also affect a few genuine lakes that lie far from modern glaciers, sit in complex terrain, or are sustained by special supply conditions. A third consideration concerns the training data. The semi-automatic annotation adopted here improves labeling efficiency and boundary consistency, although the delineation of very small, semi-frozen, or turbid lakes still involves some subjectivity. Relatedly, because cloud-covered lakes are relatively under-represented in the training set, the model recovers some cloud-covered lakes through their SAR response but misses others (Figure 13); this can be alleviated by enriching the training samples with more cloud-affected scenes. This may mainly reflect the scarcity of cloud-covered annotations, although further validation is needed to separate data limitations from model limitations. A fourth consideration concerns the interpretation of individual-channel contributions. The present SHAP analysis characterizes how attribution is distributed among the nine channels within the complete model, but it does not by itself establish that every channel is indispensable or non-redundant. Future single-channel removal or modality-specific perturbation experiments could examine redundancy and interactions among individual inputs more directly.
An additional limitation concerns the validation of the complete regional product. The geographically independent test subset provides sample-level evidence of model performance, and the comparisons with the Wang and Tian inventories assess cross-product consistency; however, neither constitutes a dedicated validation of the 2024 inventory against a manually checked, independently sampled, and temporally matched high-resolution reference dataset. Regional-product omission, commission, and boundary uncertainty may therefore remain, especially for very small, shadowed, riverine-wetland, and glacier-margin lakes.
Future research could improve this work along three directions. First, multi-temporal Sentinel-1/2 imagery could be used to compare extraction across different seasons, years, and lake-surface conditions, so as to evaluate the temporal stability of the model. Second, higher-resolution imagery, existing inventories, and manual verification in key regions could be combined to independently validate small, shadowed, and riverine-wetland lakes, allowing the uncertainty of the regional products to be assessed more accurately; in particular, constructing more complex training samples that include cloud-covered scenes, together with cloud-aware augmentation, would further improve the robustness of the model under cloud cover. Third, building on the current boundaries, area classes, and elevation-zone statistics, future work could integrate glacier retreat, downstream topography, moraine-dam type, basin convergence, and historical change, extending automated mapping from boundary extraction to dynamic monitoring and glacial lake outburst flood risk identification, and improving the updatability and practical value of glacial lake inventories for southeastern Tibet.

5. Conclusions

This study developed an improved YOLO11-seg framework for automatic glacial lake extraction in southeastern Tibet using nine-channel optical, water-index, topographic, and SAR inputs. In both the input-combination and module ablation experiments, the complete configuration achieved the highest F1-Score and mAP50(M) among the tested configurations. The model achieved Precision, Recall, F1-Score, and mAP50(M) values of 0.9714, 0.8746, 0.9205, and 0.9331, respectively. On the independent test set, Recall, F1-Score, and IoU reached 0.9283, 0.9209, and 0.8533, respectively.
Applied to imagery acquired in 2024, the model mapped 4766 glacial lakes covering 428.13 km2 in southeastern Tibet. Most were distributed between 4000 and 5500 m, and glacial lakes smaller than 0.10 km2 accounted for 80.84% of the total number, although larger lakes dominated the total area. The resulting inventory provides updated spatial information for regional glacial lake characterization and subsequent GLOF hazard assessment, while the multimodal deep-learning framework offers a methodological reference for glacial lake inventorying in environmentally similar high-mountain regions.

Author Contributions

K.L.: conceptualization, methodology, formal analysis, investigation, data curation, and writing—original draft. T.G.: conceptualization, methodology, supervision, project administration, funding acquisition, and writing—review and editing. S.Y.: conceptualization, methodology, supervision, and writing—review and editing. X.L.: methodology, resources, supervision, and writing—review and editing. Y.C.: resources and supervision. S.H.: investigation and validation. H.Z.: data curation and formal analysis. Z.S.: data curation and validation. H.L.: data curation and investigation. M.L.: data curation and validation. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the Science and Technology Projects of Xizang Autonomous Region, China (Grant No. XZ202501ZY0004, Grant No. XZ202401JD0001 and Grant No. XZ202601ZD0009), the Key Research and Development Program of Xizang Autonomous Region, China (Grant No. XZ202502ZY0021), and the Graduate Education Innovation Program of Xizang Agriculture and Animal Husbandry University (Grant No. SGYC2026001).

Data Availability Statement

The data presented in this study are available from the corresponding author upon reasonable request.

Acknowledgments

We thank the European Space Agency for providing Sentinel-1 SAR and Sentinel-2 MSI data, the Copernicus programme for providing Copernicus DEM GLO-30 data, and the Google Earth Engine platform for data access and processing support. We also thank the providers of the Second Chinese Glacier Inventory and Randolph Glacier Inventory 7.0 for sharing the glacier boundary data used in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Wang, X.; Guo, X.; Yang, C.; Liu, Q.; Wei, J.; Zhang, Y.; Liu, S.; Zhang, Y.; Jiang, Z.; Tang, Z. Glacial lake inventory of High Mountain Asia (1990–2018) derived from Landsat images. Earth Syst. Sci. Data 2020, 12, 2169–2182. [Google Scholar] [CrossRef] [Scilit]
  2. Sahu, R.; Singh, D.P. Conventional and deep learning approaches for glacial lake mapping using remote sensing data: A comprehensive review. Int. J. Remote Sens. 2025, 46, 3992–4020. [Google Scholar] [CrossRef] [Scilit]
  3. Khan, I.; Vardaman, M.; Johnston, J.; Jacobs, J.M. Glacial lake mapping: Progress, challenges, and future directions. Arct. Antarct. Alp. Res. 2025, 57, 2592365. [Google Scholar] [CrossRef] [Scilit]
  4. Song, C.; Fan, C.; Ma, J.; Zhan, P.; Deng, X. A Spatially Constrained Remote Sensing-Based Inventory of Glacial Lakes Worldwide. Sci. Data 2025, 12, 464. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Lesi, M.; Nie, Y.; Shugar, D.H.; Wang, J.; Deng, Q.; Chen, H.; Fan, J. Landsat- and Sentinel-Derived Glacial Lake Dataset in the China–Pakistan Economic Corridor from 1990 to 2020. Earth Syst. Sci. Data 2022, 14, 5489–5512. [Google Scholar] [CrossRef] [Scilit]
  6. Chand, M.B.; Watanabe, T. Development of Supraglacial Ponds in the Everest Region, Nepal, between 1989 and 2018. Remote Sens. 2019, 11, 1058. [Google Scholar] [CrossRef] [Scilit]
  7. Tong, J.Z. Sentinel-1/2 Remote Sensing Monitoring and Analysis of the Dynamic Evolution of Ice Lakes in the Parlung Zangbo on the Qinghai-Tibet Plateau. Master’s Thesis, Southwest Jiaotong University, Chengdu, China, 2023. (In Chinese) [Google Scholar]
  8. Hu, J.; Zhang, T.; Zhou, X.; Yi, G.; Bie, X.; Li, J.; Chen, Y.; Lai, P. A glacial lake mapping framework in high mountain areas: A case study of the southeastern Tibetan Plateau. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4300212. [Google Scholar] [CrossRef] [Scilit]
  9. Zhang, Y.; Zhao, J.; Yao, X.; Duan, H.; Yang, J.; Pang, W. Inventory of glacial lake in the southeastern Qinghai-Tibet Plateau derived from Sentinel-1 SAR image and Sentinel-2 MSI image. Remote Sens. 2023, 15, 5142. [Google Scholar] [CrossRef] [Scilit]
  10. Yu, M.; Guo, Y.; Zhang, J.; Li, F.; Su, L.; Qin, D. Spaciotemporal Distribution Characteristics of Glacial Lakes and the Factors Influencing the Southeast Tibetan Plateau from 1993 to 2023. Sci. Rep. 2025, 15, 1966. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Tian, Z.; Yao, X.; Chang, H.; Duan, H. A Dataset of the Glacial Lake Inventory on the Tibetan Plateau in 2023. China Sci. Data 2026, 11, 1–16. [Google Scholar] [CrossRef] [Scilit]
  12. Li, H.; Dou, J.; Kusky, T.; Dong, S.; Shi, Z.; Li, J.; Xiang, X.; Ding, F. SETP_GLI: An Annual 10–30 m Glacial Lake Inventory for the Southeastern Tibetan Plateau from 1990 to 2025. Earth Syst. Sci. Data Discuss. 2026. [Google Scholar] [CrossRef] [Scilit]
  13. Zhao, H.; Chen, F.; Zhang, M. A systematic extraction approach for mapping glacial lakes in high mountain regions of Asia. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2018, 11, 2788–2799. [Google Scholar] [CrossRef] [Scilit]
  14. Zhang, M.; Chen, F.; Tian, B. An automated method for glacial lake mapping in High Mountain Asia using Landsat 8 imagery. J. Mt. Sci. 2018, 15, 13–24. [Google Scholar] [CrossRef] [Scilit]
  15. Li, Y.; Zhang, J.; Liu, C. Extraction method of alpine small glacial lake in Qianhu Mountain area of Yunnan Province based on Sentinel-2 image. Sci. Surv. Mapp. 2021, 46, 114–120. [Google Scholar] [CrossRef]
  16. McFeeters, S.K. The use of the normalized difference water index (NDWI) in the delineation of open water features. Int. J. Remote Sens. 1996, 17, 1425–1432. [Google Scholar] [CrossRef] [Scilit]
  17. Xu, H. Modification of normalised difference water index (NDWI) to enhance open water features in remotely sensed imagery. Int. J. Remote Sens. 2006, 27, 3025–3033. [Google Scholar] [CrossRef] [Scilit]
  18. Li, X.; Huang, Z.; Xiao, C.; Lan, Z.; Liu, B. A semi-automatic method for extracting glacial lake boundary based on MNDWI: A case study of TM images of Sangwang Co. In Proceedings of the 2013 Annual Conference of Chinese Hydraulic Engineering Society, Guangzhou, China, 26–27 November 2013; pp. 1–7. [Google Scholar]
  19. Yan, B.; Jia, H.; Ren, W.; Wu, R.; Huang, X. Extraction and variation analysis of Bugagangri glacial lake based on NDWI-NDSI combined threshold method. Natl. Remote Sens. Bull. 2022, 26, 2344–2353. [Google Scholar] [CrossRef] [Scilit]
  20. Li, M.; Zheng, J.; Qian, A.; Li, J.; Parhat, A.; Wang, Z.; Ma, L.; Wang, N. Research on the extraction method of Tianshan glacial lake based on decision tree. Arid Zone Res. 2024, 41, 1699–1707. [Google Scholar]
  21. Li, J.; Sheng, Y.; Luo, J. Automatic extraction of Himalayan glacial lakes with remote sensing. J. Remote Sens. 2011, 15, 29–43. [Google Scholar] [CrossRef] [Scilit]
  22. Chen, F.; Wang, J.; Zhang, M.; Yu, B. Comparative study on the extraction methods of Himalayan glacial lakes based on historical boundaries. J. Glaciol. Geocryol. 2023, 45, 1413–1427. [Google Scholar] [CrossRef]
  23. Zhao, H.; Chen, F.; Zhang, M. Research of glacial lake contour extraction method based on improved C-V model. Remote Sens. Technol. Appl. 2018, 33, 177–184. [Google Scholar] [CrossRef]
  24. Wang, Z.; Ran, Y. Glacial lake information extraction method in mountainous areas by removing cloud shadow from remote sensing images. Beijing Surv. Mapp. 2019, 33, 1236–1239. [Google Scholar] [CrossRef]
  25. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015; Springer: Cham, Switzerland, 2015; Volume 9351, pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
  26. Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Computer Vision—ECCV 2018; Springer: Cham, Switzerland, 2018; pp. 833–851. [Google Scholar] [CrossRef] [Scilit]
  27. Cao, Y. Research on Automatic Detection Method of Himalayan Glacial Lakes Based on Deep Convolutional Neural Network. Master’s Thesis, University of Electronic Science and Technology of China, Chengdu, China, 2022. [Google Scholar]
  28. Wang, J.; Chen, F.; Zhang, M.; Yu, B. NAU-Net: A new deep learning framework in glacial lake detection. IEEE Geosci. Remote Sens. Lett. 2022, 19, 1–5. [Google Scholar] [CrossRef] [Scilit]
  29. Yin, L.; Wang, X.; Yin, Y.; Wang, Q.; Lei, D.; Lian, W.; Zhang, Y.; Wei, J. Automatic extraction of glacial lakes based on deep learning and Sentinel-2 imagery. Remote Sens. Technol. Appl. 2024, 39, 1319–1329. [Google Scholar]
  30. Sharma, A.; Thakur, V.; Prakash, C.; Sharma, A.; Sharma, R. Deep learning-based glacial lakes extraction and mapping in the Chandra-Bhaga Basin. J. Indian Soc. Remote Sens. 2024, 52, 435–447. [Google Scholar] [CrossRef] [Scilit]
  31. Basit, A.; Bhatti, M.K.; Ali, M.; Fatima, T.; Minchew, B.; Siddique, M.A. Deep learning for monitoring glacial lakes formation using Sentinel 2 multispectral data. In Proceedings of the IGARSS 2022—2022 IEEE International Geoscience and Remote Sensing Symposium, Kuala Lumpur, Malaysia, 17–22 July 2022; pp. 179–182. [Google Scholar] [CrossRef] [Scilit]
  32. Siddique, M.A.; Basit, A.; Qayyum, N.; Naseer, E.; Bhatti, M.K.; Minchew, B.; Ali, M.; Perez, C.S.; Marino, A. Towards automated monitoring of glacial lakes in Hindu Kush and Himalayas using deep learning. In Proceedings of the IGARSS 2023—2023 IEEE International Geoscience and Remote Sensing Symposium, Pasadena, CA, USA, 16–21 July 2023; pp. 2165–2168. [Google Scholar] [CrossRef] [Scilit]
  33. Thati, J.; Ari, S. A systematic extraction of glacial lakes for satellite imagery using deep learning based technique. SSRN Electron. J. 2022. [Google Scholar] [CrossRef] [Scilit]
  34. Yang, N.; Nie, Y. An improved deep learning method for mapping glacial lakes using satellite observation and its application. Spacecr. Recovery Remote Sens. 2024, 45, 41–52. [Google Scholar] [CrossRef]
  35. Tang, Q.; Zhang, G.; Yao, T.; Wieland, M.; Liu, L.; Kaushik, S. Automatic extraction of glacial lakes from Landsat imagery using deep learning across the Third Pole region. Remote Sens. Environ. 2024, 315, 114413. [Google Scholar] [CrossRef] [Scilit]
  36. Zhao, H.; Wang, S.; Liu, X.; Chen, F. Exploring contrastive representation for weakly-supervised glacial lake extraction. Remote Sens. 2023, 15, 1456. [Google Scholar] [CrossRef] [Scilit]
  37. Tan, H.; Jiang, L.; Liu, H.; Zhang, T.; Cheng, I. Prior knowledge-informed semantic segmentation framework for precise glacial lake mapping from multimodal imagery. ISPRS J. Photogramm. Remote Sens. 2025, 230, 630–643. [Google Scholar] [CrossRef] [Scilit]
  38. Zhang, S.; Li, Q.; Li, H.; Xiang, X.; Dong, A.; Dou, J. Intelligent recognition of glacial lakes in complex plateau terrain by integrating multi-source remote sensing data and an improved Mask R-CNN deep learning model. Earth Sci. 2025, 50, 3132–3143. [Google Scholar] [CrossRef] [Scilit]
  39. Yin, L.; Wang, X.; Du, W.; Yang, C.; Wei, J.; Wang, Q.; Lei, D.; Xiao, J. Using the improved YOLOv5-Seg network and Sentinel-2 imagery to map glacial lakes in High Mountain Asia. Remote Sens. 2024, 16, 2057. [Google Scholar] [CrossRef] [Scilit]
  40. Wan, D.; Lu, R.; Wang, S.; Shen, S.; Xu, T.; Lang, X. YOLO-HR: Improved YOLOv5 for object detection in high-resolution optical remote sensing images. Remote Sens. 2023, 15, 614. [Google Scholar] [CrossRef] [Scilit]
  41. Maurya, L.; Kaushik, S.; Tellman, B. GLACIA: Instance-aware positional reasoning for glacial lake segmentation via multimodal large language model. arXiv 2025, arXiv:2512.09251. [Google Scholar]
  42. Wu, R.; Liu, G.; Zhang, R.; Wang, X.; Li, Y.; Zhang, B.; Cai, J.; Xiang, W. A deep learning method for mapping glacial lakes from the combined use of synthetic-aperture radar and optical satellite images. Remote Sens. 2020, 12, 4020. [Google Scholar] [CrossRef] [Scilit]
  43. Wu, R. Deep Learning Models and Methods for Extracting Spatio-Temporal Information of Alpine Glacial Lakes by Combining Optical and SAR Remote Sensing Images. Master’s Thesis, Southwest Jiaotong University, Chengdu, China, 2021. [Google Scholar] [CrossRef]
  44. Guo, S. Water-Land Boundary Extraction and Evolution Analysis of Volcanic Islands and Glacial Lakes by Combining Optical and SAR Images. Master’s Thesis, Chengdu University of Technology, Chengdu, China, 2022. [Google Scholar] [CrossRef]
  45. Wu, R.; Liu, G.; Bao, X.; Lv, J.; Shama, A.; Zhang, B.; Mao, W.; Chen, J.; Yang, Z.; Zhang, R. Eliminating geometric distortion with dual-orbit Sentinel-1 SAR fusion for accurate glacial lake extraction in Southeast Tibet Plateau. Int. J. Appl. Earth Obs. Geoinf. 2025, 136, 104329. [Google Scholar] [CrossRef] [Scilit]
  46. Wang, J.; Chen, F.; Zhang, M.; Yu, B. ACFNet: A feature fusion network for glacial lake extraction based on optical and synthetic aperture radar images. Remote Sens. 2021, 13, 5091. [Google Scholar] [CrossRef] [Scilit]
  47. Ali, S.A.; Ali, S.; Wang, L.; Hassan, S.R.U.; Khan, G.; Ali, S.; Hussain, T.; Ali, D. Deep learning model to detect glacial lakes using high-resolution optical and radar satellite images. SSRN Electron. J. 2024. [Google Scholar] [CrossRef] [Scilit]
  48. Gorelick, N.; Hancher, M.; Dixon, M.; Ilyushchenko, S.; Thau, D.; Moore, R. Google Earth Engine: Planetary-scale geospatial analysis for everyone. Remote Sens. Environ. 2017, 202, 18–27. [Google Scholar] [CrossRef] [Scilit]
  49. Chen, F.; Zhang, M.; Tian, B.; Li, Z. Extraction of glacial lake outlines in Tibet Plateau using Landsat 8 imagery and Google Earth Engine. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2017, 10, 4002–4009. [Google Scholar] [CrossRef] [Scilit]
  50. Wu, R.; Zhang, R.; Cai, J.; Cai, K.; Wu, T.; Lv, J.; Shi, Y.; Liu, G. Glacial lake extraction framework based on coupling of GEE and historical glacial lake position. IEEE Geosci. Remote Sens. Lett. 2024, 21, 1–5. [Google Scholar] [CrossRef] [Scilit]
  51. Ma, D.; Li, J.; Jiang, L. Efficient glacial lake mapping by leveraging deep transfer learning and a new annotated glacial lake dataset. J. Hydrol. 2025, 657, 133072. [Google Scholar] [CrossRef] [Scilit]
  52. Thati, J.; Ari, S. GLeSI: A system for extraction of glacial lakes using satellite imagery. Concurr. Comput. Pract. Exp. 2022, 34, e7184. [Google Scholar] [CrossRef] [Scilit]
  53. Jiang, X.; Gu, C.; Nie, Y.; Hu, M.; Lyu, Q.; Wang, W. A new method for detecting automated mapping anomalies in Himalayan glacial lakes from satellite images. Remote Sens. 2026, 18, 61. [Google Scholar] [CrossRef] [Scilit]
  54. Liu, S.; Yao, X.; Guo, W.; Xu, J.; Shangguan, D.; Wei, J.; Bao, W.; Wu, L. The Contemporary Glaciers in China Based on the Second Chinese Glacier Inventory. Acta Geogr. Sin. 2015, 70, 3–16. [Google Scholar] [CrossRef]
  55. RGI Consortium. Randolph Glacier Inventory—A Dataset of Global Glacier Outlines, Version 7.0; National Snow and Ice Data Center: Boulder, CO, USA, 2023. [CrossRef]
  56. Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.-Y.; et al. Segment Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 3992–4003. [Google Scholar] [CrossRef] [Scilit]
  57. Wang, W. X-AnyLabeling: Advanced Auto Labeling Solution with Added Features. 2023. Available online: https://github.com/CVHub520/X-AnyLabeling (accessed on 24 June 2026).
  58. Sunkara, R.; Luo, T. No more strided convolutions or pooling: A new CNN building block for low-resolution images and small objects. In Machine Learning and Knowledge Discovery in Databases; Springer: Cham, Switzerland, 2023; pp. 443–459. [Google Scholar] [CrossRef] [Scilit]
  59. Guo, M.-H.; Lu, C.-Z.; Hou, Q.; Liu, Z.-N.; Cheng, M.-M.; Hu, S.-M. SegNeXt: Rethinking convolutional attention design for semantic segmentation. arXiv 2022, arXiv:2209.08575. [Google Scholar]
  60. Ding, Q.; Li, W.; Xu, C.; Zhang, M.; Sheng, C.; He, M.; Shan, N. GMS-YOLO: An algorithm for multi-scale object detection in complex environments in confined compartments. Sensors 2024, 24, 5789. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
  62. Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual state space model. arXiv 2024, arXiv:2401.10166. [Google Scholar]
  63. Wang, J.; Chen, K.; Xu, R.; Liu, Z.; Loy, C.C.; Lin, D. CARAFE: Content-aware reassembly of features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 3007–3016. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Location and environmental setting of the study area in southeastern Tibet, together with the spatial distribution of the ten sample regions used for dataset construction and partitioning. Black arrows indicate the nested geographic relationship from China to Xizang and from Xizang to southeastern Xizang. The national boundary shown in the China inset was based on the standard map with approval number GS (2024) 0650.
Figure 1. Location and environmental setting of the study area in southeastern Tibet, together with the spatial distribution of the ten sample regions used for dataset construction and partitioning. Black arrows indicate the nested geographic relationship from China to Xizang and from Xizang to southeastern Xizang. The national boundary shown in the China inset was based on the standard map with approval number GS (2024) 0650.
Remotesensing 18 02769 g001
Figure 2. Overall workflow of the proposed multimodal remote sensing framework for glacial lake extraction and regional inventory generation in southeastern Tibet. The framework comprises five phases: (1) acquisition and preprocessing of Copernicus DEM GLO-30, Sentinel-1 SAR, and Sentinel-2 MSI data, followed by band fusion; (2) sample creation through the combination of manual and deep-learning-model-assisted annotation; (3) training and evaluation of the glacial lake segmentation model; (4) nine-channel tile inference and post-processing, including glacier-buffer filtering, slope filtering, and morphological processing; and (5) generation of glacial lake vector, raster, and mapping products. Black arrows indicate sequential data-processing steps and network feature propagation within a phase; light-gray arrows indicate local annotation and post-processing operations; red arrows indicate the transfer of datasets, labels, training inputs, trained models, inference results, and mapping products across the major workflow stages; and blue arrows indicate the transfer of multimodal input tensors into the network and the inference/output flow in Phases 3 and 4. Plus signs indicate the integration of manual and model-assisted annotation and the combination of image tiles with their corresponding label masks.
Figure 2. Overall workflow of the proposed multimodal remote sensing framework for glacial lake extraction and regional inventory generation in southeastern Tibet. The framework comprises five phases: (1) acquisition and preprocessing of Copernicus DEM GLO-30, Sentinel-1 SAR, and Sentinel-2 MSI data, followed by band fusion; (2) sample creation through the combination of manual and deep-learning-model-assisted annotation; (3) training and evaluation of the glacial lake segmentation model; (4) nine-channel tile inference and post-processing, including glacier-buffer filtering, slope filtering, and morphological processing; and (5) generation of glacial lake vector, raster, and mapping products. Black arrows indicate sequential data-processing steps and network feature propagation within a phase; light-gray arrows indicate local annotation and post-processing operations; red arrows indicate the transfer of datasets, labels, training inputs, trained models, inference results, and mapping products across the major workflow stages; and blue arrows indicate the transfer of multimodal input tensors into the network and the inference/output flow in Phases 3 and 4. Plus signs indicate the integration of manual and model-assisted annotation and the combination of image tiles with their corresponding label masks.
Remotesensing 18 02769 g002
Figure 3. Representative annotated samples illustrating three glacial lake settings: supraglacial lakes, glacier-contact lakes, and non-glacier-contact lakes. Red outlines delineate the manually annotated glacial lake instance boundaries used as reference labels.
Figure 3. Representative annotated samples illustrating three glacial lake settings: supraglacial lakes, glacier-contact lakes, and non-glacier-contact lakes. Red outlines delineate the manually annotated glacial lake instance boundaries used as reference labels.
Remotesensing 18 02769 g003
Figure 4. Architecture of the enhanced YOLO11 glacial lake segmentation network. Yellow blocks denote standard convolutional, C3k2/C3k, and downsampling operations; green, blue, cyan, orange, and purple blocks denote SPDConv, MSCA/C2PSA_MSCA, CARAFE upsampling, C3k2_Mamba/bottleneck, and SPPF operations, respectively. Pink blocks denote feature concatenation, and light-blue blocks denote segmentation heads. Black arrows indicate feature propagation, the circled plus sign indicates residual addition, and dashed outlines group the input/output examples and the backbone, neck, and head components.
Figure 4. Architecture of the enhanced YOLO11 glacial lake segmentation network. Yellow blocks denote standard convolutional, C3k2/C3k, and downsampling operations; green, blue, cyan, orange, and purple blocks denote SPDConv, MSCA/C2PSA_MSCA, CARAFE upsampling, C3k2_Mamba/bottleneck, and SPPF operations, respectively. Pink blocks denote feature concatenation, and light-blue blocks denote segmentation heads. Black arrows indicate feature propagation, the circled plus sign indicates residual addition, and dashed outlines group the input/output examples and the backbone, neck, and head components.
Remotesensing 18 02769 g004
Figure 5. Structure of the SPDConv module. Blue, yellow, red, and green cells denote the four interleaved spatial offsets sampled from the input feature map, yielding X00, X10, X01, and X11, respectively. The four sub-feature maps are concatenated along the channel dimension and then processed by convolution, batch normalization (BN), and SiLU activation to produce the output feature Y. Black arrows indicate the processing sequence, and gray dashed lines separate the major processing stages.
Figure 5. Structure of the SPDConv module. Blue, yellow, red, and green cells denote the four interleaved spatial offsets sampled from the input feature map, yielding X00, X10, X01, and X11, respectively. The four sub-feature maps are concatenated along the channel dimension and then processed by convolution, batch normalization (BN), and SiLU activation to produce the output feature Y. Black arrows indicate the processing sequence, and gray dashed lines separate the major processing stages.
Remotesensing 18 02769 g005
Figure 6. Structures of the (a) MSCA and (b) C2PSA_MSCA modules. In (a), the blue dashed region denotes multi-scale convolutional attention and the green dashed region denotes spatial attention. Green, yellow, and purple blocks distinguish the strip-convolution branches with kernel extents of 7, 11, and 21, respectively, while blue blocks denote the direct and convolutional feature paths. In (b), blue and orange feature maps distinguish the two split branches, and the orange dashed box denotes the repeated PSABlock_MSCA unit. Circled plus and multiplication signs denote elementwise addition and multiplication, respectively, and arrows indicate feature flow.
Figure 6. Structures of the (a) MSCA and (b) C2PSA_MSCA modules. In (a), the blue dashed region denotes multi-scale convolutional attention and the green dashed region denotes spatial attention. Green, yellow, and purple blocks distinguish the strip-convolution branches with kernel extents of 7, 11, and 21, respectively, while blue blocks denote the direct and convolutional feature paths. In (b), blue and orange feature maps distinguish the two split branches, and the orange dashed box denotes the repeated PSABlock_MSCA unit. Circled plus and multiplication signs denote elementwise addition and multiplication, respectively, and arrows indicate feature flow.
Remotesensing 18 02769 g006
Figure 7. Structure of the C3k2_Mamba module.
Figure 7. Structure of the C3k2_Mamba module.
Remotesensing 18 02769 g007
Figure 8. Structure of the CARAFE module. Blue feature maps denote the input and intermediate features used for channel compression and local patch extraction, whereas green feature maps denote the predicted reassembly kernels and the upsampled output. The upper and lower dashed boxes identify the kernel-prediction and content-aware reassembly modules, respectively. Black arrows indicate feature flow, and the circled multiplication symbol denotes the weighted reassembly of each local patch using the normalized kernel weights.
Figure 8. Structure of the CARAFE module. Blue feature maps denote the input and intermediate features used for channel compression and local patch extraction, whereas green feature maps denote the predicted reassembly kernels and the upsampled output. The upper and lower dashed boxes identify the kernel-prediction and content-aware reassembly modules, respectively. Black arrows indicate feature flow, and the circled multiplication symbol denotes the weighted reassembly of each local patch using the normalized kernel weights.
Remotesensing 18 02769 g008
Figure 9. Channel-level and input-group SHAP attribution of the complete 9-channel model based on 1000 validation samples. (a) Mean absolute SHAP values and relative contributions of the nine input channels, together with the sample-level signed SHAP distributions. Dot colors indicate normalized input values within each channel, ranging from low to high. (b) Absolute SHAP composition of the five predefined input groups, calculated by summing the mean absolute SHAP values of their constituent channels.
Figure 9. Channel-level and input-group SHAP attribution of the complete 9-channel model based on 1000 validation samples. (a) Mean absolute SHAP values and relative contributions of the nine input channels, together with the sample-level signed SHAP distributions. Dot colors indicate normalized input values within each channel, ranging from low to high. (b) Absolute SHAP composition of the five predefined input groups, calculated by summing the mean absolute SHAP values of their constituent channels.
Remotesensing 18 02769 g009
Figure 10. Overlay of the 9-channel multimodal remote sensing features: (A) Sentinel-2 B4/B3/B2 RGB, (B) B8 NIR, (C) B11 SWIR, (D) NDWI, (E) MNDWI, (F) slope derived from Copernicus DEM GLO-30, and (G) Sentinel-1 VV backscatter. For visualization, RGB is shown in natural color; NIR and SWIR use false-color rendering; NDWI, MNDWI, and Sentinel-1 VV are displayed in grayscale; and slope is shown using a pseudocolor gradient. Brighter tones in NDWI and MNDWI indicate stronger water-index responses.
Figure 10. Overlay of the 9-channel multimodal remote sensing features: (A) Sentinel-2 B4/B3/B2 RGB, (B) B8 NIR, (C) B11 SWIR, (D) NDWI, (E) MNDWI, (F) slope derived from Copernicus DEM GLO-30, and (G) Sentinel-1 VV backscatter. For visualization, RGB is shown in natural color; NIR and SWIR use false-color rendering; NDWI, MNDWI, and Sentinel-1 VV are displayed in grayscale; and slope is shown using a pseudocolor gradient. Brighter tones in NDWI and MNDWI indicate stronger water-index responses.
Remotesensing 18 02769 g010
Figure 11. Spatial distribution of extracted glacial lakes in southeastern Tibet. (Red points indicate extracted glacial lakes, blue areas indicate glacier boundaries, and light-blue areas indicate the 10 km glacier buffer.)
Figure 11. Spatial distribution of extracted glacial lakes in southeastern Tibet. (Red points indicate extracted glacial lakes, blue areas indicate glacier boundaries, and light-blue areas indicate the 10 km glacier buffer.)
Remotesensing 18 02769 g011
Figure 12. Object-level agreement between the regional glacial lake inventory produced in this study and the independent inventories of Wang et al. (2020) [1] and Tian et al. (2026) [11]. (a) Overall reference-object match rate and matched reference-area proportion; (b) empirical cumulative distribution of the maximum Jaccard index; (c) reference-object match rate by lake-area class; (d) median maximum Jaccard index and interquartile range by lake-area class; (e) reference-object match rate by elevation zone; and (f) reference-object match rate by distance to the nearest glacier. Glacier distance was calculated for both reference inventories using the same merged glacier boundaries from the Second Chinese Glacier Inventory and RGI 7.0 [54,55].
Figure 12. Object-level agreement between the regional glacial lake inventory produced in this study and the independent inventories of Wang et al. (2020) [1] and Tian et al. (2026) [11]. (a) Overall reference-object match rate and matched reference-area proportion; (b) empirical cumulative distribution of the maximum Jaccard index; (c) reference-object match rate by lake-area class; (d) median maximum Jaccard index and interquartile range by lake-area class; (e) reference-object match rate by elevation zone; and (f) reference-object match rate by distance to the nearest glacier. Glacier distance was calculated for both reference inventories using the same merged glacier boundaries from the Second Chinese Glacier Inventory and RGI 7.0 [54,55].
Remotesensing 18 02769 g012
Figure 13. Qualitative validation of channel effectiveness in cloud- and shadow-affected scenes. From left to right: Sentinel-2 RGB; RGB overlaid with reference lake boundaries (RGB-Label); Sentinel-1 SAR (VV); the extraction result of the RGB setting (S1); and the result of the full multimodal setting (S5). Red boxes highlight representative locations where the full multimodal model (S5) improves over the RGB setting (S1).
Figure 13. Qualitative validation of channel effectiveness in cloud- and shadow-affected scenes. From left to right: Sentinel-2 RGB; RGB overlaid with reference lake boundaries (RGB-Label); Sentinel-1 SAR (VV); the extraction result of the RGB setting (S1); and the result of the full multimodal setting (S5). Red boxes highlight representative locations where the full multimodal model (S5) improves over the RGB setting (S1).
Remotesensing 18 02769 g013
Figure 14. Elevation-zone and area-class distribution characteristics of glacial lakes in southeastern Tibet. (a) Percentage of glacial lake number and area across elevation zones; (b) percentage of glacial lake number and area across area classes; (c) number and total area of glacial lakes across elevation zones; (d) number and total area of glacial lakes across area classes.
Figure 14. Elevation-zone and area-class distribution characteristics of glacial lakes in southeastern Tibet. (a) Percentage of glacial lake number and area across elevation zones; (b) percentage of glacial lake number and area across area classes; (c) number and total area of glacial lakes across elevation zones; (d) number and total area of glacial lakes across area classes.
Remotesensing 18 02769 g014
Figure 15. Aggregate comparison of the glacial lake inventory produced in this study (2024) with the inventories of Wang et al. (2020; inventory epoch 2018) [1] and Tian et al. (2026; inventory epoch 2023) [11] within the study area and the common 10 km glacier-buffer extent. (a) Number of lakes by area class; (b) total lake area by area class; (c) number of lakes by elevation zone; and (d) total lake area by elevation zone.
Figure 15. Aggregate comparison of the glacial lake inventory produced in this study (2024) with the inventories of Wang et al. (2020; inventory epoch 2018) [1] and Tian et al. (2026; inventory epoch 2023) [11] within the study area and the common 10 km glacier-buffer extent. (a) Number of lakes by area class; (b) total lake area by area class; (c) number of lakes by elevation zone; and (d) total lake area by elevation zone.
Remotesensing 18 02769 g015
Figure 16. Qualitative comparison across representative glacial lake scenarios, showing RGB reference imagery, ground-truth masks, and segmentation results from the proposed model, YOLO11-seg, U-Net, and DeepLabV3+. Red boxes highlight representative omissions, false detections, or boundary differences.
Figure 16. Qualitative comparison across representative glacial lake scenarios, showing RGB reference imagery, ground-truth masks, and segmentation results from the proposed model, YOLO11-seg, U-Net, and DeepLabV3+. Red boxes highlight representative omissions, false detections, or boundary differences.
Remotesensing 18 02769 g016
Table 1. Main data sources used in this study. (Sentinel-2 B11 and Copernicus DEM GLO-30 were resampled and spatially aligned to a 10 m grid.)
Table 1. Main data sources used in this study. (Sentinel-2 B11 and Copernicus DEM GLO-30 were resampled and spatially aligned to a 10 m grid.)
Data TypeProduct/Dataset IDSpatial
Resolution
Data SourcePurpose
Sentinel-2 MSICOPERNICUS/S2_SR_HARMONIZED10 m/20 mGoogle Earth Engine Data CatalogOptical, multispectral, and water-index features
Sentinel-1 SAR (VV)COPERNICUS/S1_GRD10 mSAR backscatter feature
Copernicus DEM GLO-30COPERNICUS/DEM/GLO3030 mTopographic constraint
Second Chinese Glacier InventoryThe Second Glacier Inventory Dataset of China, Version 1.0Vector dataNational Tibetan Plateau Data CenterSample checking and 10 km glacier-buffer filtering
Randolph Glacier InventoryRGI 7.0/NSIDC-0770Vector dataNSIDC
Table 2. Cloud-cover thresholds, temporal windows, patch counts, and dataset split for the ten sample regions.
Table 2. Cloud-cover thresholds, temporal windows, patch counts, and dataset split for the ten sample regions.
RegionStepwise Cloud-Cover Threshold Range/%Time WindowPatchesDataset Split
Region 11–51 September 2024–30 November 20241920Train
Region 20.1–0.51 January 2024–31 December 20246741
Region 35–101 September 2024–30 November 20241653
Region 40.1–51 May 2024–30 November 20242200
Region 50.1–51 May 2024–30 November 2024663
Region 60.1–21 May 2024–30 November 20242223
Region 70.1–21 May 2024–30 November 2024798Val
Region 80.1–11 September 2024–30 November 2024592
Region 90.5–11 May 2024–31 December 20242479Test
Region 100.5–1.51 May 2024–30 November 2024104
Total--19,373-
Table 3. Performance comparison of mainstream segmentation models.
Table 3. Performance comparison of mainstream segmentation models.
ModelPrecisionRecallF1-ScoremAP50(M)
YOLOv8-seg0.88030.81640.84710.8798
YOLO11-seg0.91700.81290.86180.8898
YOLO12-seg0.84790.82230.83490.8773
YOLO26-seg0.82790.79770.81250.8544
U-Net0.94980.83160.8868-
DeepLabV3+0.92660.86850.8966-
Proposed0.97140.87460.92050.9331
Table 4. Ablation experiment of the proposed modules.
Table 4. Ablation experiment of the proposed modules.
SettingSPDConvMSCAC3k2_MambaC2PSA_MSCACARAFEPrecisionRecallF1-ScoremAP50(M)
Baseline0.91700.81290.86180.8898
10.96620.71960.82490.8499
20.93140.82930.87740.9045
30.97100.86630.91570.9284
40.96960.86820.91610.9287
50.96390.87430.91700.9301
60.95960.87760.91670.9326
70.95510.80610.87430.8926
80.96770.87220.91750.9301
90.97140.87460.92050.9331
Table 5. Ablation experiment of the 9-channel input combination.
Table 5. Ablation experiment of the 9-channel input combination.
SettingInput CombinationChannelsPrecisionRecallF1-ScoremAP50(M)
S1RGB30.90070.85090.87510.8998
S2RGB + NIR + SWIR50.91460.86420.88870.9155
S3RGB + NIR + SWIR + NDWI + MNDWI70.90580.89870.90220.9273
S4RGB + NIR + SWIR + NDWI + MNDWI + Slope80.90400.88930.89660.9299
S5RGB + NIR + SWIR + NDWI + MNDWI + Slope + SAR90.97140.87460.92050.9331
Table 6. Quantitative comparison of the main models on the test set.
Table 6. Quantitative comparison of the main models on the test set.
ModelPrecisionRecallF1-ScoreIoU
U-Net0.92940.85570.89100.8035
DeepLabV3+0.92640.87500.89990.8181
YOLOv8-seg0.94360.80570.86920.7687
YOLO11-seg0.94030.82640.87960.7851
YOLO12-seg0.94380.79340.86210.7577
YOLO26-seg0.93490.77880.84980.7388
Proposed0.91350.92830.92090.8533
Table 7. Elevation-zone distribution statistics of glacial lakes in southeastern Tibet.
Table 7. Elevation-zone distribution statistics of glacial lakes in southeastern Tibet.
Elevation ZoneNumber of Glacial LakesTotal Area/km2Percentage of Number/%Percentage of Area/%
<4000 m535120.601711.2328.17
4000–4500 m1549152.135232.535.54
4500–5000 m1487115.475531.226.97
5000–5500 m118039.177824.769.15
>5500 m150.73540.310.17
Total4766428.1257100100
Table 8. Area-class distribution statistics of glacial lakes in southeastern Tibet.
Table 8. Area-class distribution statistics of glacial lakes in southeastern Tibet.
Area ClassNumber of Glacial LakesTotal Area/km2Percentage of Number/%Percentage of Area/%
<0.01 km29145.864319.181.37
0.01–0.05 km2219955.669946.1413
0.05–0.10 km274052.719815.5312.31
0.10–0.50 km2812161.386317.0437.7
>0.50 km2101152.48542.1235.62
Total4766428.1257100100
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Luo, K.; Gong, T.; Yang, S.; Liu, X.; Cidan, Y.; Hao, S.; Zheng, H.; Su, Z.; Liu, H.; Li, M. A Multimodal Remote Sensing Framework Based on an Improved YOLO Instance Segmentation Model for Automatic Glacial Lake Extraction in Southeastern Tibet. Remote Sens. 2026, 18, 2769. https://doi.org/10.3390/rs18162769

AMA Style

Luo K, Gong T, Yang S, Liu X, Cidan Y, Hao S, Zheng H, Su Z, Liu H, Li M. A Multimodal Remote Sensing Framework Based on an Improved YOLO Instance Segmentation Model for Automatic Glacial Lake Extraction in Southeastern Tibet. Remote Sensing. 2026; 18(16):2769. https://doi.org/10.3390/rs18162769

Chicago/Turabian Style

Luo, Kaipeng, Tongliang Gong, Shengtian Yang, Xiaoli Liu, Yangzong Cidan, Shouning Hao, Hao Zheng, Zexi Su, Hanwen Liu, and Mingzhu Li. 2026. "A Multimodal Remote Sensing Framework Based on an Improved YOLO Instance Segmentation Model for Automatic Glacial Lake Extraction in Southeastern Tibet" Remote Sensing 18, no. 16: 2769. https://doi.org/10.3390/rs18162769

APA Style

Luo, K., Gong, T., Yang, S., Liu, X., Cidan, Y., Hao, S., Zheng, H., Su, Z., Liu, H., & Li, M. (2026). A Multimodal Remote Sensing Framework Based on an Improved YOLO Instance Segmentation Model for Automatic Glacial Lake Extraction in Southeastern Tibet. Remote Sensing, 18(16), 2769. https://doi.org/10.3390/rs18162769

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop