Next Article in Journal
Building Footprint Extraction in High-Density Urban Areas Based on Multi-Source Remote Sensing Data Fusion and ACM-PSPNet
Previous Article in Journal
Land-Cover and Land-Use Mapping Under Limited Data Highlights Hyperparameter Stability and Predictor Design
Previous Article in Special Issue
Spatial Prompt and Wavelet Mamba-Based Multi-Scale Cross-Domain Feature Fusion Network for Segmentation of Mining-Disturbed Land
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Deep Learning for Water Body Segmentation in Remote Sensing Imagery: A Review

1
School of Surveying and Geoinformation Engineering, East China University of Technology, Nanchang 330013, China
2
Key Laboratory of Mine Environmental Monitoring and Improving around Poyang Lake of Ministry of Natural Resources, East China University of Technology, Nanchang 330013, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(17), 2972; https://doi.org/10.3390/rs18172972
Submission received: 6 July 2026 / Revised: 7 August 2026 / Accepted: 26 August 2026 / Published: 2 September 2026
(This article belongs to the Special Issue Deep Learning for Remote Sensing Image Segmentation)

Highlights

What are the main findings?
  • This review summarizes recent deep learning architectures for water body segmentation in remote sensing imagery, including U-Net, DeepLabv3+, Transformer, Mamba, and SAM-based models. A neural network is proposed for covariance-mismatch-tolerant phase-only beamforming.
  • It also reviews commonly used water body datasets and evaluation metrics, and discusses current challenges and future research directions.
What are the implications of the main findings?
  • This review provides a structured reference for selecting, comparing, and improving deep learning models for water body segmentation.
  • It offers practical guidance for dataset selection, metric evaluation, and future methodological development in remote sensing water body segmentation.

Abstract

Water body segmentation is essential for flood monitoring, water resource management, and ecological and environmental protection. Recent advances in deep learning and remote sensing have substantially improved the automation and accuracy of water body segmentation. However, existing methods remain susceptible to false positive predictions in complex landscapes containing shadows and dense vegetation. Their ability to delineate fragmented water surfaces and small water bodies is also limited, and their robustness to seasonal variations and transferability across regions require further improvement. This paper presents a comprehensive review of water body segmentation methods, datasets, evaluation metrics, current challenges, and future research directions. It examines representative models derived from U-Net, DeepLabv3+, Transformer, Mamba and the Segment Anything Model (SAM), with particular emphasis on their architectural modifications and improvement strategies. Representative datasets covering rivers, lakes, and reservoirs are summarized, and commonly used evaluation metrics are reviewed and interpreted. Future research should focus on spatiotemporal adaptive multimodal fusion, segmentation strategies for large model integration, fine-grained water body segmentation, and the integration of explainability and physical mechanisms. This review provides a useful reference for improving water body segmentation models, developing high-quality water body datasets, and supporting practical applications in water resource management and environmental monitoring.

1. Introduction

Water body segmentation is a critical task in remote sensing image interpretation. Its primary goal is to extract water pixels from remote sensing imagery and separate them from non-water land-cover features, thereby producing a water extent map with well-defined spatial boundaries. Driven by the combined effects of human activities and climate change, surface water bodies are expanding or shrinking to varying degrees, potentially leading to floods [1], disease transmission [2], and drought [3], which can severely compromise human well-being. To mitigate and prevent these impacts, researchers have adopted water body segmentation methods to address this challenge. Water body segmentation methods enable the accurate extraction of water body information and are of critical importance for applications such as water resources monitoring [4], environmental protection [5], urban planning [6], and agricultural irrigation [7].
Over the past few decades, water body segmentation has evolved through three major stages. Early efforts primarily relied on field observations and hydrographic surveys. These approaches could enhance the accuracy of water body boundary positioning; they were constrained by long survey cycles, high costs, and safety risks. Remote sensing technology effectively addresses the aforementioned shortcomings and advances water body segmentation to a new stage. Traditional water body segmentation methods based on optical imagery, synthetic aperture radar (SAR) imagery, and thermal infrared observations emerged and promoted development in this field. During this period, research using remote sensing imagery led to a wide range of approaches, including threshold-based methods [8], spectral index methods [9], support vector machines [10], decision trees [11], random forest methods [12], and object-oriented methods [13]. Despite these performance improvements, traditional remote sensing methods continue to face multiple challenges, and the issues of optimal threshold selection, high model complexity, cumbersome workflows, and high labor costs remain unresolved. By contrast, deep learning approaches use the integration of modules such as skip connections and attention mechanisms, leading to numerous variants and derivative models that more effectively exploit spectral and spatial information in remote sensing imagery. Existing studies have shown that deep learning methods achieve improved segmentation performance compared with spectral index-based and conventional machine learning approaches.
To clarify the distinctions between this review and existing studies, Table 1 systematically compares representative reviews of water body segmentation in terms of literature coverage, remote sensing data types, model categories, datasets, evaluation metrics, and emerging research directions. Previous reviews have mainly focused on convolutional neural networks and their variants, including fully convolutional networks and U-Net architectures, while some have been limited to either optical or synthetic aperture radar imagery. In contrast, this review covers not only conventional convolutional architectures but also recent models based on Transformers, Mamba, and the Segment Anything Model. It further classifies model optimization strategies according to their target components, including encoders, decoders, loss functions, and additional modules. This review also provides a systematic overview of 12 water body segmentation datasets, including their spatial resolution, scene characteristics, image quantity and extent, annotation quality, and access links. Key considerations for dataset selection are also discussed. In addition, the mathematical definitions and practical interpretations of commonly used evaluation metrics are summarized. The proposed research agenda extends beyond widely discussed topics such as multisource data fusion, model interpretability, and small water body extraction. Future studies should investigate adaptive spatiotemporal multimodal fusion by integrating data from multiple sources and developing dynamic weighting mechanisms suited to different application scenarios. Model interpretability should also be strengthened by incorporating physical mechanisms and establishing quantitative criteria for evaluating interpretability. Further attention should be given to high-resolution water body segmentation and segmentation methods based on large foundation models. These advances may enable future systems not only to delineate water bodies but also to predict their evolution, estimate water quality parameters, and identify periodic transitions among water body types. Large foundation models may further improve model generalization and zero-shot learning performance.
This review explores deep learning–based water body segmentation in remote sensing imagery from multiple perspectives, including model architectures, datasets, evaluation metrics, prevailing challenges, and future research directions. Section 2 describes methods for literature search and screening. Section 3 categorizes representative models by architectural design, introduces widely used water body segmentation frameworks, and further examines the effectiveness of common enhancement strategies for each model. Section 4 summarizes publicly available datasets for water body segmentation, describing their spatial resolution, category definitions, and access information, and then reviews commonly used evaluation metrics with detailed interpretations. Section 5 discusses current challenges and combs prospective research opportunities. Section 6 is a summary of the entire article.

2. Methods for Literature Search and Screening

The literature search was conducted using two major databases: Web of Science for English publications and the China National Knowledge Infrastructure for Chinese publications. Web of Science was used to retrieve high-quality and recent international studies, whereas the China National Knowledge Infrastructure was used to identify Chinese studies published in journals included in the Peking University Core Journal List, the Chinese Science Citation Database, and the Engineering Index. This strategy ensured comprehensive coverage of high-quality research published both internationally and in China. Given the rapid development of deep learning technologies, the search focused primarily on studies published between 2021 and 2025 to capture the latest advances and research trends in this field. A combination of subject terms and keywords was used. The English search terms included “water body segmentation,” “U-Net water body segmentation,” “DeepLabv3+ water body segmentation,” “Transformer water body segmentation,” “Mamba water body segmentation,” and “SAM water body segmentation.” The Chinese search terms included “water body segmentation,” “water body segmentation model,” and “water segmentation algorithm.”
During literature retrieval and screening, studies on core water body segmentation, closely related applications, and downstream monitoring and identification tasks were included to provide a comprehensive overview of current research. Core water body segmentation refers to studies primarily designed to generate pixel-level water masks or water and land boundaries. These studies include the segmentation of rivers, lakes, coastlines, and other open water bodies from satellite, aerial, unmanned aerial vehicle, and close-range imagery. Flood mapping studies that generate pixel level inundation boundaries were also included in this category. The extraction of aquaculture areas, the segmentation of wetlands and mangroves, and the segmentation or classification of sea ice were regarded as closely related applications. These tasks share spectral, textural, geometric, and boundary characteristics with water body segmentation and face similar challenges, including shadow interference, reflectance variation, mixed pixels, and complex boundaries. Water level monitoring and water quality identification or retrieval were classified as downstream tasks. Such studies were included when water body masking or the extraction of water and land boundaries constituted a key processing step.
To ensure the relevance and methodological reliability of the selected literature, explicit inclusion and exclusion criteria were established. Eligible publications included research articles and review papers related to deep learning-based water body segmentation and the predefined related applications. Review papers were primarily used to compare the scope and contributions of existing reviews with those of the present study, whereas research articles were examined to extract technical details and support the subsequent analysis and discussion. As illustrated in Figure 1, literature screening was conducted in two stages. During the initial screening, titles and abstracts were examined, and duplicate records, studies outside the predefined scope, studies focusing on unrelated models, and studies using only conventional segmentation methods were excluded. During the second stage, the full texts were assessed in detail. Studies were excluded if they lacked sufficient methodological or experimental information, provided incomplete descriptions of datasets or evaluation protocols, or did not primarily address water body segmentation or the predefined related applications. The magnitude or statistical significance of the reported performance gains was not used as an exclusion criterion. Following systematic retrieval and screening, 107 eligible studies were included in this review and formed the basis for the subsequent analysis and discussion.

3. Deep Learning-Based Water Body Segmentation Methods

A growing body of work has explored the potential of deep learning for this task, and new deep learning–based water body segmentation models continue to emerge. This review provides a chronological overview of representative models and summarizes major derivatives and variants, as illustrated in Figure 2.
Owing to the strong performance of fully convolutional networks (FCNs) in semantic segmentation [21], water body segmentation models based on FCN have gradually attracted increasing attention from the research community. By refining the FCN framework, researchers have proposed a range of representative models, including U-Net [22], DeepLabv3+ [23], SegNet [24], and PSPNet [25]. U-Net was originally proposed for medical image segmentation [22]; given the robust segmentation capability of the proposed model, it has since been adapted to water body segmentation, leading to numerous improved variants such as WaterNet, D-UNet, and WEU-Net. In the DeepLab series of models, DeepLabv3+ has attracted extensive attention, resulting in many structure-optimized and fusion-based variants, including GAN+DeepLabv3 and Swin-DeepLabv3+. In recent years, the introduction of Transformer architectures [26] into remote sensing water body segmentation has further advanced the field. Models such as the Swin Transformer [27] and Mamba [28] have gained prominence, inspiring variants including MSFS win and Mf-Mamba. Hybrid designs that combine Transformers with Convolutional Neural Networks (CNNs) to exploit their complementary strengths have become an active research direction, spanning embedding-, parallel-, and cascade-based fusion strategies, as exemplified by Water Former, UTGLO, and EMRT. As research continues to progress, an increasing number of new architectures are being explored for water body segmentation. In particular, the SAM [29] was introduced in 2023 and has been rapidly adopted for water body segmentation [30,31,32], leading to SAM-based variants such as CSW-SAM and CEA-SAM. Collectively, these developments highlight the ongoing evolution and diversification of water body segmentation models.
Deep learning models for water body segmentation can be classified from several perspectives, including computational mechanisms, network architectures, data modalities, supervision strategies, and application scenarios. For example, convolutional neural networks, Transformers, and state space models are primarily distinguished by their underlying modeling mechanisms, whereas U-Net and DeepLabv3+ are characterized mainly by their network structures and functional modules. However, a single model may exhibit several attributes simultaneously, such as a U-shaped encoder and decoder structure, a hybrid architecture combining convolutional neural networks and Transformers, and multimodal inputs. Therefore, a single mutually exclusive classification scheme cannot adequately describe the diversity of existing water body segmentation methods. To establish a more systematic classification framework, this review organizes the models according to their primary architectural categories and supplements them with descriptive attribute labels, as presented in Table 2.
Based on their primary modeling mechanisms, existing methods are divided into five categories: convolutional neural network models, Transformer models, hybrid architectures combining convolutional neural networks and Transformers, Mamba models, and Segment Anything Model-based models. Both U-Net and DeepLabv3+ use convolution as their primary feature extraction mechanism and are therefore classified as convolutional neural network models. However, because of their distinctive structures and optimization strategies, they are discussed separately in detail. Transformer models primarily use self-attention to capture global contextual information, whereas hybrid architectures combine convolution and self-attention to model local and global features. Models that integrate convolutional neural networks and Transformers within a U-shaped architecture are classified primarily as hybrid models. Mamba models use selective state space mechanisms and sequence scanning operations to capture long-range dependencies, which differs fundamentally from the self-attention mechanism used by Transformers. The Segment Anything Model and its adaptations for remote sensing rely on large scale pretraining, prompt guided segmentation, and transfer across tasks. They are therefore classified as foundation models rather than conventional convolutional neural network or Transformer models.
In addition to the primary architectural categories, the reviewed studies are further classified according to data modality, supervision strategy, and application scenario. Data modalities include optical, multispectral, hyperspectral, synthetic aperture radar, aerial or unmanned aerial vehicle, close range, multimodal, and multi-temporal imagery. Supervision and adaptation strategies include fully supervised, semi-supervised, weakly supervised, unsupervised, self-supervised pretraining, domain adaptation, and prompt-based few-shot or zero-shot adaptation. Application scenarios are divided into core water body segmentation, closely related aquatic applications, and downstream tasks supported by segmentation. Further details are provided in Table 3.

3.1. Water Body Segmentation Models Based on the U-Net Model

In 2015, Ronneberger et al. [22] proposed U-Net, a symmetric encoder–decoder FCN composed of an encoder, a decoder, and skip connections. As shown in Figure 3, the encoder and decoder form the left and right branches of the U-shaped architecture, and the skip connections constitute the key innovation of the model. This symmetric design allows the network to capture contextual information while recovering fine spatial details. Skip connections fuse shallow high-resolution features from the encoder with deep semantic features from the decoder. With limited annotated data, U-Net improves segmentation performance by using data augmentation with random elastic deformations. Owing to its strong segmentation capability, U-Net has been increasingly adopted for water body segmentation in remote sensing imagery, with successful applications such as shoreline water extraction and flood inundation mapping [33,34]. Because terrestrial backgrounds in remote sensing scenes are generally more complex than those in medical images, researchers have introduced a series of improvements to the model components, as detailed in the following sections.

3.1.1. Improvements to Encoder Structure

The encoder, also referred to as the contracting path, consists of four modules, each containing two successive 3 × 3 convolutional layers followed by a rectified linear unit activation function and a 2 × 2 max pooling layer. For water body segmentation, the encoder uses CNN architecture to extract image features while reducing spatial resolution. However, remote sensing imagery typically contains complex backgrounds, whereas the original U-Net encoder is relatively lightweight. To better address these problems, researchers have proposed a range of encoder-focused improvements, through the integration of attention mechanisms, refinement of convolutional operations, and redesign or replacement of the backbone network:
  • Introduce an attention mechanism: To address unstable spectral reflectance- and reflection-induced interference in multispectral water body segmentation, Hu et al. [35] embedded a channel attention mechanism in each encoder stage. Compared with the standard U-Net, this approach strengthens channel-wise feature representation. The spatial branches not only extract shapes and boundaries but also adaptively recalibrate responses across spectral bands. Wang et al. [36] integrated a Transformer-based multi-head attention module after the U-Net encoder, enabling the joint modeling of image features across multiple subspaces and improving the model ability to identify flood-related characteristics. The AER U-Net model proposed by Jonnala et al. [37] reduces skip connection complexity by introducing a self-attention module. For water body segmentation, this architecture enables effective discrimination between water and non-water regions.
  • Optimize the convolutional layer: Zhang et al. [38] reported that D-UNet has a smaller receptive field. By replacing standard convolutions with multiscale dilated convolution modules, they enhanced feature extraction and segmentation accuracy, reducing missed detections and false positives. Many deep learning networks for water extraction rely on the translation invariance of convolution kernels, Xu et al. [39] introduced rotation-invariant convolutions within a U-Net framework to refine the architecture and enrich internal representations, enabling the robust recognition of water bodies under image rotation. To better model complex channel relationships in RGB imagery, Wang et al. [40] replaced conventional convolutional blocks with quaternion convolutions, which learn optimal weighting across the RGB channels and improve multiscale water body recognition.
  • Changes in the backbone network: Wagner et al. [41] used ResNeXt 50 as the backbone of a U-Net model and compared 32 CNN architectures for water body segmentation. The results indicate that the improved model achieved the best segmentation performance and can be used for the real-time monitoring of river water levels. To improve water body segmentation accuracy, Liu et al. [42] replaced the VGG module in U-Net with MobileNetV2 and leveraged the inverted residual structure and the linear bottleneck layer of MobileNetV2 to efficiently extract spatial features and improve water body segmentation performance. Xia et al. [43] further lightweighted U-Net by reducing the number of convolutional kernels per layer from 64–1024 to 32–512, decreasing the parameter count to approximately one quarter of the original while maintaining essentially unchanged segmentation accuracy.
  • Other improvements: Li et al. [44] proposed the GLF-MFUNet model for Sentinel-2 satellite imagery. The dual-path encoder extracts global and local features in parallel, which helps to capture local details and global contextual information and achieves strong performance in small-target water body segmentation. Cai et al. [45] improved the U-Net architecture by incorporating an image upsampling module to form a symmetric structure and further deepen the network via an “S”-shaped loop design. Their results indicate improved water body segmentation accuracy under complex backgrounds.

3.1.2. Improving Skip Connections

In the U-Net architecture, the encoder applies two convolution operations to generate feature maps before reducing their spatial resolution. In the decoder, transposed convolution increases the spatial dimensions of the feature maps, which are then concatenated with the corresponding encoder features. Skip connections transfer feature information from each encoder stage to the corresponding decoder stage. However, important information may be lost during feature fusion. To address this limitation, many studies have incorporated attention mechanisms to enhance feature representation and improve segmentation performance.
Bai et al. [46] proposed TA-UNet3+, which embeds a window-attention module within the skip connections to enrich the semantics of shallow features and strengthen the model’s comprehension capability. In the semantic segmentation of water bodies, the method can accurately detect changes in water body morphology. Xing et al. [47] introduced spatial- and channel-attention modules into U-Net skip connections and developed the MAFU-Net model, which alleviates insufficient local feature extraction, reduces interference from irrelevant features, and improves segmentation accuracy. Zhang et al. [48] incorporated an improved coordinate attention mechanism into the skip connections and fused high-level and low-level feature maps to extract the boundaries of freshwater aquaculture areas more accurately, while also reducing boundary adhesion.

3.1.3. Decoder Structure Optimization

In the decoder, a single 2 × 2 transposed convolution is used to upsample the feature maps and recover fine spatial details. The feature maps forwarded from the encoder are then cropped and concatenated, and the fused feature maps are passed to subsequent convolutional layer. A 1 × 1 convolution projects the features to the required number of target classes. Segmentation performance remains limited under complex scenarios, such as scenes with small water bodies and strong noise. Structural refinements and additional modules have been introduced to alleviate these issues, but they increase the computational complexity of the model.
Xiang et al. [49] introduced a spatial and channel squeezing activation module in the decoder, which suppresses background noise and emphasizes small water bodies during the water body segmentation of aerial imagery. Hertel et al. [50] developed two probabilistic CNNs for water body segmentation: a Bayesian CNN and a Monte Carlo Dropout network. The former replaces standard convolution layers with Bayesian convolutional layers, whereas the latter appends spatial dropout at the end of each decoder layer. Their experiments show that the Bayesian CNN provides more reliable uncertainty quantification. To alleviate the empirical threshold dependence of conventional methods on high-resolution imagery, Li et al. [51] introduced densely connected blocks in the decoder to strengthen spectral and spatial feature extraction. They also replaced batch normalization with group normalization within these blocks to improve training stability under small mini batches. Ye et al. [52] substituted conventional convolutions with depthwise separable convolutions, reducing parameters and computational cost while improving segmentation efficiency.

3.1.4. Summary of Improvements Based on the U-Net Model

The U-Net model, with its ingenious design, incorporates skip connections and a unique symmetrical structure to capture multiscale contextual information and recover fine-grained details. This design is suited to pixel-level water body segmentation and has made U-Net one of the most widely used models in this field. For many water body segmentation applications, the U-Net model and its derivatives remain reliable choices. Researchers have made improvements to the U-Net model based on image type and accuracy requirements. As summarized in Table 4, enhanced U-Net variants have been applied to flood disaster monitoring, water-level surveillance, and freshwater aquaculture zone segmentation. Modifications to the encoder, skip connections, and decoder, such as attention modules and convolutional operations, can strengthen feature extraction and improve segmentation performance. Nevertheless, repeated downsampling may remove fine details and degrade the segmentation of small water bodies; convolution and pooling can constrain the effective receptive field, which may produce discontinuities or holes when segmenting extensive floodplains; skip connections may also propagate background noise, causing shadows and man-made structures to be misclassified as water.

3.2. Water Body Segmentation Models Based on the DeepLabv3+ Model

In 2018, the Google team extended DeepLabv3 and proposed DeepLabv3+. The model adopts an asymmetric, enhanced encoder–decoder architecture and consists mainly of an encoder and a decoder. As shown in Figure 4, the encoder uses the backbone network and dilated convolution to extract high-level semantic features. The atrous spatial pyramid pooling module then aggregates contextual information at multiple scales using different dilation rates. In the decoder, the high-level features generated by this module are upsampled and concatenated with low-level features extracted from the shallow layers of the backbone network and processed through 1 × 1 convolution. Subsequent convolution operations integrate semantic information with spatial details, refine object boundaries, and restore the prediction map to the spatial resolution of the input image. As a classic image segmentation model, it has been increasingly applied to black and odorous water detection, flood disaster assessment, and water resources management [53,54,55,56]. To further improve segmentation accuracy, many studies have refined both the encoder and the decoder, thereby achieving better segmentation performance.

3.2.1. Improvements to Encoder Structure

In the DeepLabv3+ encoder, depthwise separable convolution reduces the computational cost of convolution operations; the ASPP module aggregates multiscale features, which helps to reduce missed detections of small targets. Xception [57] and ResNet [58] are commonly used backbone networks in DeepLabv3+ models. They increase network depth while helping to reduce overfitting. Despite its strong segmentation accuracy, DeepLabv3+ often requires a relatively long training time. These defects can be mitigated by refining the ASPP design and adopting more computationally efficient backbone architecture:
  • Optimize the ASPP module: Luo et al. [59] optimized the ASPP module in DeepLabv3+. Compared with the square pooling kernels used in the original model, the introduced strip pooling kernels are better suited for extracting scattered distribution water bodies over long ranges. Zhang et al. [60] introduced two modifications to the ASPP module in the encoder. The dilation rates were adjusted from 6, 12, and 18 to 2, 4, 8, and 16; this adjustment is more advantageous for capturing feature information from elongated rivers. The standard convolution in the ASPP module was replaced with depthwise separable convolution, which reduces the number of model parameters and improves training efficiency. During water body segmentation, when the river surface and its surrounding environment change, the ASPP module may not adapt well to such target variation. Therefore, Sun et al. [61] restructured the ASPP module as a densely connected DASPP module. In this design, the dilation rate is increased progressively, and dense connections allow upper-layer atrous convolutions to reuse the outputs of lower-layer atrous convolutions. This strategy expands the receptive field and enables multiscale feature extraction. Compared with the original ASPP module, the DASPP module improves water body segmentation accuracy.
  • Changes in the backbone network: Huang et al. [62] employed an enhanced DeepLabv3+ model to segment black and odorous water in Gaofen-2 remote sensing imagery. Compared to DeepLabv3+, they replaced the original backbone network with MobileNetV2. This modification reduced the model size from 226 MB to 23.3 MB. Xue et al. [63] also adopted MobileNetV2 as the backbone network for urban waterlogging monitoring. Zhang et al. [64] employed ResNet101 as the backbone network of DeepLabv3+ to segment duckweed-type black and odorous water. This design mitigated vanishing and exploding gradients while enhancing model generalization. Chen et al. [65] incorporated Swin Transformer as an additional backbone network in DeepLabv3+. In this configuration, the same remote sensing image is fed into two backbone networks to generate feature maps. The output features from the two parallel backbones are then upsampled and fused to enable the recognition and segmentation of specular water bodies on the land surface.

3.2.2. Decoder Structure Optimization

The decoder concatenates the deep feature maps after 4× upsampling with the shallow feature maps processed by a 1 × 1 convolution. A 3 × 3 convolution and a 4× upsampling are then applied to recover spatial feature information and produce the water body segmentation feature map. However, the decoder depends strongly on the quality of the encoder output, and the feature fusion strategy remains relatively simple. In future water body segmentation studies, accuracy may be improved by introducing attention mechanisms or modifying the upsampling module.
Chen et al. [66] incorporated an attention modulation module into the decoder to adaptively weight low-level features. The reweighted features were then fused with high-level features from the encoder to obtain more accurate boundary information for water bodies, which was used for water body segmentation in multisource SAR imagery. The DeepLabv3+ decoder includes only a single skip connection; its ability to recover fine details is limited. To address this issue, Lv et al. [67] introduced three upsampling modules in the decoder to fuse multilevel encoder features, which improved segmentation performance.

3.2.3. Summary of Improvements Based on the DeepLabv3+ Model

Compared with U-Net, DeepLabv3+ strengthens multiscale feature extraction, enhances contextual information capture, and enlarges the receptive field. This design helps to reduce detail loss caused by excessive downsampling and pooling.
However, it may also lead to insufficient preservation of local details, which is a fundamental reason why the resulting segmentation boundaries tend to become coarse and blurred. The encoder architecture is relatively simple, which results in insufficient capability for processing high-resolution imagery and degrades the segmentation performance for small targets and elongated water bodies. To address these issues, researchers have proposed a series of targeted improvements. To date, strategies such as backbone replacement, the integration of attention mechanisms, and modifications to the ASPP module have improved both efficiency and accuracy, thereby broadening the application scope of DeepLabv3+, as shown in Table 5. Beyond encoder and decoder modifications, segmentation performance can also be enhanced through loss function redesign or by combining multiple loss functions; these approaches are not discussed further in this paper. These improvements can enhance water body segmentation accuracy: the models remain convolutional neural network–based and therefore retain the inherent limitations of CNN in modeling long-range dependencies.

3.3. Water Body Segmentation Models Based on Transformer Models

In 2017, Vaswani et al. proposed the Transformer, a deep learning architecture based entirely on attention mechanisms, as shown in Figure 5. The Transformer consists of an encoder and a decoder. The encoder comprises multi-head attention layers and feedforward layers, whereas the decoder additionally includes a masked multi-head attention layer. Multi-head attention assigns different weights to information at different positions, which improves the ability to capture subspace features. Because of its strong segmentation performance, the Transformer has gradually become a mainstream choice for image segmentation. In 2021, the Swin Transformer was introduced as a further development. In 2024, the Mamba model was proposed to address bottleneck issues in Transformer-based architecture. Meanwhile, researchers have leveraged the complementary strengths of Transformers and CNN and applied them to water body segmentation through a range of integration strategies.
The Swin Transformer comprises four stages, as indicated by the dashed boxes in Figure 6. Stage 1 includes a linear embedding module and a Swin Transformer block module, whereas the other three stages each consist of a patch merging module followed by a Swin Transformer block module. The input image first undergoes linear patch partitioning and embedding in Stage 1. In each subsequent stage, patch merging reduces spatial resolution and increases the number of channels, and a series of Swin Transformer blocks performs feature transformation. The Swin Transformer block module is repeated twice in Stage 2, six times in Stage 3, and twice in Stage 4, yielding the final feature map.
As shown in Figure 7, the Swin Transformer block is the core module of the model. It replaces the original multi-head self-attention with window-based multi-head self-attention and shifted window multi-head self-attention. Consider a window size of 4 × 4 and a remote sensing image of size H × W. After the image is input to the Swin Transformer block, window-based multi-head self-attention partitions it into H/4 × W/4 windows of size 4 × 4. The window features are then processed by a multilayer perceptron. Shifted window multi-head self-attention is applied to enable information exchange across windows. The output is then produced by a multilayer perceptron. Layer normalization is inserted between these components to improve model robustness.
The Swin Transformer introduces two important improvements, namely, a hierarchical feature map representation and a window attention mechanism, which reduces computational complexity and improves model performance. The issues of limited global receptive fields and increased susceptibility to overfitting, compared with CNN models, remain unresolved. These limitations warrant further investigation.
Wang et al. [68] improved the Swin Transformer model based on an analysis of existing water body segmentation methods. They introduced an ASPP module to enhance multiscale feature perception. They also refined the decoder to fuse shallow and deep features from the encoder. The improved model improves model robustness for water body segmentation in complex environments. Ma et al. [69] added an FCN as an auxiliary encoder to form a dual-branch encoder architecture. To prevent the Swin Transformer from overlooking multispectral information, they developed a prediction map integration module that combines the outputs of the Swin Transformer and the Normalized Difference Water Index using a Bayesian averaging strategy. The proposed model outperformed other segmentation models and is suitable for hydrological research. Sudakow et al. [70] replaced the convolutional layers in U-Net with Swin Transformer modules and introduced a cross-channel attention mechanism in the decoder to classify different aquatic surface types, including sea ice, water, and snow.

3.4. Water Segmentation Based on a Hybrid Model Combining Transformers and CNN

At present, researchers have leveraged the complementary strengths of Transformers and CNN and applied them to water body segmentation through a range of integration strategies. Existing studies can be grouped into four categories: embedding a Transformer within a CNN, using a parallel hybrid architecture, combining CNN feature extraction with Transformer-based feature enhancement, and adopting a Transformer encoder with a CNN decoder. These strategies offer several advantages, including multiscale feature extraction, fusion of local and global representations, and improved computational efficiency.

3.4.1. Embedding Transformers into CNN

Zhang et al. [71] reported that the automatic classification of sea ice and open water is often degraded by noise and insufficient multiscale representation. To address these issues, they introduced two Transformer-based attention modules within a hybrid framework. A multiscale spatial attention module was inserted at the bottleneck to improve the recognition of sea ice across multiple scales. Zhang et al. [72] embedded the MixFormer module into a U-Net architecture and combined CNN with MixFormer to capture rich water body details.

3.4.2. Parallel Hybrid Structure

Zhao et al. [73] proposed SPT-UNet, which integrates a CNN and a Transformer. The model fuses pixel-level features extracted by the CNN with superpixel-level features extracted by the Transformer. It achieved favorable results in both quantitative metrics and visual effects. Xiao et al. [74] designed a dual-branch parallel encoder. A ResNet 50 backbone was used to extract local features, and a deformable self-attention mechanism was introduced in the Transformer branch. The improved model achieved good segmentation performance across multiple scales and under complex backgrounds. Kang et al. [75] proposed a coupled CNN and Transformer model named WaterFormer. The dual-stream network consists of two asymmetric branches. The spatial branch captures low-order features and spatial detail information, whereas the context branch captures global contextual information. This design effectively addresses multiscale variation and boundary blurring in water body segmentation.

3.4.3. Transformer Feature Enhancement—CNN Feature Extraction

Tian et al. [76] used a CNN as a feature extractor with five downsampling modules. Detail information and semantic information were extracted by C2-blocks and C5-blocks, respectively. The features from the C2-block were used as queries, whereas the features from the C5-block were used as keys and values, which strengthens interactions between local and global features. A Vision Transformer was then used for feature enhancement. Zhang et al. [77] proposed a hybrid network for water body segmentation in high-resolution remote sensing imagery. The model combines the ability of a CNN to capture local detail features with the ability of a Transformer to model global contextual semantic relationships over large areas. Experiments on the GID dataset showed that the extracted water boundaries have high accuracy and continuity. Zhong et al. [78] designed NT-Net, a semantic segmentation network that integrates a Transformer and a CNN. The encoder uses downsampling modules to extract multiscale feature maps (D1–D5). A multistage Transformer module then enhances these feature maps, which alleviates the oversegmentation of non-lake targets. To enable accurate and efficient water body segmentation from remote sensing imagery, Yang et al. [79] proposed WatNet by combining the strengths of a CNN and a Transformer. In the encoder, ResNet50 was used to extract multiscale features. In the decoder, a global multi-attention fusion module and a water forward network module was introduced to capture contextual information. An edge focused attention module was also incorporated to mitigate the semantic gap between features.

3.4.4. Transformer Encoder—CNN Decoder

Wang et al. [80] proposed MHNet, which adopts a multiscale encoder–decoder architecture. Compared with Swin Transformer and Vision Transformer models, MHNet introduces a Restormer module to fuse local and global features more effectively and improve boundaries in water body segmentation. The decoder also incorporates resizing and pixel reassembly modules, which support a lightweight CNN encoder. Fan et al. [81] addressed interclass misclassification and intraclass discontinuity by using twelve Transformer blocks in the encoder for feature extraction and applying self-distillation for pretraining. In addition, a convolutional neural network-based decoder was used to fuse multilevel features. After these refinements, the model was suitable for the semantic segmentation of marine aquaculture areas. Fan et al. [82] designed an unsupervised UTGLO model. The encoder comprises a dual-network system, SST Teacher and SegT, and uses an iterative update mechanism to improve pseudo-label quality. The decoder uses UPerHead to process the feature maps produced by the encoder. This model was applied to marine aquaculture segmentation.

3.4.5. Summary Table of Improvements in Hybrid Transformer and CNN Models

The Swin Transformer innovatively introduces a window attention mechanism, but this design also increases the number of model parameters. Integrating convolutional neural networks and Transformers to exploit their complementary strengths may lead to vanishing gradients, while fusion across different model types increases architectural complexity. To address these challenges, researchers have investigated more effective strategies for combining convolutional neural networks with Transformers. By introducing specialized modules and developing new algorithms and network architectures, recent studies have achieved accurate water body segmentation across diverse remote sensing data sources and scene types, as summarized in Table 6. Nevertheless, Transformers remain limited in their ability to capture fine local details. In the presence of noise, they may lose high-frequency information and struggle to balance local spatial details with global semantic context. Compared with convolutional neural networks, Transformers also involve extensive multiplicative operations across sequence positions, resulting in higher computational complexity. Moreover, fully supervised learning remains the dominant paradigm in water body segmentation, and existing methods continue to rely heavily on large scale annotated datasets.

3.5. Water Body Segmentation in the Mamba Model

The self-attention mechanism used in Transformer models is effective in capturing global context and long-range dependencies. However, the computational cost of standard self-attention increases quadratically with input sequence length, which limits its efficiency when processing long sequences or high-resolution images. Gu et al. proposed Mamba, a model based on a selective state space mechanism that captures input dependent information and improves the efficiency of long sequence processing through hardware aware parallel algorithms. Because Mamba does not rely primarily on self-attention and its computational complexity increases linearly with sequence length, it should be classified as a state space model rather than as a Transformer or a Transformer variant. For certain long-sequence modeling tasks and model scales, Mamba can achieve efficiency and performance comparable to or better than those of Transformer models. However, its relative advantages remain dependent on the specific task, model architecture, and experimental conditions. As illustrated in Figure 8, Mamba combines the H3 component of a state space model with a gated multilayer perceptron to form a Mamba block, thereby simplifying conventional deep-sequence model architectures. Subsequent studies have gradually extended its application to water body segmentation. At present, excessive dependence on predefined spatial scanning patterns remains a potential limitation of the Mamba model. Because two-dimensional features must be reordered along horizontal, vertical, bidirectional, or cross-shaped paths, the scanning direction and traversal sequence may affect state propagation. A limited range of scanning paths may also hinder the preservation of the original topological relationships between adjacent elements in two-dimensional space. This limitation is particularly evident when segmenting meandering rivers and elongated water bodies, where directional sensitivity, local discontinuities, and excessive boundary smoothing may occur.
Li et al. [83] proposed the MBSSNet model based on SS2D for the semantic segmentation of optical imagery and SAR imagery. By leveraging its ability to capture contextual information, the model can better distinguish water bodies from other small land-cover targets and more accurate water boundaries. Yang et al. [84] proposed WaterMamba for water body segmentation in complex remote sensing environments. The model integrates multiple balanced-channel visual Mamba modules with Mamba boundary attention modules to suppress interference from complex backgrounds and extract more comprehensive water body features. Song et al. [85] designed a lightweight CNN and state space model architecture, evaluating it on the WorldFloods dataset. The model demonstrated favorable performance, enabling efficient and accurate floodplain segmentation.

3.6. Water Body Segmentation Based on the SAM Model

Unlike U-Net, DeepLabv3+, and Transformer models, the SAM is a large foundation model with a prompting function that achieves interactive segmentation depending on the prompt type. The SAM was pretrained on the SA-1B data set, which contains eleven million images. This pretraining improves generalization and enables segmentation across diverse scenes, making it a general-purpose image segmentation model. As shown in Figure 9, the architecture consists of an image encoder, a prompt encoder, and a mask decoder. The image encoder converts an input image into a high-dimensional and dense feature representation for subsequent processing. The prompt encoder processes both sparse and dense prompts provided by the user. Sparse prompts include foreground points, background points, and bounding boxes, whereas dense prompts are represented by input masks. The encoded prompt features are then combined with the image features and passed to the mask decoder to generate a segmentation mask for the target region. Although the original study also reported preliminary experiments using free text prompts, the publicly released standard SAM model does not natively support text as a conventional input to the prompt encoder. The mask decoder fuses image features with prompt embeddings to produce the segmentation output. Given its strong zero-shot segmentation capability, the SAM has attracted attention for water body segmentation. Although the SAM offers substantial advantages for interactive water body segmentation and assisted annotation, its general object mask generation mechanism does not fully align with the semantic characteristics of water bodies. In areas with weak textures, low shoreline contrast, or gradual transitions between land and water, such as shallow water, wetlands, and turbid water boundaries, the model may fail to generate stable contour responses. This limitation can result in incomplete segmentation, fragmented boundaries, and discontinuous masks. Moreover, specular reflection from the water surface and variations in water depth and turbidity can produce substantial reflectance differences within the same class, causing a single water body to be divided into multiple separate regions.
Moghimi et al. [86] first applied the SAM to water body segmentation in near-range remote sensing imagery. Compared with the other evaluated models, SAM achieved superior segmentation performance after partial fine-tuning, with the image encoder kept frozen and only the designated trainable components optimized. Ma et al. [87] introduced a target consistency loss and a boundary preservation loss for semantic segmentation. These loss functions refined the original SAM outputs. Experiments on public data sets showed that this strategy improved water body segmentation accuracy. Building on SAM2, Zhang et al. [88] proposed a cross-scale water body segmentation model, termed CSW-SAM. A key contribution is an optimized lightweight encoder that improves generalization across resolutions. An automatic clustering layer was also designed to improve segmentation accuracy. Qiao et al. [89] fused boundary-accurate templates from the SAM with templates that contain semantic information generated by semantic segmentation models such as ABCNet, MANet, and FT-UNetFormer. This fusion produces high-quality templates. The method is simple and efficient, and it is suitable for water body segmentation in remote sensing imagery. Zou et al. [90] combined the SAM with a CNN–Transformer hybrid architecture and proposed SAM-CTMapper. The model integrates multi-head scale-aware convolutional layers, a superpixel self-attention layer, and a dynamic selection module. This design enables the accurate classification of coastal wetlands. To support mangrove wetland conservation, Zhang et al. [91] proposed a SAM-based approach that incorporates a wetland prompt module, an adaptation module, and a low-order adaptive module to emphasize mangrove wetland regions. A parallel CNN branch was also introduced for feature extraction, providing a new strategy for wetland protection.
When the color and texture of water bodies vary because of environmental influences, SAM may not reliably distinguish water from land. This problem also arises when water targets have textures similar to the background or when labeled data are insufficient. As summarized in Table 7, researchers have introduced new modules and integrated SAM with other models to address these issues. These strategies aim to reduce boundary blurring, alleviate limited annotations, and mitigate fragmented segmentation results. SAM is a recent model; it has been rapidly adopted for water body segmentation, and further refinement remains necessary. Current challenges involve data acquisition, prompt mechanisms, and model performance. Dependence on high-quality annotations increases data collection costs and can limit training, particularly in emerging domains where labeled samples are scarce. Changes in application scenarios also impose stricter requirements on prompt design, yet developing prompt mechanisms that are flexible and diverse remains challenging. SAM supports general image segmentation, but its adaptability to remote sensing imagery remains limited. The model also requires fine-tuning in many cases, and cross-domain generalization should be further improved.

3.7. Summary of Water Body Segmentation Models

Current water body segmentation studies primarily use models such as U-Net, DeepLabv3+, Transformer, hybrid architectures combining convolutional neural networks and Transformers, Mamba, and SAM. In addition to network architecture and enhancement modules, segmentation performance is influenced by application scenarios, computational requirements, generalization across regions, and practical deployment conditions. Table 8 presents a comparison of these models. However, the computational requirements and generalization characteristics reported in the table should not be interpreted as a definitive performance ranking. The actual computational cost of a model depends on factors such as the backbone network, number of parameters, image resolution, and hardware configuration. Similarly, generalization across domains is strongly affected by the data used for pretraining, sample diversity, data augmentation, domain adaptation strategies, and external evaluation conditions.
In addition to model architecture, data leakage, annotation uncertainty, and external validation strategies can substantially affect water body segmentation results. Remote sensing images often exhibit strong spatial and temporal correlations. When overlapping image patches are first generated from the same scene or geographic area and subsequently assigned at random to the training, validation, and test sets, adjacent patches may contain identical or highly similar textures, shorelines, and background features. This practice can result in spatial data leakage and lead to an overestimation of model performance in previously unseen regions. Water body annotations are also affected by considerable uncertainty. Tidal variation, differences between high-water and low-water periods, mixed pixels, shadows, turbid water, floating vegetation, and indistinct shorelines may cause annotators to delineate water boundaries inconsistently. Some studies currently rely primarily on random partitioning within a single dataset for model evaluation. Although this design reflects the ability of a model to interpolate among samples drawn from the same distribution, it does not adequately demonstrate transferability to unfamiliar environments. Table 9 summarizes the potential biases associated with these issues and provides corresponding recommendations.
To address the limited systematic linkage between specific challenges and corresponding improvement strategies in water body segmentation research, this review further classifies existing methods according to the problems they are designed to solve. As shown in Table 10, the reviewed studies are compared in terms of target problems, typical application scenarios, improvement strategies, applicable data types, and results. A single strategy may address several problems, whereas a specific problem often requires the combined use of multiple strategies. Therefore, the relationships presented in the table are not mutually exclusive.
Overall, no single strategy is optimal for all application scenarios. The segmentation of small water bodies depends largely on preserving fine spatial details and mitigating class imbalance. Narrow and elongated water bodies require stronger modeling of long-range dependencies and topological continuity, whereas complex backgrounds demand improved discrimination between water bodies and visually similar land features. Multimodal tasks require accurate alignment across sensors and effective feature fusion, while applications across domains should place greater emphasis on domain shifts and external validation. The summary presented in Table 10 provides a reference for selecting appropriate model improvement strategies according to specific data conditions, target characteristics, and application requirements.

4. Water Body Segmentation Dataset and Evaluation Metrics

Water body segmentation has been widely used in both domestic and international research. Major applications include water resource inventory, ecological and environmental monitoring, agriculture and aquaculture, and flood hazard assessment, as illustrated in Figure 10. To better capture current and future dynamics in these domains, model-based approaches alone are insufficient. A wide range of data sets and evaluation metrics are also required. High-quality data sets are essential for model training, whereas appropriate evaluation metrics support objective assessment of model performance. Together, they provide complementary support for addressing a broad spectrum of surface water-related problems.

4.1. Water Body Segmentation Dataset

In recent years, both the quality and availability of water body segmentation datasets have improved substantially. This progress has been driven by the increasing availability of openly accessible remote sensing data, continuous advances in surface water segmentation methods, the development of large platforms such as Google Earth Engine, greater scientific collaboration, and growing attention to sustainable water resource management. Based on a comprehensive review of the literature, this study summarizes several commonly used datasets to support dataset selection in future research, as presented in Table 11. The table provides a systematic overview of their names, main categories, sensor types, data modalities, spatial resolutions, annotation quality, geographic and temporal coverage, and sample sizes.
According to their data characteristics and intended uses, the datasets are divided into three categories: general land-cover segmentation datasets, including GID, EvLab SS, DeepGlobe Land, WHU OPT SAR, LoveDA, and SEN12MS; water body mapping and lake datasets, including China Lake, GLAKES, and GLAD GSWD; and datasets designed specifically for water body segmentation, including CWaC, GLH Water, and S1S2 Water. General land-cover segmentation datasets contain multiple scene classes, among which water is one category, and are therefore suitable for training and evaluating models for multiple class segmentation. Water body mapping and lake datasets primarily consist of large-scale mapping products or vector catalogs and generally do not provide the original image tiles required for model training. These datasets are mainly used to analyze water body changes across large spatial and temporal scales or to provide reference data for large scale evaluation. By contrast, datasets specifically developed for water body segmentation provide paired images and masks and can be directly used for end-to-end training of water body extraction models.
In addition to dataset size and spatial and temporal coverage, annotation reliability, data partitioning strategies, and licensing terms influence dataset suitability and the credibility of model evaluation, as detailed in Table 12.
Current water body segmentation datasets differ considerably in sensor type, spatial resolution, and scene composition. Dataset selection should therefore account for the intended application and specific research objectives. LoveDA, DeepGlobe Land, and GLH Water contain diverse scene categories and provide imagery with high spatial resolution, allowing water boundaries and fine spatial details to be effectively preserved. These characteristics facilitate the accurate extraction of complex water body boundaries. When the influence of weather conditions on optical imagery must be reduced, WHU OPT SAR is a suitable choice because it contains both optical and synthetic aperture radar imagery. This dataset supports the fusion of complementary features from optical and radar data and can improve segmentation robustness under adverse weather conditions. China Lake and GLAKES provide extensive geographic coverage but mainly have a spatial resolution of 30 m, making them more suitable for mapping large surface water bodies, such as lakes and reservoirs. To evaluate model generalization across domains, LoveDA can be used for training and validation, whereas DeepGlobe Land or GLH Water can serve as an independent test dataset. DeepGlobe Land is particularly suitable for assessing classification stability in areas where grassland, forest, and water classes are spatially adjacent. In contrast, GLH Water enables a more detailed evaluation of model performance in delineating water boundaries and extracting narrow rivers and small water bodies.

4.2. Evaluation Metrics

To assess water body segmentation performance and to validate model robustness and effectiveness, a set of evaluation metrics is typically used for comparative analysis. This study summarizes commonly used metrics and groups them by their properties and intended uses, as reported in Table 13. The table contains the confusion matrix and evaluation metrics. The relative importance of these metrics depends on the application scenario.
The confusion matrix comprises four fundamental elements—true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). They express the relationship between predicted outcomes and actual conditions, serving as the foundation for all core metrics. Precision quantifies the reliability of water predictions. It is defined as the proportion of true water pixels among all pixels predicted as water. Low precision typically indicates that non-water regions are misclassified as water and are often accompanied by a higher false alarm rate. Pixel accuracy is the proportion of correctly classified pixels across the entire image. However, it can be difficult to interpret under severe class imbalance between water and non-water. For binary single-label semantic segmentation, pixel accuracy and overall pixel accuracy are calculated using the same formula. Mean pixel accuracy is the average of pixel accuracy computed for each class. Recall is defined as the proportion of true water pixels that are correctly identified by the model. It decreases when omission errors are frequent. Intersection over union (IoU) is defined as the ratio of the overlap area between the predicted water region and the actual water region to the area of their union. It is one of the metrics used in this field. When multiple scene classes are considered, such as water, vegetation, and cropland, mIoU is commonly reported. It is computed by calculating IoU for each class and then averaging the class-wise values to provide an overall assessment of model performance. The Dice coefficient measures the overlap between the predicted and actual segmentations. The F1-score is also widely used, particularly for class imbalanced data sets, and is defined as the harmonic mean of precision and recall. These core metrics range from 0 to 1, and larger values generally indicate better segmentation performance. It is important to note that, for binary classification, the Dice coefficient and F1 score are mathematically equivalent. The false negative rate represents the proportion of actual water pixels that are incorrectly classified as non-water, whereas the false positive rate represents the proportion of actual non-water pixels that are incorrectly classified as water. In water body segmentation, the error rate quantifies the overall proportion of misclassified pixels. The error rate is applicable to binary single-label water body segmentation when the model output has been converted into a binary prediction mask. It measures the proportion of misclassified pixels among all pixels and is mathematically equal to 1-PA. These three metrics range from 0 to 1, with lower values generally indicating better segmentation performance. The theoretical range of the Kappa coefficient is from −1 to 1, although values observed in practice usually fall between 0 and 1. In the Kappa formula, p o denotes the observed agreement between the predicted and reference labels, whereas pe denotes the agreement expected by chance based on the marginal distributions of the predicted and reference classes. A higher Kappa value indicates stronger agreement between the predicted and reference labels, whereas a value approaching −1 indicates systematic disagreement and may suggest that the class labels have been reversed. The false positive rate measures the proportion of actual non-water pixels that are incorrectly classified as water. It is useful for evaluating the model’s ability to suppress false water detections caused by shadows, dark surfaces, vegetation, and other water-like background features. Lower values indicate fewer false positive water predictions.
In water body segmentation, precision, recall, the F1-score, and mean intersection over union (mIoU) are often used together to evaluate model performance and segmentation quality. Metric selection is dependent on applications. When overall performance is the primary concern, mIoU, the Dice coefficient, and pixel accuracy are commonly emphasized. When class proportions are relatively balanced, pixel accuracy can be a useful summary measure. For small target segmentation, the Dice coefficient and IoU are often more informative. When omission errors are critical, recall and the omission rate should be prioritized. Therefore, evaluation metrics should be selected according to the application context and the assessment objectives.
Although six water body segmentation models were evaluated by Moghimi et al. [86], Table 14 summarizes only the reported results for U-Net, DeepLabv3+, and SAM to provide a concise comparison among three distinct architectural paradigms. Transformer- and Mamba-based models are not included because they were not evaluated in the cited study, and no independent comparative experiments were conducted as part of this review. These models represent three distinct architectural paradigms. U-Net adopts a classical symmetric encoder and decoder structure, DeepLabv3+ represents an advanced conventional convolutional neural network, and SAM reflects the recent shift toward foundation models based on large scale vision Transformers. Other models, including LinkNet, PSPNet, and PAN, share substantial architectural similarities with conventional convolutional neural network baselines. They were therefore excluded to ensure that the comparison remained concise and representative. The models presented in Table 14 were configured as follows. U-Net and DeepLabv3+ both used ResNet50, pretrained on ImageNet, as the backbone network, whereas SAM used the lightweight ViT B variant. Images in the four datasets, namely, Kaggle WaterNet, Elbersdorf/Wesenitz, RIWA.v1, and LuFI RiverSnap.v1, were acquired using ground-based surveillance cameras, smartphones, or low-altitude unmanned aerial vehicles and were therefore classified as close-range remote sensing images. Consequently, the conclusions regarding model performance are limited to close-range optical water body segmentation and should not be directly generalized to satellite imagery, synthetic aperture radar imagery, or multispectral water body segmentation. These data differ substantially in their imaging mechanisms, spatial resolutions, and spectral characteristics.
As shown in Table 14, in the experiments reported by Moghimi et al. [86], model performance varied across datasets with different characteristics, indicating that no single water body segmentation model achieved the best results under all conditions. On the Kaggle WaterNet dataset, the SAM achieved the highest overall accuracy, Kappa coefficient, intersection over union, precision, and F1-score, with intersection over union and F1-score values of 0.901 and 0.939, respectively. DeepLabv3+ achieved the highest recall, indicating a greater ability to identify true water regions completely. On the Elbersdorf/Wesenitz dataset, the three models produced similar results, suggesting that performance had approached saturation. In the experiments reported by Moghimi et al., U-Net achieved the highest overall accuracy, Kappa coefficient, intersection over union, recall, and F1-score. Because the image acquisition locations, viewing perspectives, and background structures in this dataset were relatively consistent, all three models were able to learn the characteristics of both water bodies and background objects effectively. These results also indicate that extremely high accuracy on datasets with relatively fixed scenes does not necessarily reflect an equivalent capacity for generalization across different scenes. On the RIWA.v1 dataset, which contains greater variation in scene types and imaging devices, U-Net demonstrated the best overall performance, with intersection over union and F1-score values of 0.878 and 0.930, respectively. The SAM achieved the highest precision, but its recall was lower than those of U-Net and DeepLabv3+. This finding suggests that its predictions were relatively conservative, reducing false positive predictions in non-water regions while potentially omitting some true water regions. DeepLabv3+ achieved intersection over union and F1-score values of 0.873 and 0.926, respectively, with only a small performance difference from U-Net and a favorable balance between precision and recall. On the LuFI-RiverSnap.v1 dataset, which includes diverse imaging devices, complex water colors, and shadows, the SAM achieved the highest Kappa coefficient, intersection over union, precision, and F1-score. Its intersection over union and F1-score values reached 0.925 and 0.957, respectively, demonstrating strong adaptability to complex scenes. U-Net achieved the highest recall, indicating that it identified water regions more completely. However, its lower precision relative to the SAM suggests that it produced more false positive predictions in background regions. DeepLabv3+ also maintained relatively balanced performance on this dataset.
Within the three-model comparison reported by Moghimi et al. [86], U-Net demonstrated strong performance on the RIWA.v1 dataset, whereas DeepLabv3+ maintained a relatively balanced relationship between precision and recall across the four evaluated datasets. SAM achieved the best reported results for several metrics on the Kaggle WaterNet and LuFI-RiverSnap.v1 datasets. However, these observations are limited to the selected close-range optical datasets and should not be generalized to Transformer- or Mamba-based models or to satellite, SAR, and multispectral water body segmentation tasks.

5. Challenges and Outlook

5.1. Existing Challenges

To date, water body segmentation has advanced substantially, with notable gains in accuracy and generalization. However, challenges associated with environmental complexity, methodological limitations, and limited data availability remain unresolved. Key issues include the following:
  • Challenges in Extracting Small and Fragmented Water Bodies: Pixels corresponding to small and fragmented water bodies usually constitute only a small proportion of remote sensing imagery and provide limited discriminative information. During model downsampling, these weak cues are attenuated or lost, which impairs reliable detection. As a result, spatially continuous water features, such as narrow streams, may be segmented into disconnected fragments, yielding scattered outputs that resemble short line segments or even isolated points and thus exhibit pronounced fragmentation. Small water bodies reflect the same underlying challenge, namely, the limited spatial extent of the targets. Current methods still provide insufficient feature extraction for small water bodies, particularly for accurate boundary delineation.
    Although extracting small targets from fragmented water bodies remains highly challenging, it is critical to human survival [106,107]. Considerable efforts have therefore been devoted to addressing this problem. Previous studies have shown that the modified normalized difference water index, calibrated for small reservoirs with areas ranging from 1 to 10 ha, exhibits high threshold stability and can reduce false negative predictions caused by shallow water and submerged vegetation. However, this method cannot completely prevent the loss of information at the subpixel scale [108]. Various deep learning methods have recently been proposed as promising solutions to this challenge. Qin et al. applied U-Net to extract small water bodies from hyperspectral imagery. By exploiting spectral differences within the imagery, their method enhanced the discrimination between water bodies and background features and used skip connections to transfer detailed information from shallow feature layers to the decoder. Nevertheless, repeated downsampling may still reduce the segmentation accuracy of narrow water bodies [109]. WaterNet is a novel deep CNN that incorporates concepts from Gaussian filtering and edge detection into a dual residual refinement module for segmenting small water bodies [110]. In addition, the AWS16K dataset, which was developed by comprehensively considering sample coverage across different regions, water body types, and complex backgrounds, provides an effective supplementary resource for improving small water body segmentation [111]. Despite these advances, further progress is needed to overcome the limitations of traditional methods, account for variability across acquisition times and seasons, balance temporal and spatial resolution, and improve the interpretability of deep learning models to better satisfy real-world application requirements.
  • Misjudgments and Missed Judgments in Complex Scenarios: Misclassification and omission in surface water body segmentation are primarily caused by the combined effects of spectral similarity between classes, spectral variability within classes, and mixed pixels. Spectral similarity between different objects occurs when water bodies exhibit spectral responses similar to those of shadows, dark roads, or saline and alkaline land. Spectral variability within the same object occurs when a single water body displays different spectral signatures because of variations in water depth, turbidity, sun glint, or season. Mixed pixels, by contrast, result from the limited spatial resolution of the sensor. When a single pixel contains water, vegetation, soil, or buildings simultaneously, its observed spectrum can be approximated as the sum of the spectra of these components weighted by their respective area proportions.
    Distinguishing different objects with similar spectral responses from the same objects with variable spectral responses remains a major challenge in land and sea segmentation. To address data diversity and the lack of prior information, some studies have introduced additional encoding branches into end-to-end PMFormer architectures and employed deep clustering to capture substantial spectral variability within classes [112]. Saline and alkaline land is particularly difficult to distinguish from water because of their similar spectral characteristics. Using Lake Aibi in the Xinjiang Uygur Autonomous Region as a case study, researchers incorporated a saline and alkaline land endmember into a spectral mixing framework with multiple endmembers and dynamically selected thresholds according to image histogram characteristics, thereby improving the separability of water bodies from saline and alkaline land [113]. For urban water body segmentation, an initial water mask can be generated using the normalized difference water index and the Otsu threshold, followed by morphological dilation to identify mixed pixels [114]. In mapping below the pixel scale, DE_MRF applies both dilation and erosion to identify potential mixed pixels inside and outside water body boundaries and represents spectral variation within classes using multiple local water and land endmembers [115]. Other studies have used CNN to learn the relationship between local seawater background components and algal components and then adjusted the Otsu threshold according to differences in coverage before and after deconvolution. This approach provides a basis for feedback calibration with mixed pixels along water body boundaries [116]. More recently, methods for increasing spatial resolution have been applied to reduce spectral mixing along river boundaries and estimate the proportion of water within mixed pixels [117]. Despite this progress, the spatial distribution of water within mixed pixels cannot yet be characterized with sufficient accuracy. Moreover, many existing methods depend heavily on imagery with high spatial resolution, and irregular water bodies with nonstationary characteristics remain insufficiently investigated.
  • The Scarcity and Cost of High-quality Annotated Data: Deep learning model training depends critically on adequate samples; however, data availability for water body segmentation remains limited. Pixel-level annotations are particularly scarce for elongated rivers and small ponds, and high-quality labels that remain representative across regions, seasons, and terrain conditions are still lacking. In addition, existing datasets are often strongly imbalanced: annotated samples are concentrated on large water bodies, such as rivers, lakes, and reservoirs, whereas samples for small water bodies are comparatively limited. These constraints can substantially hinder the generalization of water body segmentation models.
    Training with manually annotated data is highly time consuming [118]. Facing the scarcity of high-quality annotated data, transfer learning has emerged as a practical alternative. For example, transfer learning can be incorporated into Siamese network architectures to reduce the dependence on labeled data. It can also be used to train water body segmentation models and thereby alleviate performance degradation caused by limited annotations. A two-stage transfer learning strategy has also been reported, in which models are pretrained on Sentinel-2 imagery and then fine-tuned on PlanetScope data to transfer spectral representations and reduce labeling costs [119,120,121]. Although these approaches generally reduce reliance on high quality annotations, they remain subject to limitations, including sensitivity to natural factors such as occlusion and illumination, as well as challenges associated with heterogeneous remote sensing imagery. Further research is needed to address these issues in a systematic manner.
  • The Challenge of Generalization Across Seasons and Cross-Regional Scenarios: Seasonal variation, temporal consistency, regional transferability, and generalization across sensors remain major challenges in water body segmentation. Water bodies exhibit substantial morphological differences across geographic regions and seasonal conditions, such as between humid plains and arid mountainous areas, between wet and dry seasons, and between snow- and ice-covered winter landscapes and vegetation covered summer landscapes. These variations can markedly reduce model robustness and transferability. In addition, differences in spatial resolution, spectral response, and imaging mechanisms constrain generalization across sensors. Consequently, achieving reliable performance across seasons, regions, and sensor types remains an unresolved challenge.
    A contrastive learning module was incorporated into SAM to align features across synthetic images, enabling robust water body segmentation across regions without requiring annotated data [122]. Conventional methods based on thresholding or supervised classification often require thresholds or training samples to be adjusted for different regions and sensors. By contrast, a large-scale framework for water body segmentation across sensors was developed using unsupervised deep learning. This framework exploits the physical and spatial characteristics of water bodies to automatically distinguish and collect diverse samples, achieving strong overall performance across different sensor types [123]. Because lakes exhibit distinct spectral characteristics in winter and summer, a method combining K means clustering with flood fill was proposed. By selecting appropriate metrics and removing small objects along lake boundaries, this method improved the temporal continuity, accuracy, and stability of lake extraction [124]. To accurately segment areas subject to seasonal flooding, researchers developed a harmonic model based on long-term time-series dynamics of surface water. The amplitude derived from the harmonic model was used to characterize the frequency of transitions between land and water, resulting in high accuracy and robustness [125]. Lakes also undergo transitions between periods with and without ice because of seasonal variation, and conventional spectral indices and machine learning methods remain limited in distinguishing these conditions. To address this issue, large language models were combined with random forest algorithms to generate candidate indices for distinguishing water, ice, and snow. The optimal ERNIE WISI index was then selected to automatically classify water, ice, and snow without requiring seasonal threshold adjustment [126]. Despite these advances, differences in temporal dynamics caused by climatic variability and the limited generalization of models under extreme conditions or in specific geographic environments remain unresolved.

5.2. Future Research Directions

Deep learning-based approaches for water body segmentation currently face systemic challenges. Future research is expected to move beyond accuracy gains alone and toward greater model intelligence, standardization, and practical applicability. In the coming years, research directions include the following:
  • Spatiotemporal Adaptive Multimodal Fusion: Recent studies have incorporated multi-source remote sensing data for water body segmentation. For instance, Sentinel-1 data have been used to complement information missing in Sentinel-2 imagery [127]; SAR data and digital elevation models have been integrated for flood mapping [128]; and a combination of Landsat-8 and Landsat-9 with Sentinel-1 and Sentinel-2 has been applied to estimate irrigated areas in the Delingha piedmont grasslands of northwestern China [129]. Although these efforts have yielded promising results, most existing methods still rely on a single remote sensing image and often neglect multi-temporal information. Multisource data provide complementary cues for water delineation, reduce the effects of cloud cover and illumination variability, improve spatial resolution, and offer stronger robustness in complex scenes. Multi-temporal data identify the variation of periodic water dynamics and facilitate discrimination between persistent water bodies and short-term inundation, while improving the monitoring of extreme events such as floods and dam failures. Despite the clear benefits of combining multisource and multi-temporal data, further methodological improvements remain necessary. Future research is expected to move beyond conventional multimodal fusion toward spatial, temporal, and feature adaptivity. Multisource data, including optical imagery, SAR imagery, thermal infrared imagery, digital elevation models, historical time series data, LiDAR point clouds, and related sources, may be selected adaptively according to application scenarios and integrated through fusion networks with dynamic weighting. These advances are expected to improve discrimination of fine-grained water bodies.
  • Segmentation Strategies for Large Model Integration: General-purpose vision foundation models, represented by the SAM series, have attracted considerable attention in water body segmentation and are increasingly being applied in this field. The Prithvi-EO series [130] and SeaMo [131] have also been introduced for water body segmentation, and related studies are currently emerging. Although TerraMind [132], SkySense [133], and SpectralGPT [134] have not yet been specifically investigated for water body segmentation, they already possess the multimodal and cross-domain representation capabilities required for this task. However, these models remain dependent on prompts and do not fully address the challenges associated with domain shift; consequently, their cross-domain generalization performance remains unstable. With ongoing advances in science and technology, requirements for integrating large-scale foundation models are expected to become more stringent. The scope is expected to extend beyond water body delineation to forecast future dynamics, such as changes in drought- and flood-affected areas. For task-specific models, a direction is the development of geospatial foundation models that generalize across seasons and regions, thereby improving generalization and zero-shot capability. In addition, reducing parameter counts and applying model compression to support lightweight deployment are likely to remain priorities for future research.
  • Fine-grained Water Body Segmentation: Current water body segmentation methods largely remain at the level of holistic detection and lack the capability for fine-grained segmentation. This is because they fail to account for intrinsic differences in water type and composition that commonly coexist within a single water body, such as clear water, polluted water, and areas dominated by aquatic vegetation. Since these distinct water types require different management strategies and have varied environmental impacts, conventional pixel-level segmentation—which does not provide water type information—can lead to inaccurate parameterization in Earth system models, including those for water quality, ecological assessment, hydrology, and climate. To address contemporary application needs, it is essential to move beyond binary water classification. Fine-grained water body segmentation can not only delineate water extent more precisely but also support the prediction of water evolution, retrieval of water quality parameters, and identification of periodic variations in water type. Furthermore, it enables interdisciplinary applications and provides critical support for climate research, fisheries management, and water resources planning. Future research is expected to advance toward finer water body segmentation, with emphasis on pixel-level boundary delineation and the retrieval of higher-level water body-related semantics and parameters. This shift moves beyond locating water bodies toward identifying water types, enabling a more refined semantic segmentation of turbid waters, black and malodorous waters, aquatic vegetation-covered waters, shadow-affected waters, and ice surfaces. These advances are expected to promote a transition from perception to cognition and thereby better utilize effective water resource use. To achieve more precise water body segmentation, future studies could incorporate neural networks informed by physical principles to constrain the transport, diffusion, and decay of pollutants, suspended matter, and chlorophyll along flow fields. In addition, shallow water equations or shoreline kinematic conditions could be used to constrain variations in water body boundaries. On this basis, a joint representation framework could be developed to delineate water body boundaries while simultaneously predicting semantic categories, such as clear and turbid water, or estimating continuous parameters, including chlorophyll a concentration, suspended solids concentration, and turbidity.
  • Integration of Explainability and Physical Mechanisms: Although recent improvements in water body segmentation models have enhanced segmentation performance, the mechanisms underlying these gains remain insufficiently understood. Moreover, when models are applied to imagery acquired at different spatial resolutions or by different sensors, nonlinear variations in segmentation accuracy cannot be explained solely by differences in the input data. Future research should therefore adopt an approach jointly driven by data and physical knowledge. Physical constraints, such as bidirectional reflectance distribution functions and residuals derived from the two- dimensional shallow water equations, could be incorporated into segmentation models to improve their interpretability. However, three factors should be considered when integrating physical mechanisms. First, the required physical variables must be clearly defined. For example, combining atmospheric aerosol optical thickness with water quality parameters may help models to distinguish water more accurately from shadows and vegetation. Second, the temporal requirements of the observations must be satisfied. Image sequences should cover a complete evolution cycle so that the model captures the actual dynamics of water bodies rather than merely fitting the static characteristics of individual images. Third, the applicable conditions of each physical mechanism must be specified. For example, the shallow water equations are effective for open water bodies over gently varying terrain but may be less reliable in mountainous regions or urban areas affected by complex drainage systems and frequent flooding. Under these conditions, physical mechanisms can provide objective criteria for model evaluation, while interpretability methods can offer more reliable visual evidence. A close integration of these approaches may transform water body segmentation from opaque inference into transparent decision making. This shift would facilitate a clearer understanding of why models confuse water with non-water surfaces and support the diagnosis and reduction of false classifications and omissions in complex environments.

6. Summary

The rapid and accurate extraction of water bodies, including rivers, reservoirs, lakes, and flood-affected areas, from complex remote sensing scenes is essential for flood disaster management, water resource investigation and management, ecological and environmental monitoring and protection, agriculture, and aquaculture. This paper reviews water body segmentation methods based on U-Net, DeepLabv3+, Transformer, Mamba, and the SAM and evaluates their performance in preserving local details, modeling multiscale contextual information, capturing long-range dependencies, and maintaining computational efficiency, thereby providing new perspectives for future research. Overall, each model has distinct advantages and limitations. By integrating shallow spatial details with deep semantic information, U-Net is well suited to segmenting narrow rivers and small ponds. However, it is susceptible to confusion between water bodies and shadows or dark objects. DeepLabv3+ employs dilated convolutions with different dilation rates and is suitable for multiscale scenes in which lakes, rivers, and ponds coexist. Nevertheless, it may produce fragmented predictions for elongated rivers or irregular shoreline boundaries. Transformer-based models are effective for segmenting large water bodies across different seasons and regions, but they may lose fine spatial details when training data are limited or downsampling is excessive. Mamba balances the modeling of long-range dependencies with computational efficiency, making it suitable for continuous river networks and large-scale remote sensing imagery. However, it requires a high-resolution decoder to reconstruct shorelines and small water bodies accurately. The SAM is suitable for interactive extraction, assisted annotation, and few-shot transfer, but it may produce fragmented masks or incomplete water body segmentation.
In addition to selecting and improving existing water body segmentation models, future research should strengthen the fusion of optical, SAR, and multispectral imagery with data acquired at multiple times and with topographic and hydrological information. The contributions of different data sources should also be adjusted dynamically under adverse conditions, such as cloud cover, shadows, and noise. Further research should investigate foundation models, including Prithvi-EO and SeaMo, to reduce annotation costs and improve transfer learning and zero-shot water body segmentation. The conventional binary classification of water and other surfaces should be extended through a collaborative representation framework for multiple tasks. Such a framework could separate shoreline and geometric boundary extraction from water type identification, including clear water, turbid water, and aquatic vegetation, while also supporting the estimation of water quality parameters and the prediction of water body evolution. In addition, the bidirectional reflectance distribution function of the water surface, mass conservation, and the shallow water equations could be incorporated into the loss function as physical constraints to further improve segmentation performance. In the future, water body segmentation based on deep learning is expected to evolve from an isolated task into an integrated research framework that combines remote sensing, computer vision, hydrology, and environmental science.

Author Contributions

Conceptualization, D.L. and F.Y.; methodology, F.Y.; validation, Y.H., H.H. and F.Y.; formal analysis, D.L.; investigation, F.Y.; resources, F.Y.; data curation, D.L.; writing—original draft preparation, D.L.; writing—review and editing, F.Y.; visualization, H.H.; supervision, Y.H.; project administration, F.Y.; funding acquisition, F.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the National Natural Science Foundation of China under Grant 42461054, in part by Natural Science Foundation of Jiangxi Province under Grant 20202BABL202030, in part by Science and Technology project plan of Jiangxi Provincial Water Resources Department, China under Grant 202426ZDKT16, and in part by Jiangxi Provincial Training Project of Disciplinary, Academic, and Technical Leader under Grant 20232BCJ22002.

Data Availability Statement

The datasets used or analyzed during the current study are available from the corresponding author on reasonable request.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT-5.5 to improve English expression and readability, as well as to enhance the clarity of a figure. After using this tool, the authors carefully reviewed and edited all text and images. The authors confirm that the AI tool was not used to alter the original scientific data or generate fabricated figures, and they take full responsibility for the final content of the publication.

Conflicts of Interest

The authors declare that this study was conducted in the absence of any commercial or financial relationships that could be considered potential conflicts of interest.

References

  1. He, X.; Zhang, S.; Xue, B.; Zhao, T.; Wu, T. Cross-modal change detection flood extraction based on convolutional neural network. Int. J. Appl. Earth Obs. Geoinf. 2023, 117, 103197. [Google Scholar] [CrossRef] [Scilit]
  2. Gupta, S. Infectious disease: Something in the water. Nature 2016, 533, S114–S115. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Zhang, Z.; Chen, B.; Li, J.; Xie, W.; Yang, B.; Bao, Y.; Xie, Y.; Wang, Q.; Wei, Y.; Zhang, W.; et al. Distinctive water bodies surrounding lakes: An effective indicator for drought monitoring and assessment. J. Hydrol. 2024, 645, 132179. [Google Scholar] [CrossRef] [Scilit]
  4. Fu, B.; Li, X.; Jiang, L.; Deng, J.; Gan, Y.; Yao, H.; Fan, D. Revealing two decades of water quality dynamics and driving mechanism in the Qinjiang river using multisource data and LSTM-eKAN model. Ecol. Indic. 2025, 179, 114272. [Google Scholar] [CrossRef] [Scilit]
  5. Cao, Y.; Yuan, Y.; Dong, H.; Yuan, X. Water extraction games in river basins from the perspective of complex network: A case study in the Hanjiang river basin, China. J. Hydrol. 2025, 663, 134252. [Google Scholar] [CrossRef] [Scilit]
  6. Cai, J.; Tao, L.; Li, Y. CM-UNet++: A Multi-Level Information Optimized Network for Urban Water Body Extraction from High-Resolution Remote Sensing Imagery. Remote Sens. 2025, 17, 980. [Google Scholar] [CrossRef] [Scilit]
  7. Zhou, Y.; Zaitchik, B.F.; Kumar, S.V.; Nie, W.; Loomis, B.D.; McLarty, A.S.R.; Appana, R. Satellite-informed simulation of irrigation in South Asia: Opportunities and uncertainties. J. Hydrol. 2024, 641, 131758. [Google Scholar] [CrossRef] [Scilit]
  8. Ahmad, S.K.; Hossain, F.; Eldardiry, H.; Pavelsky, T.M. A Fusion Approach for Water Area Classification Using Visible, Near Infrared and Synthetic Aperture Radar for South Asian Conditions. IEEE Trans. Geosci. Remote Sens. 2019, 58, 2471–2480. [Google Scholar] [CrossRef] [Scilit]
  9. Zhao, C.; Wei, H.; Feyisa, G.L.; Tayer, T.d.C.; Ma, G.; Wu, H.; Pan, Y. Evaluating spectral indices for water extraction: Limitations and contextual usage recommendations. Int. J. Appl. Earth Obs. Geoinf. 2025, 139, 104510. [Google Scholar] [CrossRef] [Scilit]
  10. Paul, A.; Tripathi, D.; Dutta, D. Application and comparison of advanced supervised classifiers in extraction of water bodies from remote sensing images. Sustain. Water Resour. Manag. 2017, 4, 905–919. [Google Scholar] [CrossRef] [Scilit]
  11. Yang, J.; Wang, X.; Wang, J.; Ye, C.; Xiong, J. Water extraction of hyperspectral imagery based on a fast and effective decision tree water index. J. Appl. Remote Sens. 2021, 15, 42605. [Google Scholar] [CrossRef] [Scilit]
  12. Zhai, M.; Shen, H.; Cao, Q.; Ding, X.; Xin, M. Water Body Extraction Methods for SAR Images Fusing Sentinel-1 Dual-Polarized Water Index and Random Forest. Sensors 2025, 25, 4868. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Kaplan, G.; Avdan, U. Object-based water body extraction model using Sentinel-2 satellite imagery. Eur. J. Remote Sens. 2017, 50, 137–143. [Google Scholar] [CrossRef] [Scilit]
  14. Nagaraj, R.; Kumar, L.S. Extraction of Surface Water Bodies using Optical Remote Sensing Images: A Review. Earth Sci. Inform. 2024, 17, 893–956. [Google Scholar] [CrossRef] [Scilit]
  15. Gautam, S.; Singhai, J. Critical review on deep learning methodologies employed for water-body segmentation through remote sensing images. Multimed. Tools Appl. 2024, 83, 1869–1889. [Google Scholar] [CrossRef] [Scilit]
  16. Rajeswari, S.; Rathika, P. Emerging methodologies in waterbody delineation: An In-depth review. Int. J. Remote Sens. 2024, 45, 5789–5819. [Google Scholar] [CrossRef] [Scilit]
  17. Li, J.; Ma, R.; Cao, Z.; Xue, K.; Xiong, J.; Hu, M.; Feng, X. Satellite Detection of Surface Water Extent: A Review of Methodology. Water 2022, 14, 1148. [Google Scholar] [CrossRef] [Scilit]
  18. Bijeesh, T.V.; Narasimhamurthy, K.N. Surface water detection and delineation using remote sensing images: A review of methods and algorithms. Sustain. Water Resour. Manag. 2020, 6, 399–411. [Google Scholar] [CrossRef] [Scilit]
  19. Sigopi, M.; Shoko, C.; Dube, T. Advancements in remote sensing technologies for accurate monitoring and management of surface water resources in Africa: An overview, limitations, and future directions. Geocarto Int. 2024, 39, 2347935. [Google Scholar] [CrossRef] [Scilit]
  20. Guo, Z.; Wu, L.; Huang, Y.; Guo, Z.; Zhao, J.; Li, N. Water-Body Segmentation for SAR Images: Past, Current, and Future. Remote Sens. 2022, 14, 1752. [Google Scholar] [CrossRef] [Scilit]
  21. Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3431–3440. [Google Scholar]
  22. Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), Munich, Germany, 5–9 October 2015; pp. 234–241. [Google Scholar]
  23. Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2018; pp. 801–818. [Google Scholar]
  24. Badrinarayanan, V.; Kendall, A.; Cipolla, R. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 2481–2495. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2881–2890. [Google Scholar]
  26. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5999–6007. [Google Scholar]
  27. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 9992–10002. [Google Scholar]
  28. Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
  29. Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.-Y.; et al. Segment anything. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2023; pp. 4015–4026. [Google Scholar]
  30. Shi, Z.; Wu, Y. CEA-SAM: Context-Enhanced Adaptation of Segment Anything Model for Robust Water Body Extraction from UAV Imagery. In 2025 4th International Conference on Artificial Intelligence, Internet of Things and Cloud Computing Technology (AIoTC); IEEE: New York, NY, USA, 2025; pp. 87–91. [Google Scholar]
  31. Zamboni, P.A.P.; Gorriz, X.B.; Junior, J.M.; Gonçalves, W.N.; Eltner, A. Do we need to label large datasets for river water segmentation? Benchmark and stage estimation with minimum to non-labeled image time series. Int. J. Remote Sens. 2025, 46, 2719–2747. [Google Scholar] [CrossRef] [Scilit]
  32. Osco, L.P.; Wu, Q.; de Lemos, E.L.; Gonçalves, W.N.; Ramos, A.P.M.; Li, J.; Marcato, J. The Segment Anything Model (SAM) for remote sensing applications: From zero to one shot. Int. J. Appl. Earth Obs. Geoinf. 2023, 124, 103540. [Google Scholar] [CrossRef] [Scilit]
  33. O’sUllivan, C.; Kashyap, A.; Coveney, S.; Monteys, X.; Dev, S. Enhancing coastal water body segmentation with Landsat Irish Coastal Segmentation (LICS) dataset. Remote Sens. Appl. Soc. Environ. 2024, 36, 101276. [Google Scholar] [CrossRef] [Scilit]
  34. Konapala, G.; Kumar, S.V.; Ahmad, S.K. Exploring Sentinel-1 and Sentinel-2 diversity for flood inundation mapping using deep learning. ISPRS J. Photogramm. Remote Sens. 2021, 180, 163–173. [Google Scholar] [CrossRef] [Scilit]
  35. Hu, H.; He, Z.; Zheng, H. Multispectral water recognition algorithm for complex environments. J. Beijing Univ. Aeronaut. Astronaut. 2025, 1–14. (In Chinese) [Google Scholar] [CrossRef]
  36. Wang, F.; Feng, X. Flood change detection model based on an improved U-net network and multi-head attention mechanism. Sci. Rep. 2025, 15, 3295. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Jonnala, N.S.; Siraaj, S.; Prastuti, Y.; Chinnababu, P.; Babu, B.P.; Bansal, S.; Upadhyaya, P.; Prakash, K.; Faruque, M.R.I.; Al-Mugren, K.S. AER U-Net: Attention-enhanced multi-scale residual U-Net structure for water body segmentation using Sentinel-2 satellite images. Sci. Rep. 2025, 15, 16099. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Zhang, J.; Wang, H.; Li, J.; Bao, A.; Wu, H.; Li, S.; Shen, Z. Information extraction of typical floodplain wetlands based on an improved D-UNet model. Natl. Remote Sens. Bull. 2025, 29, 300–313. (In Chinese) [Google Scholar]
  39. Xu, X.; Zhang, T.; Liu, H.; Guo, W.; Zhang, Z. An Information-Expanding Network for Water Body Extraction Based on U-Net. IEEE Geosci. Remote Sens. Lett. 2024, 21, 1502205. [Google Scholar] [CrossRef] [Scilit]
  40. Wang, M.; Li, C.; Yang, X.; Ban, Y.; Chu, D.; Zhou, Z.; Lau, R.Y.K. QTU-Net: Quaternion Transformer-Based U-Net for Water Body Extraction of RGB Satellite Image. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5634816. [Google Scholar] [CrossRef] [Scilit]
  41. Wagner, F.; Eltner, A.; Maas, H.-G. River water segmentation in surveillance camera images: A comparative study of offline and online augmentation using 32 CNNs. Int. J. Appl. Earth Obs. Geoinf. 2023, 119, 103305. [Google Scholar] [CrossRef] [Scilit]
  42. Liu, D.; Wang, W.; Shao, L.; Wang, X.; Zhang, Z.; Yang, L.; Du, Y.; Ma, Y.; Fu, D.; Zhang, X. Image monitoring method for open-channel water level based on semantic segmentation. J. China Agric. Univ. 2025, 30, 241–252. (In Chinese) [Google Scholar]
  43. Xia, M.; Cui, Y.; Zhang, Y.; Xu, Y.; Liu, J.; Xu, Y. DAU-Net: A novel water areas segmentation structure for remote sensing image. Int. J. Remote Sens. 2021, 42, 2594–2621. [Google Scholar] [CrossRef] [Scilit]
  44. Li, Y.; Liu, X.; Ferreira, V.; Balzter, H.; Zhou, H.; Ge, Y.; Lai, M.; Chu, S.; Ding, H.; Gu, Z. Surface water mapping from remote sensing in Egypt’s dry season using an improved U-Net model with multi-scale information and attention mechanism. Int. J. Appl. Earth Obs. Geoinf. 2025, 142, 104666. [Google Scholar] [CrossRef] [Scilit]
  45. Cai, H.; Gong, J.; Zhang, Y.; Wang, J.; Hu, W. Water body segmentation method based on an improved U-Net model. Remote Sens. Inf. 2024, 39, 140–147. (In Chinese) [Google Scholar]
  46. Bai, Q.; Luo, X.; Mu, S. Surface water extraction from high-resolution remote sensing images based on TA-UNet3+. Comput. Eng. Appl. 2024, 61, 245. (In Chinese) [Google Scholar]
  47. Xing, G.; Lu, G.; Han, B. MAFUNet: SAR image water segmentation algorithm combining an attention mechanism and active contour loss. Acta Geod. Cartogr. Sin. 2025, 54, 924–936. (In Chinese) [Google Scholar]
  48. Zhang, X.; Dai, P.; Li, W.; Ren, N.; Mao, X. Extraction of freshwater aquaculture areas based on improved coordinate attention and a U-Net neural network. Trans. Chin. Soc. Agric. Eng. 2023, 39, 153–162. (In Chinese) [Google Scholar]
  49. Xiang, D.; Zhang, X.; Wu, W.; Liu, H. DensePPMUNet-a: A Robust Deep Learning Network for Segmenting Water Bodies From Aerial Images. IEEE Trans. Geosci. Remote Sens. 2023, 61, 4202611. [Google Scholar] [CrossRef] [Scilit]
  50. Hertel, V.; Chow, C.; Wani, O.; Wieland, M.; Martinis, S. Probabilistic SAR-based water segmentation with adapted Bayesian convolutional neural network. Remote Sens. Environ. 2023, 285, 113388. [Google Scholar] [CrossRef] [Scilit]
  51. Li, M.; Wu, P.; Wang, B.; Park, H.; Hui, Y.; Yanlan, W. A Deep Learning Method of Water Body Extraction From High Resolution Remote Sensing Images With Multisensors. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2021, 14, 3120–3132. [Google Scholar] [CrossRef] [Scilit]
  52. Ye, F.; Zhang, R.; Xu, X.; Wu, K.; Zheng, P.; Li, D. Water Body Segmentation of SAR Images Based on SAR Image Reconstruction and an Improved UNet. IEEE Geosci. Remote Sens. Lett. 2023, 21, 4010005. [Google Scholar] [CrossRef] [Scilit]
  53. Liu, T.; Zhang, M.; Feng, J.; Li, Q.; Zhu, Y. Remote sensing identification and dynamic monitoring of rural black-odorous water bodies based on GF-2 imagery. Bull. Surv. Mapp. 2025, 5, 8–14. (In Chinese) [Google Scholar]
  54. Pinheiro, M.M.F.; de Oliveira, L.Y.D.; Venancio, T.E.B.; Nogueira, K.; Júnior, J.M.; Gonçalves, W.N.; Pereira, D.R.; Osco, L.P.; Ramos, A.P.M. Deep learning on segmenting large and narrow rivers with aerial RGB imagery: A comparison of convolutional and Vision-Transformer networks. Remote Sens. Appl. Soc. Environ. 2026, 42, 101970. [Google Scholar] [CrossRef] [Scilit]
  55. Quang, N.H.; Lee, H.; Ahn, S.; Kim, G. Sprawling lake segmentations from space-bone SAR imagery by fine-tuning the DeepLabV3+ model. Adv. Space Res. 2025, 76, 6042–6065. [Google Scholar] [CrossRef] [Scilit]
  56. Wang, J.; Liu, X.; Wang, J.; Yang, M. An improved DeepLabv3+ network-based deep learning segmentation method for thermal image water-shorelines. Digit. Signal Process. 2025, 167, 105461. [Google Scholar] [CrossRef] [Scilit]
  57. Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1800–1807. [Google Scholar]
  58. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
  59. Luo, Y.; Feng, A.; Li, H.; Li, D.; Wu, X.; Liao, J.; Zhang, C.; Zheng, X.; Pu, H. New deep learning method for efficient extraction of small water from remote sensing images. PLoS ONE 2022, 17, e0272317. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Zhang, H.; Hu, R.; Jiang, Y.; Hu, Y. River water extraction based on an improved DeepLabv3+ model. Remote Sens. Inf. 2023, 38, 146–152. (In Chinese) [Google Scholar] [CrossRef] [Scilit]
  61. Sun, D.; Gao, G.; Huang, L.; Liu, Y.; Liu, D. Extraction of water bodies from high-resolution remote sensing imagery based on a deep semantic segmentation network. Sci. Rep. 2024, 14, 14604. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Huang, J.; Xu, J.; Yan, W.; Wu, P.; Xing, H. Detection of Black and Odorous Water in Gaofen-2 Remote Sensing Images Using the Modified DeepLabv3+ Model. Sustainability 2023, 16, 92. [Google Scholar] [CrossRef] [Scilit]
  63. Xue, F.; Lv, X.; Chen, X. Urban waterlogging monitoring method based on an improved DeepLabv3+ model. Sci. Surv. Mapp. 2023, 48, 216–224. (In Chinese) [Google Scholar]
  64. Zhang, C.; Ge, Y.; Ren, Y.; Gao, F.; Han, Y. Segmentation of duckweed-type rural black-odorous water bodies from high-resolution imagery using an optimized DeepLabv3+ network. Remote Sens. Technol. Appl. 2023, 38, 1433–1444. (In Chinese) [Google Scholar]
  65. Chen, Y.; He, J.; Liu, G.; Xiong, R.; Chen, N. Extraction of water-highlight regions using a Swin Transformer network model. Remote Sens. Inf. 2023, 38, 129–136. (In Chinese) [Google Scholar]
  66. Chen, L.; Long, F.; Li, Z.; Yuan, Z.; Zhu, W.; Cai, X. Multi-level feature-attention water extraction network for multi-source SAR images. Geomat. Inf. Sci. Wuhan Univ. 2025, 50, 1339–1345. (In Chinese) [Google Scholar]
  67. Lv, S.; Meng, L.; Edwing, D.; Xue, S.; Geng, X.; Yan, X.-H. High-Performance Segmentation for Flood Mapping of HISEA-1 SAR Remote Sensing Images. Remote Sens. 2022, 14, 5504. [Google Scholar] [CrossRef] [Scilit]
  68. Wang, L.; Liu, W.; Zhang, L.; Li, E.; Guo, F.; Lu, Q. MSFSwin: A multi-feature fusion water extraction method combining SAR imagery and an improved Swin Transformer. J. Geo-Inf. Sci. 2025, 27, 1638–1655. (In Chinese) [Google Scholar]
  69. Ma, D.; Jiang, L.; Li, J.; Shi, Y. Water index and Swin Transformer Ensemble (WISTE) for water body extraction from multispectral remote sensing images. GIScience Remote Sens. 2023, 60, 2251704. [Google Scholar] [CrossRef] [Scilit]
  70. Sudakow, I.; Asari, V.K.; Liu, R.; Demchev, D. MeltPondNet: A Swin Transformer U-Net for Detection of Melt Ponds on Arctic Sea Ice. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 8776–8784. [Google Scholar] [CrossRef] [Scilit]
  71. Zhang, Z.; Deng, G.; Luo, C.; Li, X.; Ye, Y.; Xian, D. A Multiscale Dual Attention Network for the Automatic Classification of Polar Sea Ice and Open Water Based on Sentinel-1 SAR Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 5500–5516. [Google Scholar] [CrossRef] [Scilit]
  72. Zhang, Y.; Lu, H.; Ma, G.; Zhao, H.; Xie, D.; Geng, S.; Tian, W.; Sian, K.T.C.L.K. MU-Net: Embedding MixFormer into Unet to Extract Water Bodies from Remote Sensing Images. Remote Sens. 2023, 15, 3559. [Google Scholar] [CrossRef] [Scilit]
  73. Zhao, T.; Du, X.; Xu, C.; Jian, H.; Pei, Z.; Zhu, J.; Yan, Z.; Fan, X. SPT-UNet: A Superpixel-Level Feature Fusion Network for Water Extraction from SAR Imagery. Remote Sens. 2024, 16, 2636. [Google Scholar] [CrossRef] [Scilit]
  74. Xiao, T.; Liu, Y.; Huang, Y.; Li, M.; Yang, G. Enhancing Multiscale Representations With Transformer for Remote Sensing Image Semantic Segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 3256064. [Google Scholar] [CrossRef] [Scilit]
  75. Kang, J.; Guan, H.; Ma, L.; Wang, L.; Xu, Z.; Li, J. WaterFormer: A coupled transformer and CNN network for waterbody detection in optical remotely-sensed imagery. ISPRS J. Photogramm. Remote Sens. 2023, 206, 222–241. [Google Scholar] [CrossRef] [Scilit]
  76. Tian, Y.; Cao, H.; Liu, Y.; Tian, C.; Wang, R. WB-Former: A Hybrid Model of CNN and Transformer for Water Body Extraction in Complex Scenes. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4210918. [Google Scholar] [CrossRef] [Scilit]
  77. Zhang, Q.; Hu, X.; Xiao, Y. A novel hybrid model based on cnn and multi-scale transformer for extracting water bodies from high resolution remote sensing images. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2023, 10, 889–894. [Google Scholar] [CrossRef] [Scilit]
  78. Zhong, H.-F.; Sun, Q.; Sun, H.-M.; Jia, R.-S. NT-Net: A Semantic Segmentation Network for Extracting Lake Water Bodies From Optical Remote Sensing Images Based on Transformer. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5627513. [Google Scholar] [CrossRef] [Scilit]
  79. Yang, Q.; Rao, L.; Fan, G.; Chen, N.; Cheng, S.; Song, X.; Yang, D. WatNet: A high-precision water body extraction method in remote sensing images under complex backgrounds. J. Appl. Remote Sens. 2024, 18, 44515. [Google Scholar] [CrossRef] [Scilit]
  80. Wang, S.; Wei, B.; Shi, B.; Wang, N.; Zhang, Y.; Zhu, Y. MHNet: A Masked Hybrid Network for Robust Water Body Segmentation From Aerial Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4208015. [Google Scholar] [CrossRef] [Scilit]
  81. Fan, J.; Zhou, J.; Wang, X.; Wang, J. A Self-Supervised Transformer With Feature Fusion for SAR Image Semantic Segmentation in Marine Aquaculture Monitoring. IEEE Trans. Geosci. Remote Sens. 2023, 61, 4207915. [Google Scholar] [CrossRef] [Scilit]
  82. Fan, J.; Li, M.; Wang, X. Unsupervised Transformer With Generative Label Optimization for Marine Aquaculture Segmentation. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4205714. [Google Scholar] [CrossRef] [Scilit]
  83. Li, J.; Liu, Z.; Liu, S.; Wang, H. MBSSNet: A Mamba-Based Joint Semantic Segmentation Network for Optical and SAR Images. IEEE Geosci. Remote Sens. Lett. 2025, 22, 6004305. [Google Scholar] [CrossRef] [Scilit]
  84. Yang, D.; Gao, X.; Tao, Y.; Yang, Y.; Guo, K.; Han, K.; Xu, L.; Li, H. Integrity Extraction of Water Bodies in Complex Scenes Based on the Adaptive Watermamba Framework. IEEE Trans. Geosci. Remote Sens. 2026, 64, 4207416. [Google Scholar] [CrossRef] [Scilit]
  85. Song, X.; Huang, H.; Zhang, L.; Xu, H.; Bai, Y. Enabling multi-decadal braided river monitoring through FloodMamba-Net and task-equivalent Landsat-to-Sentinel-2 data synthesis. J. Hydrol. 2025, 665, 134737. [Google Scholar] [CrossRef] [Scilit]
  86. Moghimi, A.; Welzel, M.; Celik, T.; Schlurmann, T. A Comparative Performance Analysis of Popular Deep Learning Models and Segment Anything Model (SAM) for River Water Segmentation in Close-Range Remote Sensing Imagery. IEEE Access 2024, 12, 52067–52085. [Google Scholar] [CrossRef] [Scilit]
  87. Ma, X.; Wu, Q.; Zhao, X.; Zhang, X.; Pun, M.-O.; Huang, B. SAM-Assisted Remote Sensing Imagery Semantic Segmentation With Object and Boundary Constraints. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5636916. [Google Scholar] [CrossRef] [Scilit]
  88. Zhang, T.; Ren, Y.; Li, W.; Qin, C.; Jiao, L.; Su, H. CSW-SAM: A cross-scale algorithm for very-high-resolution water body segmentation based on segment anything model 2. ISPRS J. Photogramm. Remote Sens. 2025, 228, 208–227. [Google Scholar] [CrossRef] [Scilit]
  89. Qiao, Y.; Zhong, B.; Du, B.; Cai, H.; Jiang, J.; Liu, Q.; Yang, A.; Wu, J.; Wang, X. SAM Enhanced Semantic Segmentation for Remote Sensing Imagery Without Additional Training. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5610816. [Google Scholar] [CrossRef] [Scilit]
  90. Zou, J.; He, W.; Wang, H.; Zhang, H. SAM-CTMapper: Utilizing segment anything model and scale-aware mixed CNN-Transformer facilitates coastal wetland hyperspectral image classification. Int. J. Appl. Earth Obs. Geoinf. 2025, 139, 104469. [Google Scholar] [CrossRef] [Scilit]
  91. Zhang, Y.; Wang, X.; Cai, J.; Yang, Q. MW-SAM:Mangrove wetland remote sensing image segmentation network based on segment anything model. IET Image Process. 2024, 18, 4503–4513. [Google Scholar] [CrossRef] [Scilit]
  92. Tong, X.-Y.; Xia, G.-S.; Lu, Q.; Shen, H.; Li, S.; You, S.; Zhang, L. Land-cover classification with high-resolution remote sensing images using transferable deep models. Remote Sens. Environ. 2020, 237, 111322. [Google Scholar] [CrossRef] [Scilit]
  93. Zhang, M.; Hu, X.; Zhao, L.; Lv, Y.; Luo, M.; Pang, S. Learning Dual Multi-Scale Manifold Ranking for Semantic Segmentation of High-Resolution Images. Remote Sens. 2017, 9, 500. [Google Scholar] [CrossRef] [Scilit]
  94. Demir, I.; Koperski, K.; Lindenbaum, D.; Pang, G.; Huang, J.; Basu, S.; Hughes, F.; Tuia, D.; Raskar, R. DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA, 18–22 June 2018; pp. 172–181. [Google Scholar]
  95. Li, X.; Zhang, G.; Cui, H.; Hou, S.; Wang, S.; Li, X.; Chen, Y.; Li, Z.; Zhang, L. MCANet: A joint semantic segmentation framework of optical and SAR images for land use classification. Int. J. Appl. Earth Obs. Geoinf. 2022, 106, 102638. [Google Scholar] [CrossRef] [Scilit]
  96. Wang, J.; Zheng, Z.; Ma, A.; Lu, X.; Zhong, Y. LoveDA: A remote sensing land-cover dataset for domain adaptive semantic segmentation. arXiv 2021, arXiv:2110.08733. [Google Scholar]
  97. Schmitt, M.; Hughes, L.H.; Qiu, C.; Zhu, X.X. SEN12MS—A Curated Dataset of Georeferenced Multi-Spectral Sentinel-1/2 Imahery for Deep Learning and Data Fusion. arXiv 2019, arXiv:1906.07789. [Google Scholar]
  98. Zhang, G.; Yao, T.; Chen, W.; Zheng, G.; Shum, C.K.; Yang, K.; Piao, S.; Sheng, Y.; Yi, S.; Li, J.; et al. Regional differences of lake evolution across China during 1960s–2015 and its natural and anthropogenic causes. Remote Sens. Environ. 2019, 221, 386–404. [Google Scholar] [CrossRef] [Scilit]
  99. Pi, X.; Luo, Q.; Feng, L.; Xu, Y.; Tang, J.; Liang, X.; Ma, E.; Cheng, R.; Fensholt, R.; Brandt, M.; et al. Mapping global lake dynamics reveals the emerging roles of small lakes. Nat. Commun. 2022, 13, 5777. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  100. Pickens, A.H.; Hansen, M.C.; Hancher, M.; Stehman, S.V.; Tyukavina, A.; Potapov, P.; Marroquin, B.; Sherani, Z. Mapping and sampling to characterize global inland water dynamics from 1999 to 2018 with full Landsat time-series. Remote Sens. Environ. 2020, 243, 111792. [Google Scholar] [CrossRef] [Scilit]
  101. Li, Y.; Niu, Z. Systematic method for mapping fine-resolution water cover types in China based on time series Sentinel-1 and 2 images. Int. J. Appl. Earth Obs. Geoinf. 2022, 106, 102656. [Google Scholar] [CrossRef] [Scilit]
  102. Li, Y.; Dang, B.; Li, W.; Zhang, Y. GLH-Water: A Large-Scale Dataset for Global Surface Water Detection in Large-Size Very-High-Resolution Satellite Imagery. Proc. AAAI Conf. Artif. Intell. 2024, 38, 22213–22221. [Google Scholar] [CrossRef] [Scilit]
  103. Wieland, M.; Fichtner, F.; Martinis, S.; Groth, S.; Krullikowski, C.; Plank, S.; Motagh, M. S1S2-Water: A Global Dataset for Semantic Segmentation of Water Bodies From Sentinel- 1 and Sentinel-2 Satellite Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 17, 1084–1099. [Google Scholar] [CrossRef] [Scilit]
  104. GB/T 21010-2017; Current Land Use Classification. China Standard Press: Beijing, China, 2017.
  105. GDPJ 01-2013; Content and Indicators of the National Geographic Survey. Office of the Leading Group for the First National Geographic Conditions Census of the State Council: Beijing, China, 2013.
  106. Shen, W.; Zhang, L.; Ury, E.A.; Li, S.; Xia, B.; Basu, N.B. Restoring small water bodies to improve lake and river water quality in China. Nat. Commun. 2025, 16, 294. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  107. Smith, S.; Renwick, W.; Bartley, J.; Buddemeier, R. Distribution and significance of small, artificial water bodies across the United States landscape. Sci. Total Environ. 2002, 299, 21–36. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  108. Ogilvie, A.; Belaud, G.; Massuel, S.; Mulligan, M.; Le Goulven, P.; Calvez, R. Surface water monitoring in small water bodies: Potential and limits of multi-sensor Landsat time series. Hydrol. Earth Syst. Sci. 2018, 22, 4349–4380. [Google Scholar] [CrossRef] [Scilit]
  109. Qin, P.; Cai, Y.; Wang, X. Small Waterbody Extraction With Improved U-Net Using Zhuhai-1 Hyperspectral Remote Sensing Images. IEEE Geosci. Remote Sens. Lett. 2021, 19, 5502705. [Google Scholar] [CrossRef] [Scilit]
  110. Farooq, B.; Manocha, A. Small water body extraction in remote sensing with enhanced CNN architecture. Appl. Soft Comput. 2024, 169, 112544. [Google Scholar] [CrossRef] [Scilit]
  111. Liu, B.; Zhao, Q.; Wang, C.; Li, M.; Xie, H.; Chen, L. African water body segmentation with cross-layer information separability based feature decoupling transformer. Int. J. Appl. Earth Obs. Geoinf. 2025, 142, 104741. [Google Scholar] [CrossRef] [Scilit]
  112. Ji, Y.; Wu, W.; Nie, S.; Wang, J.; Liu, S. Sea–Land Segmentation of Remote-Sensing Images with Prompt Mask-Attention. Remote Sens. 2024, 16, 3432. [Google Scholar] [CrossRef] [Scilit]
  113. Fu, Z.; Zhang, Z.; Liu, J.; Sun, Q. Optimizing dynamic thresholds with salt endmember constraints for lake mapping in arid regions. J. Arid Environ. 2025, 232, 105499. [Google Scholar] [CrossRef] [Scilit]
  114. Xu, J.; Xue, Y.; Liu, J.; Yin, W.; Li, P.; Zeng, Y.; Hou, H.; Varotsos, C.A. Urban Surface Water Extraction Based on Sentinel 2 Images. In IGARSS 2024—2024 IEEE International Geoscience and Remote Sensing Symposium; IEEE: New York, NY, USA, 2024; pp. 3286–3289. [Google Scholar]
  115. Jiang, L.; Zhou, C.; Li, X. Sub-Pixel Surface Water Mapping for Heterogeneous Areas from Sentinel-2 Images: A Case Study in the Jinshui Basin, China. Water 2023, 15, 1446. [Google Scholar] [CrossRef] [Scilit]
  116. Li, D.; Yu, D.; Xu, Y.; Jia, P.; Xue, W. Automatic Extraction Method of Green Tide Based on Mixed Pixel Decomposition Feedback Adjustment. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 18, 5975–5989. [Google Scholar] [CrossRef] [Scilit]
  117. Niroumand-Jadidi, M.; Vitti, A. Reconstruction of River Boundaries at Sub-Pixel Resolution: Estimation and Spatial Allocation of Water Fractions. ISPRS Int. J. Geo-Inf. 2017, 6, 383. [Google Scholar] [CrossRef] [Scilit]
  118. Sree, J.M.; R., C.D.; Elena, Z. Machine learning-based classification of lake ice and open water from Sentinel-3 SAR altimetry waveforms. Remote Sens. Environ. 2023, 299, 113891. [Google Scholar] [CrossRef] [Scilit]
  119. Vandaele, R.; Dance, S.L.; Ojha, V. Deep learning for automated river-level monitoring through river-camera images: An approach based on water segmentation and transfer learning. Hydrol. Earth Syst. Sci. 2021, 25, 4435–4453. [Google Scholar] [CrossRef] [Scilit]
  120. Zhao, B.; Sui, H.; Liu, J. Siam-DWENet: Flood inundation detection for SAR imagery using a cross-task transfer siamese network. Int. J. Appl. Earth Obs. Geoinf. 2023, 116, 103132. [Google Scholar] [CrossRef] [Scilit]
  121. Li, Y.; Zhou, P.; Wang, Y.; Li, X.; Zhang, Y.; Li, X. Deep Learning Small Water Body Mapping by Transfer Learning from Sentinel-2 to PlanetScope. Remote Sens. 2025, 17, 2738. [Google Scholar] [CrossRef] [Scilit]
  122. Yang, L.; Liu, P.; Zhang, G.; Zhao, H.; Zhao, C. Domain-Adaptive Segment Anything Model for Cross-Domain Water Body Segmentation in Satellite Imagery. J. Imaging 2025, 11, 437. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  123. Zhao, J.; Xiao, P.; Dong, Y.; Geiß, C.; Zhong, Y.; Taubenböck, H. Large-scale mapping of water bodies across sensors using unsupervised deep learning. Remote Sens. Environ. 2025, 328, 114877. [Google Scholar] [CrossRef] [Scilit]
  124. Xu, Y.; Lin, J.; Zhao, J.; Zhu, X. New method improves extraction accuracy of lake water bodies in Central Asia. J. Hydrol. 2021, 603, 127180. [Google Scholar] [CrossRef] [Scilit]
  125. Zhao, B.; Wu, J.; Chen, M.; Lin, J.; Du, R. Seasonally inundated area extraction based on long time-series surface water dynamics for improved flood mapping. ISPRS J. Photogramm. Remote Sens. 2024, 217, 32–52. [Google Scholar] [CrossRef] [Scilit]
  126. Wang, Y.; Liang, C.; Zhang, H.; Li, Q.; Huang, X. Integrating Large Language Models and Random Forest for Water-Ice-Snow Classification in Cold and Arid Region Lakes to Support Sustainable Water Management. Sustainability 2026, 18, 6209. [Google Scholar] [CrossRef] [Scilit]
  127. Li, J.; Li, L.; Song, Y.; Chen, J.; Wang, Z.; Bao, Y.; Zhang, W.; Meng, L. A robust large-scale surface water mapping framework with high spatiotemporal resolution based on the fusion of multi-source remote sensing data. Int. J. Appl. Earth Obs. Geoinf. 2023, 118, 103288. [Google Scholar] [CrossRef] [Scilit]
  128. Li, Z.; Li, R.; Zhang, Z.; Tian, F. Flood mapping in SAR images via threshold segmentation with hydrological-hydrodynamic modeling. J. Hydrol. 2025, 667, 134825. [Google Scholar] [CrossRef] [Scilit]
  129. Fu, D.; Jin, X.; Jin, Y.; Mao, X. Extraction of grassland irrigation information in arid regions based on multi-source remote sensing data. Agric. Water Manag. 2024, 302, 109010. [Google Scholar] [CrossRef] [Scilit]
  130. Szwarcman, D.; Roy, S.; Fraccaro, P.; Gíslason, Þ.E.; Blumenstiel, B.; Ghosal, R.; de Oliveira, P.H.; Almeida, J.L.d.S.; Sedona, R.; Kang, Y.; et al. Prithvi-EO-2.0: A Versatile Multitemporal Foundation Model for Earth Observation Applications. IEEE Trans. Geosci. Remote Sens. 2025, 64, 4400120. [Google Scholar] [CrossRef] [Scilit]
  131. Li, X.; Li, C.; Vivone, G.; Hong, D. SeaMo: A season-aware multimodal foundation model for remote sensing. Inf. Fusion 2026, 125, 103334. [Google Scholar] [CrossRef] [Scilit]
  132. Jakubik, J.; Yang, F.; Blumenstiel, B.; Scheurer, E.; Sedona, R.; Maurogiovanni, S.; Bosmans, J.; Dionelis, N.; Marsocci, V.; Kopp, N.; et al. TerraMind: Large-Scale Generative Multimodality for Earth Observation. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2025; pp. 7383–7394. [Google Scholar]
  133. Guo, X.; Lao, J.; Dang, B.; Zhang, Y.; Yu, L.; Ru, L.; Zhong, L.; Huang, Z.; Wu, K.; Hu, D.; et al. SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 27662–27673. [Google Scholar]
  134. Hong, D.; Zhang, B.; Li, X.; Li, Y.; Li, C.; Yao, J.; Yokoya, N.; Li, H.; Ghamisi, P.; Jia, X.; et al. SpectralGPT: Spectral Remote Sensing Foundation Model. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5227–5244. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Diagram of the filtering process for Web of Science and CNKI.
Figure 1. Diagram of the filtering process for Web of Science and CNKI.
Remotesensing 18 02972 g001
Figure 2. Water body segmentation-derived and variant models.
Figure 2. Water body segmentation-derived and variant models.
Remotesensing 18 02972 g002
Figure 3. Schematic diagram of the U-Net model structure [22].
Figure 3. Schematic diagram of the U-Net model structure [22].
Remotesensing 18 02972 g003
Figure 4. Schematic diagram of the DeepLabv3+ model structure [23].
Figure 4. Schematic diagram of the DeepLabv3+ model structure [23].
Remotesensing 18 02972 g004
Figure 5. Schematic diagram of the Transformer model structure [27].
Figure 5. Schematic diagram of the Transformer model structure [27].
Remotesensing 18 02972 g005
Figure 6. Schematic diagram of the Swin Transformer model structure [28].
Figure 6. Schematic diagram of the Swin Transformer model structure [28].
Remotesensing 18 02972 g006
Figure 7. Schematic diagram of the Swin Transformer block structure [28].
Figure 7. Schematic diagram of the Swin Transformer block structure [28].
Remotesensing 18 02972 g007
Figure 8. Mamba model structure diagram [28].
Figure 8. Mamba model structure diagram [28].
Remotesensing 18 02972 g008
Figure 9. Schematic diagram of the SAM structure [30].
Figure 9. Schematic diagram of the SAM structure [30].
Remotesensing 18 02972 g009
Figure 10. Schematic diagram of water body segmentation application areas.
Figure 10. Schematic diagram of water body segmentation application areas.
Remotesensing 18 02972 g010
Table 1. Comparison between existing review articles and our survey.
Table 1. Comparison between existing review articles and our survey.
ReferencesLiterature
Coverage Period
Model
Category
Remote Sensing
Modality
DatasetEvaluation
Metrics
Emerging Topics
Nagaraj et al. [14]2014–2023CNN, U-Net, FCN, DeepLab, PSPNet, etc.Focus on optical imageryPrimarily summarizes sensor specificationsEvaluation metrics were mentioned without systematic formulationMultisource data fusion, model interpretability, and mixed-pixel processing
Swati Gautam et al. [15]2012–2021CNN, FCN, U-Net, multiscale FCN, etc.Primarily optical imagery, with a small amount of SAR imageryLimited dataset coverageMathematical definitions were not systematically providedBuilding a large-scale labeled dataset, identification of small water bodies, and model interpretability
S. Rajeswari et al. [16]1994–2024CNN, U-Net, etc.Optical and radar imageryLists websites providing raw remote-sensing dataProvides a detailed list of 12 evaluation metricsMultisource data fusion, model interpretability, uncertainty quantification, and hybrid methods combining traditional indices with deep learning
Li et al. [17]2010–2022CNN, U-Net, DeepLab, etc.Optical and radar imageryNo dedicated dataset summaryMathematical definitions were not systematically providedMultisource data fusion, mixed-pixel extraction, and unified evaluation criteria
Bijeesh et al. [18]2010–2019FCN, etc.Optical imageryHighlights the parameters of satellite sensorsMathematical definitions were not systematically providedMultisource data fusion, mixed-pixel processing, and the use of motion profiles
Sigopi et al. [19]1990–2024CNN, FCN, etc.Optical and radar imageryLists satellite parameters and data-download linksNo mathematical formulas or standardized definitions of evaluation metrics were providedOptical–radar fusion, the Google Earth Engine cloud-computing platform, and integration of physical models with deep learning
Guo et al. [20]1991–2021CNN, FCN, U-Net, etc.SAR imageryNo dedicated dataset summaryNo mathematical formulas or standardized definitions of evaluation metrics were providedDevelopment and integration of network modules better suited to SAR water body extraction
This review2021–2025U-Net, DeepLabv3+, Transformer, Mamba, SAM, etc.Optical, SAR, and multimodal remote-sensing imageryProvides detailed information on spatial resolution, scene categories, numbers of images, download links, and other dataset characteristicsProvides mathematical definitions, interpretations, and applicability analyses of commonly used metricsSpatiotemporally adaptive multimodal fusion, segmentation strategies for large-model integration, fine-grained water body segmentation, and integration of explainability with physical mechanisms
Table 2. Primary architecture taxonomy based on dominant feature-modeling mechanisms.
Table 2. Primary architecture taxonomy based on dominant feature-modeling mechanisms.
ModelModeling
Mechanism
Representative
Models
Classification
Principles
CNNCentered on convolution operations, with an emphasis on extracting local spatial features and multiscale semantic featuresFCN, U-Net, DeepLabv3+, and their variantsModels that use convolution as their primary feature-modeling operator are classified in this category. U-Net and DeepLabv3+, as representative CNN architectures, are discussed separately.
TransformerCentered on self-attention mechanisms, with an emphasis on modeling global context and long-range dependenciesViT, Swin Transformer, SegFormer, and their variantsWhen self-attention is the primary feature-modeling mechanism and CNNs serve only as embedding modules or local auxiliary components, the model is classified in this category.
CNN–TransformerBy cascading, parallelizing, or fusing convolutional and self-attention modules across layers, the architecture simultaneously models local features and global dependenciesCNN encoder–Transformer modules, dual-branch networks, Transformer encoder–CNN decoder architectures, etc.When both CNNs and Transformers perform substantial feature-modeling functions, the model is classified as a hybrid architecture. U-shaped structures are treated only as auxiliary topologies.
MambaCentered on selective state space modeling and spatial sequence scanning, enabling the modeling of long-range dependenciesVision Mamba, VM-UNet, and other Mamba variantsModels that use state space modules, rather than self-attention, as their primary long-range modeling mechanism are classified in this category.
SAMRelies on large-scale pre-training, prompt-driven segmentation, and cross-task transfer capabilitiesSAM and its remote-sensing-adapted variantsModels centered on pre-trained foundation models and prompt mechanisms are classified in this category. Adapters or CNN modules are recorded as auxiliary components.
Table 3. Complementary classification based on data modality, supervision and adaptation regime, network topology, and application scenario.
Table 3. Complementary classification based on data modality, supervision and adaptation regime, network topology, and application scenario.
Classification
Dimensions
Main
Categories
Criteria for
Classification
Data ModalityOptical RGB imagery; panchromatic imagery; multispectral imagery; hyperspectral imagery; SAR imagery; aerial or UAV imagery; close-range imagery; multimodal imagery (e.g., optical–SAR); and multi-temporal imageryAnnotated according to the sensors, spectral bands, imaging platforms, and time-series types actually used in the study.
Supervision TypeFully supervised learning; semi-supervised learning; weakly supervised learning; unsupervised learning; self-supervised pre-training; transfer learning; domain adaptation; few-shot learning; zero-shot learning; and prompt-driven adaptationAnnotated according to the availability of training labels and the methods used for model pre-training, fine-tuning, and task adaptation.
Network TopologyFCN-style architecture; U-shaped encoder–decoder architecture; DeepLab-style dilated convolutional architecture; multi-branch architecture; cascaded architecture; parallel architecture; and prompt-based encoder–decoder architectureUsed to supplement the description of the model’s macro-level network organization; it is not the sole criterion for determining the primary architecture family.
Application ScenariosSegmentation of rivers, lakes, reservoirs, coastlines, and other open water bodies; extraction of flood inundation areasThe primary outputs are pixel-level water masks, waterlines, or water–land boundaries.
Table 4. Summary of improvements based on the U-Net model.
Table 4. Summary of improvements based on the U-Net model.
AuthorYearStructure TypeImproved StructureImprovement MethodApplication
Hu et al. [35]2025EncoderAttention mechanismIntroduction of channel attention mechanismMultispectral water body segmentation
Wang et al. [36]2025EncoderAttention mechanismIntegration of transformer multi-head attention mechanismFlood damage change detection in SAR images
Jonnala et al. [37]2025EncoderAttention mechanismIntroduction of self-attention mechanismWater body segmentation in Sentinel-2 satellite imagery
Zhang et al. [38]2025EncoderConvolution layerIntroduction of a Multiscale Expanding Convolution ModuleWater body segmentation in flooded wetlands
Xu et al. [39]2024EncoderConvolution layerIntroduction of rotation-invariant convolutionWater body segmentation in optical remote sensing images
Wang et al. [40]2024EncoderConvolution layerIntroduction of quaternion convolutionWater body segmentation in RGB satellite images
Wagner et al. [41]2023EncoderBackbone networkResNeXt 50Real-time river water level monitoring
Liu et al. [42]2025EncoderBackbone networkMobileNetV2Open-channel water level monitoring
Xia et al. [43]2021EncoderBackbone networkReduction in convolution kernel countSmall tributary water body segmentation
Bai et al. [46]2025Skip connectionsAttention mechanismIntroduction of window attention embedding moduleHigh-resolution image water body segmentation
Xing et al. [47]2025Skip connectionsAttention mechanismIntroduction of spatial and channel attention moduleSAR image water body segmentation
Zhang et al. [48]2023Skip connectionsAttention mechanismIntroduction of enhanced coordinate attention moduleRemote sensing image segmentation of aquaculture area
Xiang et al. [49]2023DecoderAttention mechanismIntroduction of channel and spatial attention moduleUAV image water body segmentation
Hertel et al. [50]2023DecoderConvolution layerBayesian convolution layerFlood submergence segmentation in SAR images
Li et al. [51]2021DecoderIntroducing a new moduleDenseBlock moduleHigh-resolution image water body segmentation
Ye et al. [52]2024DecoderConvolution layerDepthwise separable convolutionSAR image water body segmentation
Table 5. Summary of improvements based on the DeepLabv3+ model.
Table 5. Summary of improvements based on the DeepLabv3+ model.
AuthorYearStructure TypeImproved StructureImprovement MethodApplication
Luo et al. [59]2022EncoderASPP moduleStrip pooling moduleDiscrete distributed water body segmentation
Zhang et al. [60]2023EncoderASPP moduleASPP module’s dilation ratesRiver water body segmentation of high-resolution image
Sun et al. [61]2024EncoderASPP moduleDensely connected DASPP moduleWater body segmentation in complex urban environments
Huang et al. [62]2023EncoderBackbone networkMobileNetV2Segmentation of black and odorous water body
Xue et al. [63]2023EncoderBackbone networkMobileNetV2The division of urban water accumulation
Zhang et al. [64]2023EncoderBackbone networkResNet101Segmentation of duckweed-type black and odorous water body
Chen et al. [65]2023EncoderBackbone networkDual backbone networkSegmentation of water body highlight areas
Chen et al. [66]2025DecoderAttention mechanismIntroducing multi-level feature attention modulesSAR image water body segmentation
Lv et al. [67]2022DecoderUpsampling moduleIncreasing the number of upsampling operationsFlood monitoring and emergency management
Table 6. Summary of improvements based on Transformer and state space models.
Table 6. Summary of improvements based on Transformer and state space models.
AuthorYearModel TypeImprovement MethodApplication
Zhang et al. [71]2024CNN–TransformerIntroduced two key modules: MSSA and PDAMSea ice and water body segmentation
Zhang et al. [72]2023CNN–TransformerEmbedded MixFormer into U-Net and introduced the AMM moduleWater body segmentation in optical images
Zhao et al. [73]2024CNN–TransformerIntroduced the C-MLP module and superpixel segmentation moduleWater body segmentation in SAR images
Xiao et al. [74]2023CNN–TransformerDual-branch parallel encoder; introduced a deformable self-attention mechanismMultiscale and complex-background water body segmentation
Kang et al. [75]2023CNN–TransformerCross-level visual transformer; sub-pixel upsampling moduleWater body segmentation at different spatiotemporal scales
Tian et al. [76]2025CNN–TransformerDesigned AIEM, edge-enhancement module, and LSTM redundancy decay moduleSegmentation of small, narrow, and irregular water bodies
Zhang et al. [77]2023CNN–TransformerDesigned a multiscale transformation module and introduced a hybrid attention mechanismWater body segmentation in high-resolution remote sensing images
Zhong et al. [78]2022CNN–TransformerDesigned interference suppression and multi-level transformation modulesLake water body segmentation
Yang et al. [79]2024CNN–TransformerIntroduced GMAF, WFN, and EFA modulesWater body segmentation on the Qinghai–Tibet Plateau
Wang et al. [80]2025CNN–TransformerProposed the MCFF module; integrated the MAE mechanism; combined the model with GISAccurate water body segmentation in aerial images
Fan et al. [81]2023CNN–TransformerProposed a new STFF algorithm and designed a hybrid loss functionWater body segmentation in marine aquaculture areas
Fan et al. [82]2025CNN–TransformerDesigned a novel label generator; constructed a local feature fine-grained discriminative complementary moduleWater body segmentation in marine aquaculture areas
Table 7. Summary of improvements based on the Segment Anything Model.
Table 7. Summary of improvements based on the Segment Anything Model.
AuthorYearProblemImprovement MethodApplication
Moghimi et al. [86]2024The river water exhibits varying colors and texturesThe VIT-B backbone was introduced, and the SAM encoder was frozenRiver water body segmentation
Ma et al. [87]2023Issues of fragmentation and imprecise boundariesSAM was leveraged to generate two new concepts: SGO and SGBOptical image abstract semantic water body segmentation
Zhang et al. [88]2025Overreliance on accuracy labelsAn auto-clustering layer was designed, and the lightweight encoder was optimizedLarge-scale high-resolution image water body segmentation
Qiao et al. [89]2025Problems of fragmentation and imprecise boundariesA simple post-processing method and a new framework were proposedWater body segmentation under multiple land-cover types
Zou et al. [90]2025Precise classification and manual data annotationSAM and a scale-aware hybrid CNN–Transformer were usedCoastal wetland classification
Zhang et al. [91]2024Inability to achieve large-scale, high-precision recognitionA wetland prompt module, adaptation module, and CNNs were introducedMangrove wetland segmentation
Table 8. Comparative analysis of models for water body segmentation based on deep learning.
Table 8. Comparative analysis of models for water body segmentation based on deep learning.
Model
Categories
Main
Applications
Computational
Cost
Cross-Regional
Generalization Characteristics
Proposed External
Validation
U-Net and its variantsSuitable for small- to medium-sized datasets, boundary-detail restoration, and pixel-level segmentation of objects such as rivers and pondsClassic architectures are typically relatively lightweight; however, the actual computational cost increases significantly with encoder depth, the number of channels, and the number of attention modulesWithout large-scale pre-training or domain adaptation, the model is susceptible to variations in sensors, regions, and seasonsUse region-specific datasets and conduct independent testing across different regions, seasons, and sensor imagery
DeepLabv3+ and its variantsSuitable for water body scenes with significant scale differences, complex backgrounds, and a need for multiscale contextual modelingGenerally moderate to high, depending on the backbone network, output stride, and atrous spatial pyramid pooling configurationMultiscale features aid in scene adaptation but do not guarantee cross-regional generalizationReport performance variations across different geographic regions and spatial resolutions
Swin Transformer and its variantsSuitable for high-resolution imagery, scenes with large-scale differences, and scenarios requiring local-to-global contextual modelingWindow-based attention reduces the cost of global self-attention, but large-scale models and high-resolution inputs may still result in high video-memory requirementsLarge-scale pre-training may enhance transferability, but it remains constrained by domain shifts in remote sensing and the coverage of the training dataConduct transfer tests from seen to unseen regions and from single-sensor to heterogeneous-sensor settings
CNN–Transformer hybrid architecturesSuitable for complex water body scenes requiring both local boundary localization and global contextual modelingTypically higher than that of a single lightweight CNN, depending on the number of branches, feature-fusion operations, and the scale of the attention modulesThese architectures have the potential to improve robustness, but their generalization capability depends on fusion mechanisms, data diversity, and pre-training methodsConduct cross-region ablation comparisons with CNN and Transformer baselines of comparable model scale
Mamba or visual state space modelsSuitable for high-resolution imagery and scenarios involving river networks, shorelines, and continuous water bodies that require the modeling of long-range spatial dependenciesState space core computations scale linearly with sequence length; the overall cost depends on the two-dimensional scanning strategy, network scale, and decoder designThese models have potential for efficient long-range modeling, but evidence of their cross-regional applicability in water body remote sensing remains relatively limitedConduct unified comparisons with Transformers using data from different watersheds, sensors, and spatial resolutions
SAM modelsSuitable for prompt-driven segmentation, interactive annotation, sample generation, and human-assisted extraction in disaster-response scenariosThe original large-scale model has high computational and video-memory requirements; however, encoded features can be reused, and lightweight and adapted versions are availableSAM has potential for zero-shot transfer but is constrained by differences in scale, spectrum, target semantics, and domain-specific characteristics of remote-sensing imageryReport external regional performance separately under zero-shot, few-shot fine-tuning, and fully supervised conditions
Table 9. Evaluation risks and recommendations for improvement in water body segmentation studies.
Table 9. Evaluation risks and recommendations for improvement in water body segmentation studies.
IssuesPotential Sources
of Bias
Recommended
Practices
Spatial data leakageSlices from the same original image or adjacent regions are included in both the training and test sets, causing the model to exploit spatial similarities and overestimate its generalization performanceFirst partition the dataset at the level of raw imagery, scenes, watersheds, or administrative regions, and then perform cropping; if necessary, establish spatial buffers between the training and test sets
Temporal data leakageImages of the same area acquired on adjacent dates or at different stages of the same event are assigned to both the training and test setsPerform independent partitioning by year, season, or flood event to prevent temporally similar samples from being distributed across different datasets
Preprocessing leakageNormalization parameters are calculated, thresholds are selected, or hyperparameters are tuned using the entire datasetCalculate normalization statistics using only the training set; use the validation set for model selection and the test set solely for final evaluation
Annotation uncertaintyLabel noise is caused by blurred shorelines, mixed pixels, tidal variations, cloud shadows, and differences in human interpretationUse multi-rater review, soft labels, boundary exclusion zones, boundary-tolerance metrics, and label-quality grading
Reliance on internal testingReporting results only on a test set randomly divided from the same dataset fails to demonstrate the model’s true transferabilityIncrease external validation across different regions, sensors, seasons, and spatial resolutions
Single-precision metricsOverall accuracy or IoU may mask segmentation failures involving small rivers, boundaries, and rare scenariosReport IoU, F1-score, precision, recall, boundary metrics, and results for different target scales, together with confidence intervals
Table 10. Relationships among target problems, improvement strategies, applicable data types, and effects in water body segmentation.
Table 10. Relationships among target problems, improvement strategies, applicable data types, and effects in water body segmentation.
Target
Problem
Typical Application
Scenarios
Improvement
Strategy
Applicable
Data Type
Results
Significant variations in water body scaleAreas containing large lakes, rivers, and small pondsMultiscale context aggregation, feature pyramids, ASPP, and multiscale attentionHigh-resolution optical, multispectral, and SAR imageryImproves the consistency of water body detection across different scales and increases the recall rate for small-scale targets
Loss of small- and micro-scale featuresSmall ponds, minor tributaries, and scattered waterlogged areasShallow high-resolution feature preservation, deep supervision, boundary-aware loss, and fine-grained decodersAerial or UAV imagery and high-resolution satellite imageryHelps to restore small targets and local boundaries, thereby reducing false negatives
Fragmentation of narrow, elongated water bodiesNarrow rivers, ditches, and complex river networksLong-range modeling using Transformers or Mamba, direction-aware scanning, topological constraint loss, and connectivity constraintsOptical, multispectral, and SAR imageryImproves river continuity and reduces breaks and fragmentation
Interference from complex background features similar to water bodiesUrban shadows, dark roofs, roads, mountain shadows, and dark vegetationSpectral–spatial joint modeling, dual attention, hard-negative sample mining, and boundary enhancementOptical, multispectral, and hyperspectral imageryReduces false-positive detections caused by background features that are similar to water bodies
Clouds, shadows, noise, and poor imaging conditionsCloudy and rainy regions, flood events, and low-quality imageryOptical–SAR fusion, mask-aware feature weighting, denoising modules, and robust loss functionsOptical, SAR, and optical–SAR multimodal imageryEnhances model robustness under conditions involving missing or degraded data
Insufficient multimodal feature fusionJoint segmentation using optical imagery, SAR imagery, DEMs, and other auxiliary dataCross-modal attention, feature alignment, dual-branch fusion, and decision-level fusionOptical–SAR, optical–DEM, and other multisource dataFully exploits complementary information provided by different sensors and auxiliary data sources
Insufficient generalization across regions and sensorsModel transfer from training basins to unknown regions or imagery acquired by other sensorsDomain adaptation, domain generalization, sensor-invariant representations, style transfer, self-supervised pre-training, and test-time adaptationMulti-region, multi-season, and multi-sensor dataReduces performance degradation caused by regional differences and sensor-domain shifts
Temporal variations and real-time monitoring requirementsDynamic flood monitoring, seasonal changes in water bodies, and analysis of video or continuous imageryTemporal feature fusion, state space models, change detection, lightweight networks, and prompt-based temporal segmentationMulti-temporal satellite imagery, video, and continuous aerial imageryEnhances temporal consistency and improves the efficiency of dynamic water body updates
Ambiguous shorelines and annotation uncertaintiesTidal zones, wetlands, shallow-water areas, and mixed-pixel regionsSoft labels, boundary-tolerance loss, probabilistic segmentation, uncertainty estimation, and multi-annotator consistency analysisVarious types of remote-sensing imagery, particularly medium- and low-resolution dataMinimizes the influence of uncertain boundaries on model training and provides confidence estimates for segmentation predictions
Table 11. Basic characteristics of representative datasets for water body segmentation and related remote sensing applications.
Table 11. Basic characteristics of representative datasets for water body segmentation and related remote sensing applications.
DatasetMain
Categories
Sensor
Type
Data
Modality
Spatial
Resolution
CoverageSample
Size
GID
[92]
5 land-cover segmentation classes and 15 subclasses; water bodies are subdivided into rivers, lakes, and pondsGF-2 PMSPanchromatic and multispectral imagery1 m (panchromatic); 4 m (multispectral)More than 60 cities in China, covering an area exceeding 50,000 km2150 images of 6800 × 7200 pixels; 30,000 multiscale patches and 10 pixel-level validation images
EvLab-SS
[93]
11 high-resolution semantic segmentation classes, including water bodies, farmland, orchards, forested areas, and roadsWorldView-2, GeoEye, QuickBird, GF-2, and aerial platformsHigh-resolution optical and aerial imagerySatellite: 0.2, 0.5, 1, and 2 m; aerial: 0.1 or 0.25 mDerived from the National Survey and Mapping Project on China’s Geography60 images, averaging approximately 4500 × 4500 pixels; 48,622 training patches and 13,539 validation patches
DeepGlobe_Land
[94]
7 land-cover semantic segmentation classes; water bodies include rivers, oceans, lakes, wetlands, and pondsDigitalGlobe Vivid+RGB optical imagery0.5 mPrimarily rural scenes, covering an area of 1716.9 km21146 images of 2448 × 2448 pixels
WHU-OPT-SAR
[95]
7 land-use semantic segmentation classes: water bodies, croplands, urban areas, villages, forests, roads, and othersGF-1 and GF-3 FS IIRGB and NIR optical imagery; C-band SAR imagery5 mHubei Province, China (30°–33°N, 108°–117°E), covering 51,448.56 km2100 sets of paired optical–SAR images of 5556 × 3704 pixels
LoveDA
[96]
7 land-cover semantic segmentation classes divided into urban and rural areasGoogle Earth historical imageryRGB optical imagery0.3 m18 administrative districts in Nanjing, Changzhou, and Wuhan, China; imagery acquired in July 2016; total area of 536.15 km25987 images of 1024 × 1024 pixels
SEN12MS
[97]
Provides four MODIS land-cover labeling systems, including water bodies and wetlandsSentinel-1, Sentinel-2, and MODIS MCD12Q1Dual-polarization SAR, full multispectral imagery, and land-cover products10 mAll inhabited continents, covering four meteorological seasons180,662 registered triplets, each containing 256 × 256 pixels
China Lake
[98]
Long-term vector inventory of natural lakes in China (≥1 km2), excluding wetlands, reservoirs, and riversHistorical topographic maps; Landsat 1–4 MSS, Landsat 5 TM, Landsat 7 ETM+, and Landsat 8 OLIMultispectral optical imagery and historical mapsMSS: approximately 80 m; other Landsat imagery: 30 mEntirety of China; imagery acquired from the 1960s to 2015 at approximately 5-year intervalsMore than 3831 Landsat scenes; the 2015 dataset ultimately included 2407 lakes
GLAKES
[99]
Global lake and reservoir boundaries and water-probability-weighted areas for three periods, with a minimum area of 0.03 km2GSWOWater-presence probability raster and vector lake products30 m60°S–80°N; 1984–1999, 2000–2009, and 2010–2019Approximately 3.4 million lakes; 754 sample plots for the baseline model and 90,512 lake labels
GLAD-GSWD
[100]
Global monthly, seasonal, annual, and dynamic-type products for inland open water bodiesLandsat 5, Landsat 7, and Landsat 8Multispectral imagery and topographic auxiliary data30 mGlobal land areas excluding Greenland; imagery acquired from 1999 to 2018Approximately 3.4 million Landsat scenes were processed; 165, 164, and 120 manually classified scenes were used for training according to sensor type
CWaC
[101]
Six-category water-coverage mapping for China: rivers, lakes, reservoirs, aquaculture ponds, seasonal wetlands, and paddy fieldsSentinel-1 and Sentinel-2C-band SAR and multispectral imagery10 mEntire territory of China, covering an area of 240,383 km218,402 Sentinel-1 and 64,924 Sentinel-2 scenes; 7952 training samples and 5681 validation samples
GLH-water
[102]
Global ultra-high-resolution binary surface-water segmentation, including rivers, ponds, glacial lakes, and diverse background scenesGoogle EarthRGB optical imageryApproximately 0.3 mImagery acquired from 2011 to 2022, covering approximately 3686 km2250 large images; 156,250 non-overlapping tiles of 512 × 512 pixels
S1S2-Water
[103]
Binary semantic segmentation of normal water bodiesSentinel-1, Sentinel-2, and Copernicus DEMC-band SAR, multispectral imagery, and elevation data10 m29 countries; imagery acquired from 2018 to 2020, covering approximately 650,000 km265 data triplets; more than 50,000 training tiles and more than 25,000 tiles in each of the validation and test sets
Table 12. Annotation and usage characteristics of representative datasets for water body segmentation and related remote sensing applications.
Table 12. Annotation and usage characteristics of representative datasets for water body segmentation and related remote sensing applications.
DatasetAnnotation Sources
and Quality
Data
Uncertainty
Data
Split
Class
Balance
Licensing
and Access
GID
[92]
The classification system was established in accordance with GB/T 21010-2017 [104], and pixel-level labels are providedLarge-scale pixel-level class proportions are not reported, and no independent test set is available120 images in the training set and 30 images in the validation setThe training dataset consists of 15 classesLicense not specified; https://x-ytong.github.io/project/GID.html, accessed on 29 August 2026
EvLab-SS
[93]
Full-pixel annotation was conducted in accordance with Content and Indicators of the National Geographic Survey (GDPJ 01-2013) [105]Significant variations exist in multisensor spatial resolution and imaging conditions37 images in the training set, 8 images in the validation set, and 15 images in the test setSignificant class imbalance exists; the “Garden” category is completely absent from the validation setLicense not specified; https://github.com/EarthVisionLab/EVLab-SS-dataset?utm_source=chatgpt.com, accessed on 29 August 2026
DeepGlobe_Land
[94]
Pixel-level masks were created by professional annotators; annotation was required for instances larger than approximately 20 m × 20 mA small amount of human error is present; roads and bridges were intentionally left unlabeled803 images in the training set, 171 images in the validation set, and 172 images in the test setAgriculture: 56.76%; forest: 13.75%; pasture: 10.21%; urban: 9.35%; bare ground: 6.14%; water: 3.74%; unknown: 0.04%License not specified; https://www.kaggle.com/datasets/balraj98/deepglobe-land-cover-classification-dataset?utm_source=chatgpt.com, accessed on 29 August 2026
WHU-OPT-SAR
[95]
Labels were derived from the 2017 National Land-Use Change Survey vector data, rasterized and resampled to 5 m; optical and SAR images were registered at the sub-pixel levelResampling may introduce boundary errors, and SAR shadowing in mountainous areas remains a limitation17,640 images in the training set, 5880 images in the validation set, and 5880 images in the test setA high degree of category imbalance exists; the class proportions originally reported in the paper are not recommended because they have subsequently been officially correctedLicense not specified; https://github.com/AmberHen/WHU-OPT-SAR-dataset.git, accessed on 29 August 2026
LoveDA
[96]
Professional remote sensing annotators used polygons to annotate six foreground classesSignificant differences exist between urban and rural areas in class distribution, target scale, and spectral characteristics2522 images in the training set, 1669 images in the validation set, and 1796 images in the test setBackground pixels dominate, and the distributions of urban and rural categories are inconsistentLicense not specified; https://github.com/Junjue-Wang/LoveDA, accessed on 29 August 2026
SEN12MS
[97]
A total of 252 scenes were retained after visual quality inspection by remote sensing experts; MODIS MCD12Q1 labels include four classification systemsOverall accuracy across the label layers ranges from approximately 67% to 87%, making the dataset unsuitable for detailed water body boundary segmentationNo fixed official training, validation, or test set split is providedPixel-level category proportions are not reported, and the MODIS categories are inherently imbalancedCC BY 4.0; https://mediatum.ub.tum.de/1474000, accessed on 29 August 2026
China Lake
[98]
Following automatic NDWI extraction, lake-by-lake visual inspection and manual editing were conducted using historical lake inventories, online maps, and original Landsat imagerySmall lakes are systematically excluded; seasonal shoreline variations and the quality of early MSS imagery introduce uncertaintyThe dataset consists of cartographic products and provides no machine-learning training, validation, or test set splitThe number and area of lakes follow long-tailed distributions, and a minimum mapping unit of ≥1 km2 was adoptedData license not specified; use is subject to the terms of the National Qinghai–Tibet Plateau Scientific Data Center; https://data.tpdc.ac.cn/zh-hans/data/fa8426c0-d3f0-4615-8e78-0465a0957891/, accessed on 29 August 2026
GLAKES
[99]
Labels were manually revised; a river mask was overlaid after global prediction, and independent test labels were evaluatedThe dataset has a relatively high omission rate for small lakesThe five sample-area categories were allocated to the training, validation, and test sets using stratified random sampling at a ratio of 60%, 20%, and 20%, respectivelyBy quantity, small, medium, and large lakes account for 94.39%, 5.56%, and 0.05%, respectivelyCC BY 4.0; https://doi.org/10.5281/zenodo.7016548, accessed on 29 August 2026
GLAD-GSWD
[100]
A hierarchical ensemble tree was trained using manually and fully classified scenes from each Landsat sensorSignificant mixed-pixel effects occur at the 30 m spatial resolution; the resulting products are primarily applicable to open water bodiesNo conventional publicly disclosed ratios or quantities are provided for the training, validation, and test setsLand and permanent water bodies together account for approximately 98.6% of the global land areaCC BY 4.0; https://www.glad.umd.edu/dataset/global-surface-water-dynamics, accessed on 29 August 2026
CWaC
[101]
Samples were manually verified using high-resolution Google Earth imagery and time-series imagery; river discontinuities were manually corrected after automatic classificationThe annotation process was not fully automaticSampling points for rivers, lakes, reservoirs, and aquaculture ponds were allocated at a ratio of 2/3 for training and 1/3 for validation; the final training set contained 7952 points, and the six independent validation categories contained 5681 pointsThe training set covered only four classes with unequal sample sizes; the validation set contained 775–1062 points per class, and the mapped-area distributions were also unevenStandard license not specified; use is subject to the terms of Science Data Bank; https://www.scidb.cn/en/detail?dataSetId=951b7d1d22304c59ae8b9a91415e550c, accessed on 29 August 2026
GLH-water
[102]
Manual fine-tuning was conductedVariations in imaging time and geographic region within Google Earth imagery may still introduce boundary uncertaintyThe training, validation, and test sets were randomly divided from the original large-scale maps at a ratio of 80%, 10%, and 10%, respectivelyThe proportions of water body and background pixels were not reportedLicense not specified; https://jack-bo1220.github.io/project/GLH-water.html, accessed on 29 August 2026
S1S2-Water
[103]
An initial water body mask was generated using Sentinel-2 imagery and NDWI and was subsequently visually verified by three image-interpretation expertsSevere weather conditions are insufficiently represented, and the dataset covers only normal water body conditionsThe training, validation, and test sets were fixed using a 100 × 100 km scene-level divisionSampling ensured that each scene contained a minimum proportion of water bodiesCC BY 4.0; https://doi.org/10.5281/zenodo.8314175, accessed on 29 August 2026
Table 13. Water body segmentation evaluation metrics.
Table 13. Water body segmentation evaluation metrics.
CategoryNameFormulaFunction
Confusion MatrixTrue Positive (TP)Predicted as water, whereas the actual class is water
True Negative (TN)Predicted as background, whereas the actual class is background
False Positive (FP)Predicted as water, whereas the actual class is background
False Negative (FN)Predicted as background, whereas the actual class is water
Core MetricsPrecision TP TP + FP The proportion of correctly predicted water pixels among all pixels predicted as water
Pixel Accuracy TP + TN TP + FP + FN + TN The proportion of correctly classified pixels among all pixels
Mean Pixel Accuracy 1 N i = 1 N TP i TP i + FN i The average pixel accuracy calculated across all classes
Recall TP TP + FN The proportion of correctly predicted water pixels among all actual water pixels
IoU TP TP + FP + FN The ratio of the intersection to the union between the predicted and actual water regions
mIoU 1 N i = 1 N TP i TP i + FP i + FN i Measures the model’s average segmentation accuracy across all classes
Dice 2 TP 2 TP + FP + FN Measures the degree of overlap between the predicted and actual segmentation regions
F1-Score 2 × Precision × Recall Precision + Recall Determines the harmonic mean of precision and recall
Omission Error Rate FN TP + FN Measures the proportion of actual water pixels that are missed by the model
Commission Error FP TP + FP The proportion of false water predictions among all pixels predicted as water; it is equal to 1 Precision
Error Rate FP + FN TP + TN + FP + FN The proportion of incorrectly classified pixels among all pixels
False Positive Rate FP FP + TN The proportion of actual non-water pixels that are incorrectly classified as water
Kappa κ = p o p e 1 p e Evaluates agreement beyond chance and measures classification reliability, particularly for imbalanced datasets
Table 14. Comparison of the performance of representative deep learning models on common river water body segmentation datasets [86].
Table 14. Comparison of the performance of representative deep learning models on common river water body segmentation datasets [86].
DatasetModelPAKIoUPrecisionRecallF1-Score
Kaggle
Water Net
U-Net0.9550.8890.8740.9260.9190.920
DeepLabv3+0.9570.9000.8860.9210.9410.929
SAM0.9630.9140.9010.9530.9320.939
Elbersdorf
Wesenitz
U-Net0.9990.9980.9980.9990.9990.999
DeepLabv3+0.9980.9970.9970.9990.9980.998
SAM0.9940.9880.9880.9960.9910.994
RIWA.v1U-Net0.9690.9070.8780.9270.9420.930
DeepLabv3+0.9660.9010.8730.9290.9320.926
SAM0.9660.8910.8600.9430.8970.914
LuFI-
RiverSnap.v1
U-Net0.9610.9000.8900.9390.9420.934
DeepLabv3+0.9610.9090.8980.9510.9380.941
SAM0.9630.9310.9250.9850.9380.957
Note: The data presented in the table were obtained from [86], and bold values indicate the best results achieved by each model on the corresponding dataset.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Li, D.; Ye, F.; Huang, Y.; He, H. Deep Learning for Water Body Segmentation in Remote Sensing Imagery: A Review. Remote Sens. 2026, 18, 2972. https://doi.org/10.3390/rs18172972

AMA Style

Li D, Ye F, Huang Y, He H. Deep Learning for Water Body Segmentation in Remote Sensing Imagery: A Review. Remote Sensing. 2026; 18(17):2972. https://doi.org/10.3390/rs18172972

Chicago/Turabian Style

Li, Donglin, Famao Ye, Yingjie Huang, and Haiqing He. 2026. "Deep Learning for Water Body Segmentation in Remote Sensing Imagery: A Review" Remote Sensing 18, no. 17: 2972. https://doi.org/10.3390/rs18172972

APA Style

Li, D., Ye, F., Huang, Y., & He, H. (2026). Deep Learning for Water Body Segmentation in Remote Sensing Imagery: A Review. Remote Sensing, 18(17), 2972. https://doi.org/10.3390/rs18172972

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop