1. Introduction
Mangrove forests are highly valuable coastal wetland ecosystems that support biodiversity and store large amounts of blue carbon [
1]. They also provide important coastal protection services by reducing shoreline erosion, attenuating waves, and buffering coastal hazards [
2,
3]. Reliable information on mangrove distribution is therefore essential for conservation planning, restoration assessment, and coastal wetland management. Accurate large-scale mangrove mapping remains challenging because mangroves often occur as narrow and fragmented intertidal belts where tidal variation, mixed pixels, and spectral confusion with adjacent land-cover types can reduce classification consistency [
4]. These difficulties are further complicated by the spectral variability of mangrove canopies and surrounding coastal backgrounds [
5,
6]. For satellite-based national-scale mapping, a stable and representative annual image composite is therefore required before reliable mangrove delineation can be achieved.
These challenges are particularly evident along the coast of China. Chinese mangroves are mainly distributed in subtropical and tropical coastal regions, where they often occur as narrow belts, small patches, or restored stands along estuaries, bays, aquaculture ponds, tidal flats, and modified shorelines [
7]. Long-term monitoring studies have shown that mangrove loss and recovery in China have been strongly shaped by conservation, restoration, and coastal land-use change [
8]. Remote sensing studies have also reported that mangroves can be difficult to separate from salt marshes and other adjacent coastal land-cover types [
9,
10]. Such classification challenges are especially pronounced for fragmented shoreline patches and recently restored stands, and their severity varies among coastal environments. Residual clouds, haze, tidal exposure, water-background effects, and mixed pixels further reduce the temporal stability and spectral separability of mangrove signals [
4,
5,
6,
9]. For national-scale Sentinel-2 mapping, the methodological challenge is therefore not only classifier design, but also the construction of annual input images that consistently represent mangrove and non-mangrove coastal surfaces.
Annual image compositing is widely used to reduce the influence of clouds, atmospheric contamination, and short-term surface variability in optical remote sensing time series [
11]. For dense Sentinel-2 and Landsat observations, compositing algorithms can substantially affect the reflectance values retained in the final image and, consequently, the information provided to land-cover classifiers [
12,
13]. Quality-mosaic approaches select one observation per pixel according to a defined quality band, so residual haze, sun glint, turbidity, or unusual tidal conditions may also be retained [
11,
12,
13]. Median compositing can reduce these short-lived extremes but may attenuate temporally specific phenological or hydrological signals [
12,
13]. In addition, KNDVI-, EVI-, MFI-, and negative-NDWI-based rules emphasize different spectral characteristics [
14,
15,
16,
17], potentially affecting the contrast between mangroves and surrounding coastal classes. Mangrove-specific spectral indices, such as the mangrove vegetation index (MVI), have also been developed to enhance the spectral discrimination of mangrove vegetation [
18]. Such compositing-related differences can alter the spectral-spatial characteristics of the resulting annual image inputs and may subsequently influence deep learning-based mangrove segmentation [
19]. However, previous regional and national studies differed in sensors, classifiers, samples, compositing strategies, and geographic settings, making it difficult to isolate the effect of the compositing rule itself. A controlled comparison under a consistent segmentation and validation framework is therefore necessary.
Compositing rules do more than produce visually different annual images. They determine which observations are selected or summarized for each pixel and can therefore alter the spectral-spatial patterns learned by downstream classification models.
Deep learning models have shown strong potential for remote sensing image segmentation because they can learn nonlinear spectral-spatial features from complex image data [
19]. Recent mangrove mapping studies have also demonstrated the ability of deep learning models to delineate mangrove boundaries and fragmented patches at high spatial resolution [
19]. Nevertheless, model performance remains constrained by the quality, stability, and representativeness of the input image composites. A strong segmentation model cannot fully compensate for residual cloud contamination, tidal artefacts, water-background effects, or inconsistent spectral signals introduced during annual compositing. Deep learning-based mangrove mapping should hence be evaluated together with the compositing strategy used to construct the annual input image. Without such control, it is difficult to determine whether accuracy differences arise mainly from the classifier, the input composite, or their interaction.
Several global and regional mangrove products have demonstrated the feasibility of large-scale mangrove monitoring at 10 m or finer spatial resolution. Global Mangrove Watch version 3.0 provides a long-term global mangrove extent dataset from 1996 to 2020 [
20]. LREIS_GLOBALMANGROVE_v2 provides 10 m global mangrove classification products for 2018–2020 [
21]. HGMF_2020 maps global mangrove forest distribution at 10 m resolution [
22]. Despite these advances, local omissions and boundary inconsistencies may still occur along China’s fragmented and tidally dynamic shorelines, particularly around aquaculture ponds, narrow mangrove belts, and recently restored stands. Multi-source approaches that combine optical imagery, synthetic aperture radar, and multi-resolution features can improve mangrove classification under complex coastal conditions [
23,
24,
25]. However, these approaches often require additional data sources, more preprocessing steps, and higher computational cost. Sentinel-2 annual compositing remains an operationally attractive strategy for national-scale mangrove monitoring, but the effects of different compositing rules have not been sufficiently tested under a controlled deep learning framework.
To address this gap, this study evaluated five rule-specific Sentinel-2 annual composites for 10 m mangrove mapping in China. The compositing rule was treated as the main experimental variable, whereas the ResNet-34 U-Net architecture, labelled sample locations, validation data, probability threshold, and post-processing settings were kept consistent. The objectives were to (1) compare maximum kernel normalized difference vegetation index-based composite (MAX-KNDVI), maximum enhanced vegetation index (MAX-EVI), negative normalized difference water index (negative-NDWI), maximum mangrove forest index (MAX-MFI), and Median compositing for deep learning-based mangrove mapping; (2) evaluate the selected output against GMW v3.0, HGMF_2020, and LREIS_v2_2020 using spatially independent validation samples; and (3) apply the selected configuration to map annual mangrove distribution in China from 2019 to 2023 and summarize mapped area changes. By isolating the compositing rule within a shared segmentation framework, this study clarifies how annual image representation affects Sentinel-2 deep learning-based mangrove mapping and provides practical guidance for national-scale coastal wetland monitoring.
2. Materials and Methods
2.1. Coastal Study Domain, Reference Samples, and Ancillary Datasets
Mangrove forests in China are mainly distributed along the subtropical and tropical southeastern coastline, including Hainan, Guangdong, Guangxi, Fujian, Zhejiang, and Taiwan [
7]. These regions include estuaries, bays, deltas, lagoons, tidal flats, aquaculture ponds, and highly modified shorelines, which together create complex coastal backgrounds for mangrove mapping [
8]. In this study, mangrove forests refer to the woody mangrove vegetation component of coastal mangrove wetland ecosystems. The coastal study domain is shown in
Figure 1.
To define the modelling domain, existing mangrove products and field-surveyed mangrove features were merged to form an envelope of known mangrove occurrence. An isotropic 10 km buffer was then applied to this envelope to include representative coastal background classes that are spectrally or spatially confusable with mangroves, including tidal flats, aquaculture ponds, coastal evergreen vegetation, water bodies, and modified shorelines. The 10 km distance was selected as a practical compromise to retain representative coastal background variability while avoiding extensive inland and open-ocean areas unrelated to mangrove mapping. The buffer was used to define the modelling domain as distinct from an ecological influence distance of mangroves. The buffered domain was clipped by coastal administrative boundaries to retain coastal land–water transition zones while reducing open-ocean areas unlikely to contain mangroves. Buffering and area calculations were performed in projected coordinate systems to preserve distance and area fidelity.
Field reference data were collected from 34 coastal survey regions (
Figure 2). A detailed survey of mangrove distribution on Hainan Island had been conducted before this study [
26]. Additional field surveys were conducted between February and June 2021 in parts of Hainan, Guangxi, Guangdong, and Fujian. Handheld GPS devices and high-resolution Google Earth imagery were used to support field positioning and reference interpretation. Training samples were collected from five survey regions, yielding 695 labelled patch locations. Each input patch contained 12 Sentinel-2 spectral bands and had a size of 128 × 128 pixels. At the 10 m spatial resolution, each patch represented an area of approximately 1.28 km × 1.28 km, providing sufficient spatial context to capture fragmented mangrove patches and their surrounding coastal backgrounds while maintaining computational efficiency. External validation samples were obtained from the remaining 29 survey regions, which were not used for model training. A total of 10,354 reference points were used for external validation, including 5208 mangrove points and 5146 non-mangrove points. The separation between the training and validation survey regions was used to reduce the risk of inflated accuracy caused by spatially proximate samples. Taiwan was included in the mapping and area summaries but was not used for field sampling or model training; its results were therefore interpreted as an external application area instead of a field-validated region. Three existing mangrove products were used for product-level comparison. Global Mangrove Watch version 3.0 provides global mangrove extent information through 2020 [
20]. LREIS_GLOBALMANGROVE_v2 provides 10 m global mangrove classification products for 2018–2020 [
21]. HGMF_2020 provides a 10 m global mangrove forest map for 2020 [
22]. Because the field reference data mainly corresponded to 2020, the 2020 layers of these products were used for external comparison.
Ancillary climate and land-use datasets were used only for exploratory association analysis and were not used for model training. Annual maximum, minimum, and mean land surface temperature were derived from the MOD11A1.061 Terra LST product [
27]. Annual total precipitation was derived from the CHIRPS daily precipitation dataset [
28]. Annual cropland and built-up areas were summarized from the Esri 10 m annual land-use dataset [
29].
2.2. Sentinel-2 Annual Compositing Under Coastal Cloud and Water-Background Effects
Sentinel-2 surface reflectance imagery from 1 January 2019 to 31 December 2023 was obtained from the COPERNICUS/S2_SR_HARMONIZED collection in Google Earth Engine [
30]. Images with CLOUD_COVERAGE_ASSESSMENT values greater than 20% were excluded before annual compositing. Cloud- and cirrus-contaminated pixels were masked using the QA60 band, in which bits 10 and 11 indicate opaque clouds and cirrus clouds, respectively [
31]. This filtering and masking procedure was applied consistently to all compositing rules to maintain comparability among the rule-specific annual composites.
Twelve Sentinel-2 spectral bands were retained for classification: B1, B2, B3, B4, B5, B6, B7, B8, B8A, B9, B11, and B12. All selected bands were prepared on a common grid and exported at 10 m spatial resolution for patch-based segmentation. The spectral indices described below were used only as quality bands for selecting annual observations during compositing. They were not appended as additional input channels to the deep learning model.
Five annual compositing rules were designed to represent different annual surface conditions relevant to mangrove mapping. Four rule-specific composites were generated using the qualityMosaic() function in Google Earth Engine. For each pixel, qualityMosaic() selects the observation with the highest value of a user-defined quality band, and the corresponding Sentinel-2 spectral bands from that same observation are retained in the output composite. KNDVI [
16], an EVI-based quality band [
17], negative NDWI [
14], and MFI [
15] were used to generate the MAX-KNDVI, MAX-EVI, negative-NDWI-based, and MAX-MFI composites, respectively. The workflow used to generate the five rule-specific annual composites is shown in
Figure 3.
KNDVI was used to emphasize annual maximum canopy greenness [
16]. The EVI-based quality band was used to rank observations according to vegetation vigour and vegetation-background contrast [
17]. MFI was used to enhance red-edge and intertidal spectral characteristics associated with mangrove environments [
15]. For the negative-NDWI-based rule, negative NDWI was used as the quality band so that the selected observation corresponded to the lowest NDWI value within the annual image stack [
14]. This design was intended to reduce the influence of water-background and tidal-water conditions in coastal pixels.
An annual Median composite was also generated from all valid observations. Median compositing summarized the central tendency of valid annual reflectance observations for each spectral band and reduced the influence of short-lived spectral extremes caused by residual clouds, haze, water-surface glint, turbidity, or unusual tidal states. The five rule-specific annual composites were then evaluated under the same segmentation framework to determine how annual compositing-rule selection affected mangrove mapping performance. The quality bands and compositing rules are summarized in
Table 1.
2.3. Common Segmentation Framework for Rule-Specific Mangrove Segmentation
Mangrove mapping was formulated as a binary semantic segmentation task. MR_DLM_2020 was designed as a common segmentation framework for evaluating the influence of Sentinel-2 annual compositing rules on mangrove forest mapping. In this framework, the rule-specific annual composite was the factor varied during inference, whereas the network architecture, labelled sample locations, validation design, training strategy, probability threshold, and post-processing settings were kept consistent.
For model development, the same 695 labelled patch locations were sampled from each of the five rule-specific 2020 annual composites, including MAX-KNDVI, MAX-EVI, the NDWI-based composite, MAX-MFI, and Median composites. Each input patch contained 12 Sentinel-2 spectral bands and had a size of 128 × 128 pixels. The corresponding binary labels represented mangrove and non-mangrove classes. Before generating the pooled training dataset, the 695 labelled spatial locations were first divided into training (70%) and internal validation (30%) subsets. This spatial partition was then consistently applied to all five rule-specific composite datasets, ensuring that the same geographic locations were assigned to the same subset across different compositing strategies. The training patches from the five composites were pooled to train a single shared ResNet-34 U-Net model, thereby avoiding additional variability from independently trained rule-specific models. After training, the model weights were fixed and applied separately to each composite, so the comparison reflects relative performance under a common model rather than rule-specific model optimization.
The segmentation model followed a ResNet-34 U-Net architecture. ResNet-34 was used as the encoder to extract hierarchical spectral-spatial features, and the U-Net decoder restored spatial detail for pixel-level prediction [
32,
33]. All input bands were normalized using statistics calculated from the training data. During prediction, patch-level probability maps were merged to reconstruct tile-level mangrove maps.
The MR_DLM_2020 model was implemented using the PyTorch 2.13.0 deep learning framework and trained on a workstation equipped with an Intel i9-13900HX CPU and an NVIDIA GeForce RTX 4060 Laptop GPU with 8 GB of VRAM. The model was optimized using the Adam optimizer for 1000 epochs [
34]. To reduce overfitting and alleviate class imbalance between mangrove and non-mangrove pixels, Mixup augmentation and weighted cross-entropy loss were applied during training [
35]. Mixup generated augmented training samples by linearly combining two image patches and their labels, whereas weighted cross-entropy reduced the dominance of the non-mangrove background class during optimization.
The core formulation for this approach is provided in Equations (1)–(4).
where
,
and
,
represent two randomly selected input samples and their corresponding labels, respectively. The mixed sample is
, and the corresponding label is
. where
is sampled from a Beta distribution,
, and
is the shape parameter.
where
is the total weighted Cross-Entropy loss, which is the average of the weighted Cross-Entropy loss over all pixels.
represents the total number of samples, i.e., the total number of pixels in the image.
represents the index of a pixel,
,
represents the number of classes in the image classification task.
denotes the weight of class
, which is primarily used to address the class imbalance problem and is calculated based on the pixel frequency of each class.
represents the true label indicating
whether the pixel belongs to class
, it is 1 if the pixel belongs to class
, otherwise it is 0.
denotes the output probability that pixel
belong to
.
After training, the same MR_DLM_2020 model was applied separately to each rule-specific 2020 annual composite, generating MR_2020_KNDVI, MR_2020_EVI, MR_2020_NDWI, MR_2020_MFI, and MR_2020_Median outputs for controlled comparison (
Figure 4). The final layer produced mangrove probability maps. A fixed probability threshold of 0.5 was applied to all rule-specific outputs and all annual maps. Although rule-specific thresholds might improve the accuracy of individual composites, a common threshold was used to maintain a controlled comparison across compositing rules and years and to avoid introducing rule-specific threshold optimization as an additional variable. Thus, the comparison reflects the relative performance of different composites under the same decision threshold in place of their maximum performance after separate threshold optimization (
Figure 5).
2.4. Independent Validation, Temporal Mapping, and Exploratory Association Analysis
Internal validation was conducted using the validation subset reserved during model training. Precision, recall, -score, and intersection over union (IoU) were calculated at the pixel level to evaluate segmentation performance. Precision measured the reliability of predicted mangrove pixels, recall measured the completeness of detected reference mangrove pixels, score balanced precision and recall, and IoU measured the spatial overlap between predicted and reference mangrove pixels. These metrics were used to assess the behaviour of the shared segmentation model during model development.
Where
,
,
, and TN denote true positives, false positives, false negatives, and true negatives, respectively.
External validation samples were obtained from the remaining 29 survey regions that were not used for model training. Reference points were selected and interpreted using field-survey information, handheld GPS records, and high-resolution Google Earth imagery. A survey-region-based spatially independent validation design was adopted, resulting in 10,354 reference points, including 5208 mangrove and 5146 non-mangrove points. The separation between the training and validation survey regions was used to reduce the risk of inflated accuracy caused by spatially proximate samples. The same set of validation points was used to evaluate all five rule-specific outputs and the three existing mangrove products. Confusion matrices were generated for each output, and producer accuracy (PA), user accuracy (UA), overall accuracy (OA), and the Kappa coefficient were calculated. The confusion matrix used for external validation is shown in
Table 2.
For the mangrove class,
and
were calculated as:
For the non-mangrove class,
and
were calculated as:
The expected agreement was calculated as:
The Kappa coefficient was calculated as:
The best-performing compositing rule was selected for annual mangrove mapping from 2019 to 2023. Because Google Earth imagery may not correspond exactly to the target classification year, UAV orthophotos were used as supplementary validation for selected mangrove patches detected by the model but absent from existing products. UAV surveys were conducted using a DJI Mavic 3 Pro between July and October 2023, and the orthophotos were compared with the 2023 model output to assess the delineation of selected small mangrove patches.
Annual mangrove maps were converted to polygons for area and patch-size analysis. The spatial data were stored in WGS-84 for display, whereas area statistics were calculated within the corresponding UTM zones to reduce area distortion. Contiguous mangrove patches were grouped into six size classes: 0–5, 5–50, 50–100, 100–200, 200–300, and >300 ha. These statistics were used to describe changes in mapped area and patch-size distribution, but they were not interpreted as a complete fragmentation assessment.
Finally, exploratory correlation analysis was conducted to examine broad linear associations between mapped mangrove area and selected climatic and anthropogenic variables. Annual total precipitation, annual mean land surface temperature, annual maximum land surface temperature, annual minimum land surface temperature, cropland area, and built-up area were aggregated to 10 km grid cells together with mapped mangrove area. Pearson correlation coefficients were calculated between mapped mangrove area and each selected variable:
where
is the covariance between variables
and
,
and
are their standard deviations. These correlations were interpreted as exploratory linear associations as opposed to evidence of causal drivers. Nonlinear effects, interaction effects, and spatially explicit driver attribution were not considered in this analysis.
4. Discussion
4.1. Effects of Compositing Rules on Mangrove Mapping
The controlled comparison shows that Sentinel-2 annual compositing-rule selection is an important source of variation in deep learning-based mangrove mapping. Because the model architecture, training samples, validation points, threshold, and post-processing settings were kept consistent, differences among the rule-specific outputs mainly reflect differences in how each annual composite represented the coastal surface. Together, these results support the central premise of the study: image compositing should be treated as an experimental variable, not merely as a neutral preprocessing step. Similar concerns have been raised in studies of maximum-value and median compositing, where the choice of temporal compositing rule can substantially alter the spectral signal passed to classification models [
11,
12,
13].
The superior performance of MR_2020_Median in this study can be explained by the coastal conditions under which mangroves were mapped. Maximum-quality-band composites select observations with extreme quality-band values from the annual image stack. This strategy can enhance vegetation greenness, water-background contrast, or intertidal spectral responses, but it can also retain observations affected by residual cloud, haze, glint, turbidity, or unusual tidal exposure. In fragmented mangrove belts, these short-lived extremes may weaken spectral continuity and increase boundary uncertainty. By contrast, median compositing suppresses transient spectral extremes and provides a more conservative annual representation of the coastal surface. Therefore, the result should not be interpreted as evidence that median compositing is universally optimal; rather, under the national-scale Sentinel-2 setting tested here, it provided the most balanced trade-off between omission and commission errors.
The comparison with GMW v3.0, HGMF_2020, and LREIS_v2_2020 further indicates that compositing design can influence product-level agreement, especially in narrow and fragmented coastal zones. MR_2020_Median showed higher agreement with independent validation samples and reduced omissions in selected examples. The UAV comparison also indicated that several small mangrove patches absent from existing products were detected by the selected configuration. These findings are particularly relevant for China, where mangroves often occur as discontinuous belts, small restored patches, or intertidal stands adjacent to aquaculture ponds, mudflats, and evergreen vegetation. In such settings, improving the stability of annual input composites can be as important as improving the classifier itself.
A common ResNet-34 U-Net was used to isolate compositing-related differences while limiting additional model-related variability, thereby allowing for the evaluation of compositing rules under a unified segmentation framework. This architecture was selected because it provides a stable and computationally practical framework for patch-based semantic segmentation [
32,
33]. More complex attention-based or Transformer-based models may further improve feature learning, but they would introduce additional model-related variables and require larger training datasets. Likewise, optical-SAR fusion and multi-source feature learning can improve mangrove classification in complex coastal environments [
23,
24,
25], but they also increase preprocessing and computational requirements. The present design therefore emphasizes methodological control: the five outputs were generated by varying the annual composite while keeping the trained model and inference settings fixed.
4.2. Mangrove Changes and Regional Differences from 2019 to 2023
The mapped mangrove forest area in China increased from 2019 to 2023. This pattern is consistent with previous observations that mangrove conservation and restoration have contributed to partial recovery along parts of the Chinese coast [
7,
8]. However, the mapped increase should be interpreted cautiously. The annual maps quantify changes in mapped forest extent, not necessarily changes in stand quality, species composition, survival rate, or ecological function. Some increases may reflect restoration and planting, whereas others may be influenced by boundary uncertainty, tidal exposure, or improved detection of small patches. For comparison, recent national statistics reported approximately 30,300 ha of mangrove land in China [
36], while a previous 10 m national-scale remote-sensing study estimated approximately 21,148–24,801 ha [
37]. Our 2023 estimate of 24,631.62 ha is lower than the official statistic but comparable in magnitude to previous high-resolution remote-sensing estimates. Because the geographic and thematic definitions are not fully equivalent, these values should be compared cautiously. Differences among estimates may reflect variations in mangrove definition, mapping resolution, classification thresholds, tidal and mixed-pixel treatment, and minimum mapped patch size.
Patch-size statistics suggest that changes in mangrove spatial pattern were more complex than total area change alone. The number of small patches increased during the study period, especially in the 0–5 ha and 5–50 ha classes. This pattern may indicate newly detected or recently established small stands, but patch counts alone cannot support a formal conclusion about fragmentation. A rigorous fragmentation assessment would require additional landscape metrics, such as patch density, mean patch size, edge density, core area, and connectivity, together with uncertainty estimates for patch boundaries.
The regional differences observed in this study likely reflect the combined influence of restoration effort, local geomorphic conditions, available intertidal space, hydrological connectivity, and coastal land-use pressure. Guangxi contributed the largest mapped area increase, whereas Guangdong showed stronger interannual fluctuation. These differences suggest that national-scale mangrove monitoring should be interpreted together with regional coastal settings and management context. Area expansion is encouraging, but conservation assessment should also consider whether mapped patches are hydrologically connected, structurally mature, and resilient to disturbance.
4.3. Exploratory Associations with Environmental and Anthropogenic Variables
The relationships between mapped mangrove area and selected environmental and land-use variables should be interpreted as exploratory linear associations as opposed to driver attribution. Pearson correlation does not establish causality, and all coefficients in this study were small. The weak negative associations with land surface temperature may reflect temperature stress or regional climatic gradients, whereas the weak positive association with precipitation may indicate generally favourable conditions in humid coastal environments. These weak relationships may also reflect spatial-scale mismatch, aggregation effects, nonlinear responses, and regional heterogeneity that are not captured by simple bivariate correlations.
The weak negative associations with cropland and built-up area may indicate potential pressure from coastal land use, but they do not provide direct evidence of habitat loss. Mangrove dynamics are shaped by interactions among tidal exchange, sediment supply, salinity, geomorphology, protection status, restoration practices, and local disturbance. Therefore, the correlation analysis provides context for interpreting broad spatial patterns but does not replace spatially explicit change analysis or field-based ecological assessment.
Future driver analysis would benefit from longer time series, restoration project records, coastal development data, hydrodynamic variables, and nonlinear modelling approaches. Generalized additive models, tree-based methods, or hierarchical spatial models could be used to examine thresholds, interactions, and regional heterogeneity. Such approaches would be more appropriate for identifying potential drivers than simple bivariate correlations over a short five-year period.
4.4. Limitations and Uncertainties
Several sources of uncertainty remain. First, although the training and validation samples covered multiple coastal regions and environmental settings, some conditions were still underrepresented, including extremely turbid waters, dense evergreen coastal vegetation, unusual phenological patterns, and highly disturbed urban shorelines. The lower performance under spatially independent external validation likely reflects the greater difficulty of transferring the model to geographically distinct coastal regions, while spatial correlation among pixels within the internal validation patches may have contributed to relatively optimistic internal metrics. Direct transfer of the model to other mangrove regions therefore requires further independent validation.
Second, classification uncertainty is concentrated along mangrove boundaries and fragmented belts. At 10 m resolution, mixed pixels are common along mangrove-water and mangrove-upland edges. Tidal stage, water turbidity, salt marshes, aquaculture ponds, and evergreen vegetation can all reduce spectral separability. Although the Median composite improved mapping stability in this study, it may also smooth short-lived seasonal or hydrological signals that are useful in some local settings. Future work should include seasonal composites, uncertainty layers, tide-level information, and higher-resolution validation data to better quantify boundary uncertainty.
Third, this study focused on the controlled comparison of Sentinel-2 annual compositing rules as opposed to feature-level ablation, adaptive rule weighting, or multi-source data fusion. Because the shared model was trained using pooled rule-specific patches, the results evaluate rule-specific annual inputs under a common model, but they do not quantify the independent contribution of each spectral band or index-derived feature. Future work should test adaptive compositing, rule weighting, SAR-optical integration, improved cloud and shadow masking, and broader independent validation to determine whether the framework can be generalized beyond the sampled coastal settings.
5. Conclusions
This study evaluated five Sentinel-2 annual compositing rules for 10 m national-scale mangrove forest mapping in China within a shared ResNet-34 U-Net framework. By keeping the model architecture, training samples, validation points, probability threshold, and post-processing settings consistent, the study isolated compositing rule as the main experimental variable. The controlled comparison showed that annual compositing-rule selection substantially affected deep learning-based mangrove segmentation. Among the tested rule-specific outputs, MR_2020_Median produced the most balanced result, with an overall accuracy of 90.1% and a Kappa coefficient of 0.801.
MR_2020_Median showed higher agreement with spatially independent validation samples than GMW v3.0, HGMF_2020, and LREIS_v2_2020, and reduced omissions in selected fragmented coastal zones. UAV orthophotos further supported the detection of several small mangrove patches absent from existing products. When the selected Median configuration was applied to annual Sentinel-2 composites from 2019 to 2023, the mapped mangrove forest area in China increased from 22,731.76 ha to 24,631.62 ha, with Guangxi contributing the largest increase. Exploratory correlations with climatic and land-use variables were weak and should be interpreted as contextual associations as opposed to causal evidence.
These results show that annual compositing design should be evaluated explicitly in Sentinel-2 deep learning-based mangrove mapping. Median compositing provided a practical and stable input strategy under the specific model, threshold, spatial resolution, and coastal conditions tested in this study; its relative advantage may differ under other experimental settings or geographic regions. Remaining uncertainties are mainly related to mixed pixels, tidal variation, cloud residuals, fragmented boundaries, and underrepresented coastal settings. Future work should incorporate more independent validation data, uncertainty mapping, seasonal or tide-aware composites, and broader tests beyond the sampled coastal regions.