Next Article in Journal
M-FSAD-KD: Full-Link Multi-Granularity Distillation for SAR Object Detection
Previous Article in Journal
Ambulance STARS: A Satellite-Driven Framework for Rapid Flood Impact Assessment and Time-Critical Ambulance Routing
Previous Article in Special Issue
Forest Soil Moisture Monitoring Using L-Band Passive Microwave and Machine Learning
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Spatial Domain Dependence Evolution of Input Parameter Importance in Soil Moisture Retrieval Under the XGBoost and SHAP Framework

1
School of Hydrology and Water Resources, Nanjing University of Information Science and Technology, Nanjing 210044, China
2
Key Laboratory of Hydrometeorological Disaster Mechanism and Warning of Ministry of Water Resources, Nanjing University of Information Science and Technology, Nanjing 210044, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(17), 3006; https://doi.org/10.3390/rs18173006
Submission received: 10 July 2026 / Revised: 31 August 2026 / Accepted: 1 September 2026 / Published: 4 September 2026

Highlights

What are the main findings?
  • Dominant predictors for soil moisture retrieval shift significantly from global to local spatial domains.
  • Inherent physical feature correlations tend to remain stable across spatial domains, while cross-category covariations vary spatially.
What are the implications of the main findings?
  • In global domain soil moisture modeling, ERA5-Land soil moisture should be retained as the core input, while local fine predictors can be simplified to minimize data redundancy.
  • Local domain soil moisture retrieval models require domain-adaptive feature screening prioritizing seasonal, topographic, and soil texture variables to enhance local retrieval accuracy.

Abstract

Soil moisture (SM) is a core state variable of terrestrial hydrological and land–atmosphere interactions. Machine learning-based downscaling and retrieval frameworks that integrate multi-source remote sensing and auxiliary datasets have become mainstream approaches for generating high-spatial-resolution SM products. However, the spatial domain dependence evolution law governing the relative importance of these multiple predictors remains insufficiently quantified. This study constructs an integrated XGBoost and SHAP interpretability framework to reveal how the contribution and driving mechanisms of predictors shift across spatial domains. We compiled global in-situ SM observations from 24 international soil moisture network (ISMN) monitoring networks spanning 2017–2024. Predictors were classified into five categories: Sentinel-1A radar backscatter, vegetation indices, ERA5-Land meteorological forcing, topographic geospatial variables, and static soil texture attributes. Two modeling approaches were adopted: independent local network models representing the regional domains and a unified composite model representing the global domain. Model performance metrics demonstrate that the global composite model yields robust generalization with minimal overfitting, while individual regional models exhibit domain-specific retrieval differences due to varying land surface conditions. Pearson correlation analyses confirm that physical covariances remain nearly consistent across different domains, whereas cross-category correlations vary with spatial domain and local landscape backgrounds. SHAP-based feature importance quantification reveals clear domain-dependent differences among dominant predictors. This work quantitatively verifies the spatial domain dependence of parameter importance in machine learning-based SM retrieval, providing guidance for domain-adaptive predictor selection and interpretable high-resolution SM modeling under diverse land surface conditions.

1. Introduction

Soil moisture (SM) is an important state variable in the terrestrial water cycle and plays a critical role in land–atmosphere interactions [1,2]. Due to the strong spatiotemporal variability of SM, traditional point-based monitoring techniques struggle to capture its spatial patterns across extensive regions over long time series [3,4]. Remote sensing, depending on its special observation ability at a large scale, has shown remarkable advantages for SM mapping [2]. Observations from both microwave and optical remote sensing have been widely used to obtain SM information across broad spatial and temporal extents [5]. However, the global and regional SM products retrieved from current satellite remote sensors generally have coarse spatial resolutions [6], which restricts their applications at fine spatial domains. For this reason, generating high-resolution SM datasets has emerged as a prominent yet challenging research focus within this discipline.
To meet the increasing demand for high-resolution SM data, numerous algorithms have been developed by relevant researchers to fulfill this demand [7]. One mainstream approach directly retrieves fine-resolution SM from high-spatial-resolution satellite measurements with radiative transfer models. Fan et al. [8] derived global 1 km SM products from Sentinel-1 observations using a semi-empirical water cloud model. This method requires relatively few auxiliary datasets, but its accuracy heavily relies on well-calibrated radiative transfer models with robust simulation performance. Spatial downscaling represents another effective strategy for generating fine-resolution SM products [4]. This technique is built on the surface correlation between coarse- and fine-resolution remote sensing retrievals. Researchers establish scale transformation functions for SM using visible bands, thermal infrared signals, or vegetation-related auxiliary variables, which compensate for the spatial detail loss inherent to coarse microwave observations [9,10,11,12]. The core principle lies in mitigating surface heterogeneity errors caused by scale discrepancies through covariates such as vegetation and land surface temperature, thereby refining coarse-resolution SM products into high-spatial-resolution datasets. Based on the thermal inertia theory, Fang et al. [9] constructed an SM downscaling framework to disaggregate the 9 km SMAP enhanced L2 radiometer SM product down to 1 km. Using machine learning methods, Zheng et al. [10] downscaled CCI SM data to a 1 km grid by integrating International Soil Moisture Network (ISMN) ground measurements and high-resolution optical imagery. The third category employs machine learning or deep learning approaches [1] to directly integrate multi-source high-resolution satellite observations or auxiliary datasets to retrieve 1 km SM [2,5]. Adopting a random forest model, Han et al. [13] generated daily global topsoil SM at 1 km resolution by combining meteorological forcing data, static soil physical parameters, and multi-source optical remote sensing records. Following a comparable modeling strategy, Li et al. [14] produced daily 1 km SM products with inputs including ERA5-Land reanalysis, leaf area index (LAI), land cover, digital elevation model (DEM) and soil property datasets. Using the XGBoost framework, Zhang et al. [15] constructed a spatiotemporally continuous, global daily dataset at 1 km resolution by merging GLASS products, ERA5-Land reanalysis datasets, environmental auxiliary datasets, and in situ SM observations from ISMN. All three of the above strategies are extensively adopted to produce the fine-resolution SM datasets for regional and global research applications.
Nearly all the SM retrieval methods outlined above integrate multi-source datasets as model inputs [16,17,18]. Such input variables fall into two distinct categories based on temporal characteristics: dynamic predictors and static predictors. Dynamic predictors include high-resolution visible, thermal infrared and active microwave remote sensing imagery, resampled atmospheric reanalysis products, and other auxiliary datasets. Static predictors consist of geographic coordinates, topographic metrics, land cover, and inherent soil attributes. Sentinel-1, Sentinel-2, and MODIS retrievals serve as primary high-resolution data sources, which can capture fine domain spatial variations in land surface conditions. ERA5 or ERA5-Land reanalysis datasets are widely incorporated as background meteorological inputs, supplying variables such as surface SM, land surface temperature, and precipitation. Coarse-resolution SM products are also occasionally incorporated into model frameworks. Furthermore, soil texture metrics (e.g., sand fraction, clay fraction) and bulk density are standard proxies to characterize soil physical properties. Collectively, these predictive covariates have become standard inputs for regional and global SM estimation. Nevertheless, the relative contribution of each individual predictor to the final SM retrieval outputs has not yet been systematically quantified.
Three mainstream frameworks exist to quantify the relative influence of input predictors [19]: built-in feature importance from tree-based ensembles, model-agnostic interpretability techniques (e.g., permutation importance, SHAP), and conventional statistical decomposition approaches. Tree-based models including random forest, XGBoost, LightGBM and CatBoost natively support feature importance scoring [20,21,22]. Random forest supplies two standard metrics, namely, Gini feature importance and permutation importance, rendering it the most prevalent tool for SM retrieval studies in remote sensing [20]. Gradient boosting frameworks such as XGBoost output gain-based importance scores, which can quantify the total contribution ratio of each predictor to SM fitting [21]. A key strength of all tree-based ensemble methods is their computational efficiency. They produce importance rankings directly after training without the need for auxiliary models and minimal extra computation. Model-agnostic global interpretability techniques form a second group, which are compatible with any machine learning or deep learning architecture. For instance, the permutation importance approach disrupts the values of a single variable and measures the resulting decline in model accuracy [23]. A larger accuracy drop indicates a higher contribution of that predictor, offering an intuitive physical interpretation. Based on game theory, SHAP quantifies positive and negative contributions for each individual sample [24]. It generates a global average importance ranking and differentiates linear versus nonlinear predictor responses, establishing it as the leading interpretability method within contemporary SM downscaling literature. Local interpretability models focus on local contributions at individual pixels, and are suitable for analyzing typical sampling points in small study areas, yet they perform less effectively than SHAP in global contribution assessment. Finally, classical statistical decomposition workflows can be paired with machine learning outputs. Variance decomposition and hierarchical partitioning (H-statistic) separately quantify each variable’s independent and interactive effects [25]. Partial dependence plots (PDP) and individual conditional expectation (ICE) curves visualize how model outputs shift alongside changing predictor values [22,26], offering indirect insights into variable influence. Standardized multiple linear regression coefficients serve only linear modeling contexts, where absolute coefficient magnitudes signal variable importance. However, this approach yields significant bias when applied to nonlinear datasets. Across all interpretability tools outlined above, SHAP remains the dominant choice within SM retrieval research for the comparative analysis of multi-source covariate weights.
Most existing machine learning-based SM retrieval studies adopt a fixed spatial resolution, ignoring how spatial differences modulate the contribution of model predictors. To address this gap, this work couples an XGBoost regression with the SHAP interpretability framework. Two modeling schemes are established, namely, independent site-specific models and a unified global model. Ultimately, this study aims to quantitatively reveal how the contribution magnitudes and driving mechanism of input predictors (including vegetation, radar, topography, and meteorological variables) for SM retrieval evolve across a varying spatial domain.

2. Materials

2.1. ISMN SM Networks

The ISMN serves as a globally accessible centralized repository for in situ SM measurements [27]. In this study, we utilized in situ SM data from the ISMN spanning the period from 2017 to 2024, which are treated as ground truth SM values for the corresponding 1 km pixels of Sentinel-1A. Our compiled database comprises 24 monitoring networks distributed worldwide (Table 1 and Figure 1), spanning a wide range of climatic regimes, terrain types, land cover types and latitudinal gradients. Only in situ observations classified as “G” (good) quality from the top soil layer (≤5 cm) were retained for our analysis. For each day, the in situ record whose timestamp most closely matched the Sentinel-1A overpass time was adopted as the reference SM value. Table 1 summarizes the key attributes of each network, such as the number of stations, relative orbit number (a fixed path indicator within the Sentinel-1A repeat cycle, where identical numbers denote consistent observational geometry), sample size, and IGBP land cover classes. In this study, the global unified model integrated samples from all 24 ISMN networks listed in Table 1. However, only eight networks with a sufficient number of matched samples were chosen for independent local domain XGBoost and SHAP modeling (marked by triangles in Figure 1). The remaining 16 networks were included in the global pooled dataset but were not trained as standalone local models (marked by circles in Figure 1), because their limited sample size could lead to unstable model convergence and unreliable SHAP outputs.

2.2. Predictor Attributes

All predictors used for XGBoost-based SM retrieval were split into two broad categories, as follows: dynamic time-varying variables and static time-invariant variables. These features were further grouped into five subcategories, as follows: Sentinel-1A observations, vegetation parameters, meteorological variables, geospatial topography, and soil texture. Table 2 provides a comprehensive overview of all input features. Based on the Sentinel-1A IW GRD products, three radar-derived dynamic predictors are retained, as follows: VV co-polarized and VH cross-polarized backscattering coefficients, and local incidence angle (Theta). Additionally, the day of year (DOY) was extracted from Sentinel-1A metadata to represent seasonal variations in vegetation phenology and surface environmental conditions. Both ascending and descending data were combined in this study. Vegetation indices including LAI, the normalized difference vegetation index (NDVI), and enhanced vegetation index (EVI) were all extracted from MODIS products. The original LAI 4-day (MCD15A3H, 500 m) and the 8-day NDVI/EVI (MCD13A2, 1 km) products were smoothed and temporally interpolated to match Sentinel-1A acquisition times. Topsoil ERA5-Land SM and soil temperature (ST) corresponding to the Sentinel-1A acquisition times were utilized as the meteorology inputs. The static geographic and terrain variables included the latitude (LAT), longitude (LON) and a digital elevation model (DEM). Three time-invariant topsoil (0–30 cm) physical attributes were extracted from the Harmonized World Soil Database (HWSD), namely, sand content (sand, mass %), clay content (clay, mass %), and topsoil bulk density (density). Further details regarding the dataset’s preprocessing can be found in our previous work [48].
It should be noted that a vertical depth mismatch exists across the multisource datasets used in this study. The in situ SM observations represent the topsoil layer (≤5 cm), while ERA5-Land SM and ST correspond to the 0–7 cm layer. These two near-surface depth ranges are physically comparable and exhibit tightly coupled temporal dynamics for surface soil conditions. Soil texture variables extracted from HWSD represent the 0–30 cm topsoil layer. They act as static spatial proxies for near-surface soil physical properties rather than exact 0–5 cm measurements. Sentinel-1A C-band backscatter is predominantly sensitive to the upper soil under most conditions, demonstraing good vertical consistency with our ≤5 cm in situ SM, though the penetration depth may slightly increase under extremely dry bare soil scenarios. While such vertical discrepancies are widely encountered in multi-source machine learning-based SM retrieval, they introduce inherent uncertainties. Therefore, these differences in depth representativeness should be carefully considered when interpreting SHAP-derived feature importance results in this domain-dependence analysis.

3. Methodology

3.1. XGBoost

XGBoost is an ensemble machine learning algorithm built upon gradient boosting decision trees frameworks [21]. It stacks decision trees sequentially through serial iteration and residual correction to construct high-accuracy and robust predictive models. Compared with traditional boosting methods, XGBoost introduces regularization terms into the objective function, which effectively improves the model’s fitting accuracy, convergence speed, and the generalization performance. Furthermore, XGBoost supports the random subsampling of instances and features, custom loss functions, and learning rate shrinkage strategies, affording it excellent stability and flexibility. Characterized by high accuracy, high computational efficiency, strong fault tolerance and low deployment costs, XGBoost has been widely adopted in classification and regression tasks, ranking among the mainstream high-efficiency algorithms for machine learning on structured data at present.
In this study, the hyperparameter optimization of the XGBoost model was automatically implemented via Grid Search (GridSearchCV) combined with 10-fold cross validation. The parameter search space mainly focuses on three core hyperparameters that exert the most significant impact on model performance: the number of decision trees was searched from 1000 to 2000 with a step size of 500; the maximum depth of trees was tuned over integer values ranging from 3 to 7; and the learning rate was selected from three candidate values (0.005, 0.01, and 0.02). In addition, the subsample ratio and column sampling ratio per tree were both fixed at 0.8 to improve the model’s generalization ability and randomness. During model training, the mean squared error was adopted as the loss function, and a random seed (seed = 42) was set to guarantee the reproducibility of experimental results. GridSearchCV exhaustively traversed all hyperparameter combinations that yielded the optimal cross-validation score, and the final model was trained under this configuration. All model experiments were performed in a Python 3.11.5 environment. The XGBoost regression model was implemented using the xgboost package (version 3.2.0), and hyperparameter optimization was conducted via GridSearchCV embedded in scikit-learn (version 1.8.0). It should be noted that we employed a 70%/30% split for training/validation datasets, and 10-fold cross validation was used exclusively on the training set for hyperparameter tuning, whereas the final performance metrics were derived from the reserved 30% validation subset.

3.2. SHAP

After completing the parameter training of the XGBoost model, this study further employed the SHAP method for interpretability analysis [24]. The contribution of each feature was quantified by calculating the mean absolute SHAP value [49], thereby revealing the influence mechanisms of individual input variables on SM retrieval. SHAP is a classic game theory-based interpretability framework for machine learning models. It effectively overcomes the critical limitation of poor interpretability inherent in black box models, and has been widely adopted for the mechanistic analysis of various tree-based and complex predictive models. Its core principle is to quantify the marginal contribution of each input feature to the model’s predictions, measuring the positive or negative effects of individual features on prediction outcomes while accounting for feature interactions. This approach addresses the shortcomings of conventional feature importance methods, which only output weight proportions and fail to indicate the direction of feature impacts or their underlying mechanisms. During model analysis, SHAP values accurately reflect the contribution magnitude of each variable for every single sample. Meanwhile, the mean absolute SHAP values enable a global ranking of feature importance, clearly revealing the contributions of core variables and the interfering effects of secondary ones. Compared with traditional interpretation techniques, SHAP offers both global and local interpretability, quantifiable outputs, and high consistency. It can objectively and comprehensively dissect the prediction logic of ensemble models such as XGBoost, providing reliable guidance for research on the predictive mechanisms of eco-environmental parameters, including SM.
The SHAP analysis was implemented by the SHAP Python package (version 0.51.0). The TreeExplainer, which is specially designed for tree-based ensemble models, was adopted to compute SHAP values. Notably, it does not require a manually defined background reference sample set. Mean absolute SHAP values were calculated to obtain the global feature importance for each model. It is important to emphasize that all SHAP values and subsequent global feature importance statistics (mean absolute SHAP values) were computed exclusively on the held-out validation dataset (30% samples), rather than the training dataset, to reduce the influence of training set overfitting noise. Two metrics were derived from SHAP outputs in this study. The first is the raw mean absolute SHAP value (mean |SHAP|) mentioned in the above. For each feature, this was calculated by taking the average of absolute SHAP values across all validation samples, serving as our primary metric for comparing feature importance. The second metric is the normalized percentage contribution. Here, each feature’s mean absolute SHAP value was divided by the sum of mean absolute SHAP values over all input features and converted into a percentage. This normalized percentage was used solely for visualization within embedded donut charts, and was not adopted for quantitative comparison in the main text discussion.

4. Results

4.1. Performance Assessment of SM Retrievals

Table 3 summarizes the statistical metrics (Bias, RMSE and R2) of the XGBoost models trained separately on eight independent monitoring networks and the combined global networks, with performance reported for both training and validation datasets. The XGBoost model demonstrates a strong capability in capturing the nonlinear relationships between input features and SM, as evidenced by the consistently high training and validation accuracies across all networks. The near-zero biases for both training and validation datasets further indicate the absence of systematic overestimation or underestimation, which provides a statistically unbiased residual background for the subsequent SHAP-based interpretability analysis. However, a pronounced performance gap exists between training and validation datasets across all networks, with validation R2 values falling within a considerably wider range of 0.764–0.945. This broad variance in validation datasets is theoretically expected, and is hypothesized to reflect the underlying spatial heterogeneity of each monitoring network. The REMEDHUS network achieves the optimal validation performance, followed by FMI, COSMOS-UK, and SMOSMANIA-SWATMEX. These four networks yield validation R2 values exceeding 0.900, demonstrating that the XGBoost model can effectively capture dominant SM controls within their respective local domains. In contrast, the SNOTEL network exhibits the weakest generalization ability among all independent datasets, recording the highest validation RMSE and the lowest R2. Moderate retrieval accuracy is recorded for SCAN, USCRN, and RSMN. Even though the XGBoost model at the RSMN network shows near perfect training performance, its validation R2 drops to 0.894, implying substantial overfitting within this local network.
For each individual network, the XGBoost model achieves exceptionally high training accuracy but consistently degraded validation accuracy. Interestingly, when all networks are aggregated into a single training set (the ISMN SM dataset in the last row), the model exhibits fundamentally different behavior. The training accuracy (R2 = 0.865) is considerably lower than that of any individual network, which is expected given the substantially increased feature space complexity and the inclusion of diverse, sometimes conflicting, land surface signatures. Although the global merged model yields a lower training set R2 (0.865) than any individual regional-domain model, it exhibits a comparatively smaller degradation from training to validation performance. This pattern indicates the reduced risk of over-fitting to site-specific local land surface signatures and better cross-domain stability across heterogeneous conditions, rather than superior absolute prediction accuracy. Indeed, seven of eight regional-domain models achieve higher validation set R2 values within their respective local domains.
To complement the quantitative statistical metrics presented in Table 3, density-colored scatterplots illustrate the consistency between the estimated and measured SM across the global merged ISMN datasets (Figure 2) and individual monitoring networks (Figure 3), respectively. In all subplots, the black dashed line denotes the 1:1 reference, while the red lines represent the linear regression fit between estimated and measured SM. Color intensity represents the local density of data points, with red indicating high-density clusters and blue denoting low-density areas.
Figure 2a presents the training datasets of the global merged ISMN networks, and Figure 2b corresponds to the validation datasets. For the training datasets, nearly all data points tightly cluster along the 1:1 reference line, and the regression line closely overlaps the 1:1 line. This visual alignment is consistent with the extremely low training bias and high R2 reported for ISMN SM in Table 3. The dense and narrow distribution of points confirms that the XGBoost model effectively captures the overall SM variation across the full combined dataset during training. The validation dataset exhibits linear trends similar to the training results, which verifies that the model retains an acceptable generalization capacity for the global merged dataset with negligible systematic over- or underestimation.
Figure 3 displays the training (left column: a, c, e, g) and validation (right column: b, d, f, h) density scatter plots for the first four regional monitoring networks, which contain a sufficiently large number of samples. The results for the remaining four networks are displayed in Figure A1 of Appendix A. For the SNOTEL network, the training scatter (Figure 3a) shows concentrated points along the 1:1 line, but the validation subset (Figure 3b) features pronounced point dispersion across the full SM range, with the red regression line clearly deviating from the 1:1 line at high SM values. This visual divergence matches the poorest generalization performance shown in Table 3. For the remaining three networks, the training datasets form extremely compact and narrow bands aligned with the 1:1 line, supporting their high training accuracy. The validation scatter plots widen moderately, but maintain tighter clustering than SNOTEL’s validation results. The regression lines remain well-matched to the 1:1 reference, demonstrating mild generalization degradation without severe prediction bias.
Across both the global merged ISMN dataset and all individual local networks, the training datasets consistently form tighter clusters along the 1:1 line, whereas the validation data display broader scattering. This phenomenon reflects a well-recognized pattern in machine learning, where the model’s performance on the validation datasets is inherently lower than on the training datasets. The local networks present drastically different degrees of validation scatter dispersion. Networks with high validation R2 (e.g., COSMOS-UK) retain compact linear alignment in validation plots, while lower-performance networks (e.g., SNOTEL) show severe point divergence. Specifically, the SNOTEL network shows obvious SM underestimation within the high SM range. Its sites are mainly distributed in mountainous cold regions, where complex terrain, snow cover and snow melt processes lead to the model’s tendency to underestimate high SM values. This visual inter-network disparity corroborates the statistical conclusion that spatial domain and site-specific environmental heterogeneity fundamentally modulate the XGBoost model’s SM retrieval capacity. In the following section, the SHAP analysis will be implemented to examine how input feature importance shifts depending on the spatial domain.

4.2. Pearson Correlation Analysis of Predictor Input Features

Prior to interpreting the spatial domain dependence of feature importance based on SHAP, we quantified pairwise linear correlations among all input predictors using Pearson correlation heatmaps (Figure 4 for the merged global ISMN networks, Figure 5 for the first four individual local networks, and Figure A2 for the remaining four local networks). The color gradient ranges from dark blue (R = −1, strong negative correlation) to dark red (R = 1, perfect positive correlation), with yellow–green intermediate tones representing near-zero linear dependence. These results help reveal spatial domain-driven differences in covariation patterns among environmental, satellite, vegetation, and meteorological predictors across global and local regions.
Figure 4 shows a strong positive linear correlation between VV co-polarized and VH cross-polarized backscattering coefficients (R = 0.739). This inherent physical covariation arises from the consistent microwave scattering responses of surface vegetation and soil roughness across both polarizations. In addition, LAI, NDVI and EVI form a tightly correlated cluster. Notably, NDVI and EVI show a near perfect positive correlation (R = 0.932), as vegetation proxy indicators capture overlapping canopy greenness and biomass signals. Furthermore, VV and VH show highly positive correlations with vegetation proxy indicators and ERA5-Land SM, underscoring the capacity of SAR observations to be used in monitoring vegetation and soil. For the remaining predictors, nearly all pairwise correlation coefficients fall within the range of −0.30~0.30, indicating weak linear covariation. It should be noted that the strong multicollinearity among these specific feature groups will influence the SHAP attribution outputs in subsequent sections.
The four local monitoring networks demonstrate dramatically different pairwise feature correlation magnitudes and spatial covariation patterns, providing direct evidence that spatial domains reshape linear dependencies between predictive variables. The SNOTEL network presents the strongest topographic correlation signal among all subplots where LAT and DEM show an intense negative correlation. While the VV and VH radar correlation remains strongly positive (consistent with the global dataset), the covariation among vegetation indices (LAI, NDVI, EVI) weakens visibly relative to Figure 4. In addition, VV and VH show a weak positive correlation with vegetation indicators compared to the global results in Figure 4. Conversely, the SCAN and USCRN networks feature the most pronounced vegetation index correlation cluster, with NDVI and EVI reaching a near-perfect positive correlation. VV and VH correlation remains stable, but they have a higher positive correlation with vegetation indices and ERA5-Land SM compared to the global dataset. This suggests that vegetation and SM dominate local surface processes in this regional network. For the COSMOS-UK network, a strong VV/VH correlation and moderate vegetation index covariation are retained. However, the correlation between VV/VH, vegetation indices, and ERA5-Land SM decreases markedly compared with those from both the global dataset and other monitoring networks. This region experiences a temperate maritime climate characterized by frequent precipitation, persistent moderate SM and continuous vegetation coverage. Consequently, Sentinel-1 C-band backscatter is largely modulated by local vegetation canopy and surface roughness variations, and no longer tightly covaries with large domain reanalysis and vegetation index products.
Overall, within physically linked feature families (VV/VH, LAI/NDVI/EVI), strong positive linear correlations persist universally across both the global merged dataset and all local networks. These inherent physical covariations are invariant across spatial domains. In contrast, the cross-category correlations are highly sensitive to the spatial domain and local environmental background. This domain-driven divergence in predictor covariation provides a critical mechanistic foundation for the upcoming SHAP-based analysis of feature importance. As spatial extents shift from small local networks to a global merged dataset, changes in inter-feature linear dependencies inherently alter the relative marginal contribution of each predictor to XGBoost SM predictions, forming the core spatial domain-dependent feature importance signal explored in the following section.

4.3. SHAP Based Feature Importance

In this section, we further employ SHAP interpretability to quantify the absolute contribution (mean |SHAP|) and marginal response direction of each predictor to SM retrieval. Figure 6 presents SHAP outputs for the merged global ISMN dataset. Figure 7 illustrates feature importance and SHAP dependence distributions for the first four individual networks (SNOTEL, SCAN, USCRN, and COSMOS-UK), while results for the remaining four networks are displayed in Figure A3. The horizontal bar charts (Figure 6a, Figure 7a,c,e,g and Figure A3a,c,e,g) rank predictors by their mean |SHAP| values, which quantify the overall predictive weight of each variable, and the embedded donut charts summarize the cumulative contribution of the top five dominant features. The corresponding SHAP summary beeswarm plots (Figure 6b, Figure 7b,d,f,h and Figure A3b,d,f,h) visualize how varying feature values shift the predicted SM. In these plots, the horizontal axis represents the SHAP values (positive = increases predicted SM, negative = reduces predicted SM) and color gradients denote raw feature magnitude (red = high feature value, blue = low feature value). It should be noted that under strong feature collinearity, the mean |SHAP| metric redistributes predictive attribution across correlated predictors. Therefore, observed shifts in individual feature importance rankings may reflect changes in inter-feature correlation structure, not solely differences in the underlying physical SM driving mechanisms. Accordingly, we avoid overly strong causal interpretations for single-feature SHAP rankings in this study.
As shown in the horizontal bar chart (Figure 6a), ERA5-Land SM acts as the overwhelmingly dominant predictor based on the raw mean |SHAP| values. The embedded donut subplot further presents normalized percentage contributions for visual reference. This result aligns with physical expectations, as ERA5-Land SM provides a reliable baseline estimate of SM across a wide range of global land surface conditions. The second to fifth most influential variables are LAT, DEM, LON, and EVI. Collectively, these top five predictors account for 68.55% of the total predictive power, confirming that large domain spatial patterns of SM are primarily governed by background meteorological moisture states and broad latitudinal and elevation climate gradients. By contrast, vegetation indices (LAI and NDVI), Sentinel-1A observations (VV, VH and incidence angle), soil texture properties (Sand, Clay and Density), ERA5-Land ST, and DOY exhibit far lower mean |SHAP| values. The limited contribution of these variables at the global domain level can be explained by the fact that their local effects tend to be smoothed or averaged out when datasets from widely different climate and terrain zones are combined.
The beeswarm plot (Figure 6b) further clarifies marginal driving effects. Elevated ERA5-Land SM values produce strongly positive SHAP outputs, representing a consistent positive linear driving relationship. LAT and DEM exhibit mixed positive and negative SHAP responses, which can be largely attributed to latitudinal gradients in temperature. Furthermore, the results show that EVI, Sentinel-1A and soil parameters exert weak impacts on the global merged dataset. It should be noted that considerable sample size imbalance exists across the individual monitoring networks. For instance, SNOTEL contributes 56,366 samples, while RSMN provides only 7977 samples. The global merged dataset directly combines all available records without down sampling or equal weight stratification across networks. Therefore, the global domain model and its SHAP-derived feature importance ranking are inherently weighted toward signals from large-sample northern mid-latitude networks, rather than representing an equal contribution from each network. For future work, stratified resampling or equal weight network ensemble strategies could mitigate this imbalance, but such balancing experiments are beyond the scope of this study.
Figure 7 shows that each network demonstrates a different ranking of controlling predictors, providing definitive evidence that feature importance is highly spatially domain-dependent and modulated by site-specific environmental conditions. For the SNOTEL network, ERA5-Land SM remains the primary predictor, yet its relative dominance is drastically weakened compared to the global dataset. DOY is identified as the second most dominant predictor, with LAT, DEM, and LON ranking subsequently. The remaining features occupy the lowest importance tiers. The beeswarm plot reveals that DOY exerts a clear directional control on SM predictions locally, even though it falls to a much lower range in the global merged model. SNOTEL stations are concentrated across high-latitude and high-elevation regions of North America, where SM variability is heavily controlled by snow accumulation and melt processes, as well as intra-annual phenology. DOY effectively represents these strong local seasonal cycles, thereby gaining high local domain importance. In contrast, the global merged dataset combines observations from both hemispheres and multiple climate zones. Consequently, seasonal signals encoded by DOY become spatially incoherent: phenological and hydrological cycles are out of phase across hemispheres, and tropical sites show weak seasonal variation. These conflicting seasonal patterns partially counteract each other, reducing DOY’s marginal global predictive contribution. This rank shift illustrates how the global model prioritizes variables that capture universally consistent large domain moisture and climate gradients when reconciling divergent local environmental signals from different monitoring networks.
The SCAN network exhibits the most extreme dominance of ERA5-Land SM among all local networks, far exceeding all other predictors and even surpassing its contribution in the global dataset. DEM ranks second, while soil texture (Sand fraction) enters the top five most influential features. This pattern reflects that the background meteorological moisture is the primary control factor, while subtle topographic and soil texture differences drive local small-domain SM variability. The other features contribute minimally. For the USCRN network, ERA5-Land SM remains the dominant predictor, but LON rises to the second most influential variable. Sand fraction and DOY also rank within the top five predictors, highlighting combined controls of soil texture and seasonal climate in local SM. Vegetation and radar features remain secondary drivers. The COSMOS-UK network delivers the most dramatic domain-dependent reversal of feature importance observed in this study. Here, LON replaces ERA5-Land SM as the most important predictor, with a mean |SHAP| value of 0.0434, while ERA5-Land SM falls to second place (0.0232). DOY and DEM occupy the third and fourth positions, and EVI enters the top five key features. The beeswarm plot confirms that LON produces a wide range of positive and negative SHAP values, representing its dominant modulating effect on SM retrieval. This is driven by systematic east-west gradients of precipitation, soil and vegetation across UK. For the SMOSMANIA-SWATMEX network (Figure A3c,d), the sand fraction ranks as the second most important predictor, contributing approximately 20% of total SHAP importance. This highlights the strong local control exerted by soil texture on surface SM variability across this French observational domain. Within this regional domain, large spatial gradients in sand content modulate SM holding capacity, infiltration, and drainage behavior. Therefore, static soil texture properties become a major driver of topsoil SM spatial heterogeneity, outweighing several dynamic satellite-derived and meteorological predictors at this local domain. This observation aligns with our core conclusion that static soil texture variables gain elevated predictive importance within local domain models.
When integrating all heterogeneous monitoring networks into a unified global dataset, ERA5-Land SM becomes the undisputed primary control. Fine domain predictors (vegetation, radar backscatter, soil texture) lose predictive weight, as their site-specific signals are smoothed and averaged across divergent environmental zones. Latitude, elevation, and longitude act as secondary drivers by capturing continental domain climate gradients. However, at individual local networks, the relative importance of ERA5-Land SM declines for some networks, and predictors tied to regionally unique landscape characteristics rise to prominence. Across different domains, Sentinel-1A radar polarizations (VV, VH), ERA5-Land ST, and LAI consistently occupy the lowest tiers of feature importance. This indicates that microwave backscatter and LAI exert weak marginal effects on the XGBoost SM retrieval, regardless of spatial extent, likely due to their collinear covariation with NDVI/EVI and ERA5-Land SM, as observed in the preceding correlation heatmaps. Collectively, the SHAP results quantitatively verify that the relative contribution of input predictors to SM retrieval based on the XGBoost model is strongly spatial-domain-dependent. The observed shifts in feature importance from global to local domains arise from the changing dominant environmental processes governing surface SM variability at different spatial extents. This provides critical guidance for optimizing predictor selection in multi-scale machine learning-based SM retrieval frameworks.

5. Discussion

5.1. Role of ERA5-SM Revealed by Ablation Experiments

In our global domain analysis and certain other local domain analyses, ERA5-Land SM exhibits the greatest SHAP-derived importance among all input predictors. However, this study should not be regarded as a conventional ERA5-Land SM downscaling exercise. The ablation experiments summarized in Table 4 provide supportive quantitative evidence for this viewpoint. When only ERA5-Land SM is used as the predictor, the model achieves a validation R2 of only 0.337. By contrast, the model excluding ERA5-Land SM still achieves considerable predictive accuracy, with validation R2 reaching 0.782, using Sentinel-1 SAR, vegetation indices, static geospatial covariates (latitude, longitude, DEM), and other terrain and soil auxiliary variables. This demonstrates that satellite-derived and surface covariates carry independent predictive capacity for surface SM retrieval. When ERA5-Land SM is incorporated into the full input model, the validation R2 is further increased to 0.820. This indicates that ERA5-Land SM delivers moderate performance improvement by acting as a coarse regional background prior. Therefore, our framework functions as a multi-source hybrid machine learning system that learns to refine the coarse reanalysis background using high-resolution satellite and surface information, rather than simply downscaling ERA5-Land SM products.

5.2. Predictor Importance on SM Retrieval

Previous SHAP-based SM modeling studies have mostly reported feature importance patterns for either a single global model or one specific study area, generally confirming the relative contributions of common SM predictors. Beyond reproducing these known patterns, the novelty of this work lies in the systematic comparisons of SHAP-derived predictor importance hierarchies across a global domain and multiple independent regional domain models. Our results demonstrate that the dominant predictors governing surface SM are not invariant, but instead change across spatial domains. We further differentiate domain stable inherent physical feature correlations from spatially variable cross-category covariations, illustrating how region-specific correlation structures can reshape SHAP-based attribution outputs. This reveals the spatial domain-dependent nature of feature importance hierarchies within multi-source SM machine learning workflows.
To mitigate interpretability challenges caused by multicollinearity among physically correlated predictors, we performed grouped feature importance analysis by summing the mean |SHAP| values for predefined collinear feature groups (Table 5). In the global domain, the geospatial covariate group (LAT+LON+DEM) and ERA5-SM are the two most influential predictors, showing comparable aggregated importance. The vegetation group (LAI+NDVI+EVI) ranks third, whereas the Sentinel-1 radar group (VV+VH) shows relatively minor importance. However, clear regional variations are observed across different datasets. The REMEDHUS, FMI, and RSMN networks show broadly consistent patterns with the global ISMN dataset. For the SCAN and USCRN networks, ERA5-Land SM supersedes the geospatial group as the most important predictor. For the COSMOS-UK network, the geospatial group (LAT+LON+DEM) dominates model predictions, followed by the vegetation indices group (LAI+NDVI+EVI). For the SMOSMANIA-SWATMEX network, the sand fraction becomes the third most influential predictor. For the SNOTEL network, the vegetation indices group occupies the third position in predictor importance. The vegetation collinear group shows high importance in COSMOS-UK and RSMN networks, but only moderate contributions in other networks. By contrast, the radar group (VV+VH) exhibits low aggregated importance across most study domains. The higher mean |SHAP| values for the geospatial covariate group among global and local domains highlights the strong control of spatial topography on SM. The moderate mean |SHAP| values for the vegetation indices group arises from multiple physical feedbacks between vegetation and SM. These findings demonstrate that the aggregated importance of collinear feature groups is strongly region dependent. The combined effect of spatially correlated predictors (LAT+LON+DEM) is amplified when strong topographic and climatic gradients exist within a study area. Ultimately, the relative priority of predictor groups shifts with local environmental conditions, suggesting that global domain feature importance cannot be directly transferred to local domain applications.
Our SHAP-based analysis reveals persistently low marginal importance for Sentinel-1 VV and VH polarizations across most spatial domains, which appears counter intuitive considering the abundance of existing literature reporting reliable SM retrieval from C-band SAR observations. This phenomenon should not be interpreted as evidence that Sentinel-1 backscatter lacks physical sensitivity to SM. Instead, it mainly stems from the multicollinearity among input features within our multi-source XGBoost model. Predictors such as NDVI, EVI and ERA5-Land SM share strong covariation with SAR backscatter signals. Under such conditions, TreeSHAP tends to redistribute predictive attribution among correlated covariates, meaning most SM-related information contained in VV/VH is captured by other co-varying predictors. Therefore, this represents weak marginal contributions only under the current feature combination. If Sentinel-1 SAR observations were used as the primary input source, their predictive power would become more prominent. In addition, site-specific surface conditions, such as vegetation canopy structure and surface roughness, can further modulate the effective contribution of SAR signals across different regional domains.
Importantly, longitude (LON) is a static geographic coordinate and not a direct physical predictor. Its elevated importance in the COSMOS-UK network does not reflect a direct physical influence of longitude itself. Instead, LON serves as a spatial proxy for the strong west-to-east environmental gradients present across the study area. Across the UK, precipitation decreases markedly from the wet western coast toward the drier eastern interior, which further drives spatial variations in SM and vegetation status. The XGBoost model captures these consistent spatial patterns and uses LON as a surrogate variable to represent these correlated physical signals. Although location memorization effects cannot be fully excluded, the dominant contribution originates from the real east–west climatic and environmental gradients. This observation highlights a key interpretive caveat, namely, that geographic coordinates may achieve high SHAP importance when they are tightly coupled to underlying physical drivers, even though they are not physically meaningful predictors themselves.

5.3. Limitations and Outlook

This study explores the shifts in predictor importance controlling surface SM across different spatial domains based on multi-site observations from the ISMN. Nevertheless, the sample size varies substantially among individual ISMN monitoring networks. Networks with larger observation volumes contribute more samples to model training and SHAP analysis, which makes our comparative results intrinsically biased toward regions with abundant in situ records. Networks with limited sample counts may not be sufficiently represented, and their feature importance patterns could be under-captured in our outputs. Although we select eight representative networks with adequate data volumes for regional domain analysis, the uneven spatial distribution of available ground observations remains an inherent constraint of this work. Future investigations could adopt stratified sampling strategies to balance sample contributions across networks, thereby achieving more equitable comparisons of driving factor patterns among diverse regions.
Given the above-mentioned limitations, several promising directions remain for future improvement. First, further work could explore objective spatial domain thresholds to quantitatively separate global domain and regional domain modeling regimes, which would help clarify at which spatial level the dominant controlling factors of SM begin to shift. Second, the reliability of SHAP attribution should be validated under more diverse scenarios, including varying seasonal and temporal contexts. Additionally, it is necessary to evaluate how different machine learning algorithms may alter SHAP-derived feature importance outputs. Third, regarding validation strategies, future studies should adopt more robust evaluation schemes such as leave-year-out, leave-station-out and spatially blocked cross-validation to mitigate spatiotemporal leakage. Grouped feature block importance analysis can also be implemented to reduce the attribution bias caused by multicollinearity. Finally, more complete observation datasets are required, and hydrometeorological time-lag effects should be incorporated into the modeling workflow. Consistent with our earlier discussion, stratified sampling strategies should also be considered to balance sample contributions across networks and achieve fairer inter-regional comparisons of SM driving patterns.

6. Conclusions

This study develops an interpretable XGBoost and SHAP framework to quantitatively describe the spatial domain dependence evolution of predictor importance for machine learning-based SM retrieval, using global multiple networks’ in situ SM observations from the ISMN covering 2017–2024. The model performance comparison between the unified global model and the independent regional monitoring networks model has verified that the globally trained model exhibits a reduced risk of over-fitting and better cross-domain stability across diverse land surface conditions. Conversely, while many regional domain models achieve higher absolute retrieval accuracy within their specific locals, their overall performance remains constrained by local surface environmental heterogeneity.
Pearson correlation analysis demonstrated that physically linked feature groups (e.g., Sentinel-1A VV/VH polarizations, and LAI/NDVI/EVI vegetation indices) maintain stable high collinearity across different domains. However, cross-category correlations among radar, vegetation and meteorological predictors are highly sensitive to study extent and regional background conditions. SHAP-based feature importance quantification further confirmed evident domain-dependent differentiations among dominant controlling factors. At the global domain level, ERA5-Land SM acts as the dominant predictor, contributing nearly 44% of total predictive power, with geospatial gradients (latitude, longitude, elevation) serving as secondary regulators. Localized covariates including radar signals, vegetation indices and soil texture may exhibit reduced marginal predictive weight. This reduction partly originates from the spatial averaging of heterogeneous surface signals and inter-feature covariation across predictors.
At regional domains, the feature importance of ERA5-Land SM exhibits a heterogeneous behavior relative to the global pooled model. It remains the dominant predictive feature for the SNOTEL, SCAN, USCRN SMOSMANIA-SWATMEN and RSMN networks, achieving even higher importance for the SCAN network than in the global model. Nevertheless, for the COSMOS-UK network, longitude (LON) overtakes ERA5-Land SM to become the leading predictor, while latitude (LAT) replaces it as the top feature for the REMEDHUS and FMI networks. Group feature importance analysis highlights that the relative contributions of predictor groups shift substantially between global and regional domains. Within our multi-source modeling framework, Sentinel-1A radar backscatter, surface soil temperature and LAI exhibit relatively low marginal independent contributions across most spatial domains, largely due to collinearity with other vegetation and SM-related covariates. However, this does not imply that these radar observations contain no SM-relevant physical signals.
These findings provide clear spatial domains and adaptive guidance for predictor selection in SM modeling. For global domain SM retrieval frameworks, ERA5-Land SM should be retained as the core input, while fine domain local predictors can be simplified to reduce computational cost. Regional downscaling models, on the other hand, require targeted feature screening that prioritizes seasonal, topographic and soil texture variables to capture local SM variability. Moreover, highly collinear vegetation and radar features can be selectively removed to streamline model training without significant accuracy loss. Ultimately, this work advances the interpretability of machine learning-based SM retrieval, and offers a scalable feature optimization strategy for generating high-resolution SM products under diverse global land surface conditions.

Author Contributions

Conceptualization, S.Z. and X.B.; methodology, X.B.; software, S.Z.; validation, S.Z. and Y.W.; formal analysis, S.Z.; investigation, S.Z.; resources, X.B.; data curation, Y.W.; writing—original draft preparation, S.Z.; writing—review and editing, X.B. and Y.W.; visualization, S.Z. and Y.W.; supervision, X.B. and W.S.; project administration, X.B. and W.S.; funding acquisition, X.B. and W.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by funding from the National Natural Science Foundation of China under Grant 42471385 and 42471032.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Acknowledgments

The authors express sincere thanks to the International Soil Moisture Network and its data providers, as listed in Table 1.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviation

SMSoil Moisture
EVIEnhanced Vegetation Index
LAILeaf Area Index
NDVINormalized Difference Vegetation Index
LATLatitude
LONLongitude
DEMDigital Elevation Model
DOYDay of Year
STSurface Temperature
VVVertical–Vertical Polarization
VHVertical–Horizontal Polarization
ERA5-LandERA5-Land Reanalysis Dataset
ISMNInternational Soil Moisture Network
SNOTELSnow Telemetry Network
USCRNUnited States Climate Reference Network
XGBoostExtreme Gradient Boosting
SHAPShapley Additive Explanations
CVCross Validation
RMSERoot Mean Square Error
R2Coefficient of Determination

Appendix A

Figure A1. Colored density scatter plots of measured and estimated SM for four representative monitoring networks. Left panels (a,c,e,g) correspond to training datasets and right panels (b,d,f,h) correspond to validation datasets: (a,b) REMEDHUS, (c,d) SMOSMANIA-SWATMEX, (e,f) FMI, and (g,h) RSMN. The black dashed line represents the 1:1 reference, and the solid red line denotes the linear regression fit between estimated and measured SM. Color bars indicate point sample density.
Figure A1. Colored density scatter plots of measured and estimated SM for four representative monitoring networks. Left panels (a,c,e,g) correspond to training datasets and right panels (b,d,f,h) correspond to validation datasets: (a,b) REMEDHUS, (c,d) SMOSMANIA-SWATMEX, (e,f) FMI, and (g,h) RSMN. The black dashed line represents the 1:1 reference, and the solid red line denotes the linear regression fit between estimated and measured SM. Color bars indicate point sample density.
Remotesensing 18 03006 g0a1
Figure A2. Pearson correlation coefficient heatmaps of input predictors for four individual monitoring networks: (a) REMEDHUS, (b) SMOSMANIA-SWATMEX, (c) FMI, and (d) RSMN. The color bar represents correlation values ranging from −1 (dark blue) to 1 (dark red).
Figure A2. Pearson correlation coefficient heatmaps of input predictors for four individual monitoring networks: (a) REMEDHUS, (b) SMOSMANIA-SWATMEX, (c) FMI, and (d) RSMN. The color bar represents correlation values ranging from −1 (dark blue) to 1 (dark red).
Remotesensing 18 03006 g0a2
Figure A3. The SHAP-based importance ranking and direction of driving factors for four individual networks: (a,b) REMEDHUS, (c,d) SMOSMANIA-SWATMEX, (e,f) FMI, and (g,h) RSMN. The horizontal axis denotes SHAP values, and color gradients represent raw feature magnitude (red = high feature value, blue = low feature value).
Figure A3. The SHAP-based importance ranking and direction of driving factors for four individual networks: (a,b) REMEDHUS, (c,d) SMOSMANIA-SWATMEX, (e,f) FMI, and (g,h) RSMN. The horizontal axis denotes SHAP values, and color gradients represent raw feature magnitude (red = high feature value, blue = low feature value).
Remotesensing 18 03006 g0a3

References

  1. Montzka, C.; Brocca, L.; Chen, H.; Das, N.N.; Dasgupta, A.; Rahmati, M.; Jagdhuber, T. AI in soil moisture remote sensing. Int. J. Appl. Earth Obs. Geoinf. 2026, 146, 105011. [Google Scholar] [CrossRef] [Scilit]
  2. Lamichhane, M.; Mehan, S.; Mankin, K.R. Soil Moisture Prediction Using Remote Sensing and Machine Learning Algorithms: A Review on Progress, Challenges, and Opportunities. Remote Sens. 2025, 17, 2397. [Google Scholar] [CrossRef] [Scilit]
  3. Kornelsen, K.C.; Coulibaly, P. Advances in soil moisture retrieval from synthetic aperture radar and hydrological applications. J. Hydrol. 2013, 476, 460–489. [Google Scholar] [CrossRef] [Scilit]
  4. Senanayake, I.P.; Pathira Arachchilage, K.R.L.; Yeo, I.-Y.; Khaki, M.; Han, S.-C.; Dahlhaus, P.G. Spatial Downscaling of Satellite-Based Soil Moisture Products Using Machine Learning Techniques: A Review. Remote Sens. 2024, 16, 2067. [Google Scholar] [CrossRef] [Scilit]
  5. Singh, A.; Gaurav, K.; Sonkar, G.K.; Lee, C.-C. Strategies to Measure Soil Moisture Using Traditional Methods, Automated Sensors, Remote Sensing, and Machine Learning Techniques: Review, Bibliometric Analysis, Applications, Research Findings, and Future Directions. IEEE Access 2023, 11, 13605–13635. [Google Scholar] [CrossRef] [Scilit]
  6. Liu, Y.; Yang, Y. Advances in the Quality of Global Soil Moisture Products: A Review. Remote Sens. 2022, 14, 3741. [Google Scholar] [CrossRef] [Scilit]
  7. Rahmati, M.; Balenzano, A.; Bechtold, M.; Brocca, L.; Fluhrer, A.; Jagdhuber, T.; Karamvasis, K.; Mengen, D.; Reichle, R.H.; Kim, S.-b.; et al. Soil moisture retrieval from Sentinel-1: Lessons learned after more than a decade in orbit. Remote Sens. Environ. 2026, 333, 115146. [Google Scholar] [CrossRef] [Scilit]
  8. Fan, D.; Zhao, T.; Jiang, X.; García-García, A.; Schmidt, T.; Samaniego, L.; Attinger, S.; Wu, H.; Jiang, Y.; Shi, J.; et al. A Sentinel-1 SAR-based global 1-km resolution soil moisture data product: Algorithm and preliminary assessment. Remote Sens. Environ. 2025, 318, 114579. [Google Scholar] [CrossRef] [Scilit]
  9. Fang, B.; Lakshmi, V.; Cosh, M.; Liu, P.W.; Bindlish, R.; Jackson, T.J. A global 1-km downscaled SMAP soil moisture product based on thermal inertia theory. Vadose Zone J. 2022, 21, e20182. [Google Scholar] [CrossRef] [Scilit]
  10. Zheng, C.; Jia, L.; Zhao, T. A 21-year dataset (2000–2020) of gap-free global daily surface soil moisture at 1-km grid resolution. Sci. Data 2023, 10, 139. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Song, P.; Zhang, Y.; Guo, J.; Shi, J.; Zhao, T.; Tong, B. A 1 km daily surface soil moisture dataset of enhanced coverage under all-weather conditions over China in 2003–2019. Earth Syst. Sci. Data 2022, 14, 2613–2637. [Google Scholar] [CrossRef] [Scilit]
  12. Shangguan, Y.; Min, X.; Shi, Z. Inter-comparison and integration of different soil moisture downscaling methods over the Qinghai-Tibet Plateau. J. Hydrol. 2023, 617, 129014. [Google Scholar] [CrossRef] [Scilit]
  13. Han, Q.; Zeng, Y.; Zhang, L.; Wang, C.; Prikaziuk, E.; Niu, Z.; Su, B. Global long term daily 1 km surface soil moisture dataset with physics informed machine learning. Sci. Data 2023, 10, 101. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Li, Q.; Shi, G.; Shangguan, W.; Nourani, V.; Li, J.; Li, L.; Huang, F.; Zhang, Y.; Wang, C.; Wang, D.; et al. A 1 km daily soil moisture dataset over China using in situ measurement and machine learning. Earth Syst. Sci. Data 2022, 14, 5267–5286. [Google Scholar] [CrossRef] [Scilit]
  15. Zhang, Y.; Liang, S.; Ma, H.; He, T.; Wang, Q.; Li, B.; Xu, J.; Zhang, G.; Liu, X.; Xiong, C. Generation of global 1 km daily soil moisture product from 2000 to 2020 using ensemble learning. Earth Syst. Sci. Data 2023, 15, 2055–2079. [Google Scholar] [CrossRef] [Scilit]
  16. Zhu, L.; Dai, J.; Jin, J.; Yuan, S.; Xiong, Z.; Walker, J.P. Are the Current Expectations for SAR Remote Sensing of Soil Moisture Using Machine Learning Overoptimistic? IEEE Trans. Geosci. Remote Sens. 2025, 63, 4501815. [Google Scholar] [CrossRef] [Scilit]
  17. Zhu, L.; Cai, Q.; Jin, J.; Yuan, S.; Shen, X.; Walker, J.P. Multi-Scale domain adaptation for high-resolution soil moisture retrieval from synthetic aperture radar in data-scarce regions. J. Hydrol. 2025, 657, 133073. [Google Scholar] [CrossRef] [Scilit]
  18. Zhu, L.; Dai, J.; Liu, Y.; Yuan, S.; Qin, T.; Walker, J.P. A cross-resolution transfer learning approach for soil moisture retrieval from Sentinel-1 using limited training samples. Remote Sens. Environ. 2024, 301, 113944. [Google Scholar] [CrossRef] [Scilit]
  19. Molnar, C. Interpretable Machine Learning. In A Guide for Making Black Box Models Explainable; 2022; Available online: https://christophm.github.io/interpretable-ml-book/ (accessed on 1 July 2026).
  20. Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  21. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]
  22. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
  23. Fisher, A.; Rudin, C.; Dominici, F. All Models are Wrong, but Many are Useful: Learning a Variable’s Importance by Studying an Entire Class of Prediction Models Simultaneously. J. Mach. Learn. Res. 2019, 20, 1–81. [Google Scholar]
  24. Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; pp. 4768–4777. [Google Scholar]
  25. Friedman, J.H.; Popescu, B.E. Predictive learning via rule ensembles. Ann. Appl. Stat. 2008, 2, 916–954. [Google Scholar] [CrossRef] [Scilit]
  26. Goldstein, A.; Kapelner, A.; Bleich, J.; Pitkin, E. Peeking Inside the Black Box: Visualizing Statistical Learning with Plots of Individual Conditional Expectation. J. Comput. Graph. Stat. 2015, 24, 44–65. [Google Scholar] [CrossRef] [Scilit]
  27. Dorigo, W.; Himmelbauer, I.; Aberer, D.; Schremmer, L.; Petrakovic, I.; Zappa, L.; Preimesberger, W.; Xaver, A.; Annor, F.; Ardö, J.; et al. The International Soil Moisture Network: Serving Earth system science for over a decade. Hydrol. Earth Syst. Sci. 2021, 25, 5749–5804. [Google Scholar] [CrossRef] [Scilit]
  28. Galle, S.; Grippa, M.; Peugeot, C.; Moussa, I.B.; Cappelaere, B.; Demarty, J.; Mougin, E.; Panthou, G.; Adjomayi, P.; Agbossou, E.K.; et al. AMMA-CATCH, a Critical Zone Observatory in West Africa Monitoring a Region in Transition. Vadose Zone J. 2018, 17, 1–24. [Google Scholar] [CrossRef] [Scilit]
  29. Cook, D.R. Soil Temperature and Moisture Profile (STAMP) System Handbook; DOE Office of Science Atmospheric Radiation Measurement (ARM) Program: Argonne, IL, USA, 2018. [Google Scholar]
  30. Dabrowska-Zielinska, K.; Musial, J.; Malinska, A.; Budzynska, M.; Gurdak, R.; Kiryla, W.; Bartold, M.; Grzybowski, P. Soil Moisture in the Biebrza Wetlands Retrieved from Sentinel-1 Imagery. Remote Sens. 2018, 10, 1979. [Google Scholar] [CrossRef] [Scilit]
  31. Cooper, H.M.; Bennett, E.; Blake, J.; Blyth, E.; Boorman, D.; Cooper, E.; Evans, J.; Fry, M.; Jenkins, A.; Morrison, R.; et al. COSMOS-UK: National soil moisture and hydrometeorology data for environmental science research. Earth Syst. Sci. Data 2021, 13, 1737–1757. [Google Scholar] [CrossRef] [Scilit]
  32. Ikonen, J.; Vehviläinen, J.; Rautiainen, K.; Smolander, T.; Lemmetyinen, J.; Bircher, S.; Pulliainen, J. The Sodankylä in situ soil moisture observation network: An example application of ESA CCI soil moisture product evaluation. Geosci. Instrum. Methods Data Syst. 2016, 5, 95–108. [Google Scholar] [CrossRef] [Scilit]
  33. Al-Yaari, A.; Dayau, S.; Chipeaux, C.; Aluome, C.; Kruszewski, A.; Loustau, D.; Wigneron, J.P. The AQUI Soil Moisture Network for Satellite Microwave Remote Sensing Validation in South-Western France. Remote Sens. 2018, 10, 1839. [Google Scholar] [CrossRef] [Scilit]
  34. Blöschl, G.; Blaschke, A.P.; Broer, M.; Bucher, C.; Carr, G.; Chen, X.; Eder, A.; Exner-Kittridge, M.; Farnleitner, A.; Flores-Orozco, A.; et al. The Hydrological Open Air Laboratory (HOAL) in Petzenkirchen: A hypothesis-driven observatory. Hydrol. Earth Syst. Sci. 2016, 20, 227–255. [Google Scholar] [CrossRef] [Scilit]
  35. Jensen, K.H.; Illangasekare, T.H. HOBE: A Hydrological Observatory. Vadose Zone J. 2011, 10, 1–7. [Google Scholar] [CrossRef] [Scilit]
  36. Osenga, E.C.; Vano, J.A.; Arnott, J.C. A community-supported weather and soil moisture monitoring database of the Roaring Fork catchment of the Colorado River Headwaters. Hydrol. Processes 2021, 35, e14081. [Google Scholar] [CrossRef] [Scilit]
  37. Smith, A.B.; Walker, J.P.; Western, A.W.; Young, R.I.; Ellett, K.M.; Pipunic, R.C.; Grayson, R.B.; Siriwardena, L.; Chiew, F.H.S.; Richter, H. The Murrumbidgee soil moisture monitoring network data set. Water Resour. Res. 2012, 48, W07701. [Google Scholar] [CrossRef] [Scilit]
  38. Larson, K.M.; Small, E.E.; Gutmann, E.D.; Bilich, A.L.; Braun, J.J.; Zavorotny, V.U. Use of GPS receivers as a soil moisture network for water cycle studies. Geophys. Res. Lett. 2008, 35, L24405. [Google Scholar] [CrossRef] [Scilit]
  39. González-Zamora, Á.; Sánchez, N.; Pablos, M.; Martínez-Fernández, J. CCI soil moisture assessment with SMOS soil moisture and in situ data under different environmental conditions and spatial scales in Spain. Remote Sens. Environ. 2019, 225, 469–482. [Google Scholar] [CrossRef] [Scilit]
  40. Ojo, E.R.; Bullock, P.R.; L’Heureux, J.; Powers, J.; McNairn, H.; Pacheco, A. Calibration and Evaluation of a Frequency Domain Reflectometry Sensor for Real-Time Soil Moisture Monitoring. Vadose Zone J. 2015, 14, vzj2014-08. [Google Scholar] [CrossRef] [Scilit]
  41. Schaefer, G.L.; Cosh, M.H.; Jackson, T.J. The USDA Natural Resources Conservation Service Soil Climate Analysis Network (SCAN). J. Atmos. Ocean. Technol. 2007, 24, 2073–2077. [Google Scholar] [CrossRef] [Scilit]
  42. Ardö, J. A 10-Year Dataset of Basic Meteorology and Soil Properties in Central Sudan. Dataset Pap. Geosci. 2013, 2013, 297973. [Google Scholar] [CrossRef] [Scilit]
  43. Zhao, T.; Shi, J.; Lv, L.; Xu, H.; Chen, D.; Cui, Q.; Jackson, T.J.; Yan, G.; Jia, L.; Chen, L.; et al. Soil moisture experiment in the Luan River supporting new satellite mission opportunities. Remote Sens. Environ. 2020, 240, 111680. [Google Scholar] [CrossRef] [Scilit]
  44. Albergel, C.; R¨udiger1, C.; Pellarin, T.; Calvet, J.-C.; Fritz, N.; Froissard, F.; Suquia, D.; Petitpa, A.; Piguet, B.; Martin, E. From near-surface to root-zone soil moisture using an exponential filter: An assessment of the method based on in-situ observations and model simulations. Hydrol. Earth Syst. Sci. 2008, 12, 1323–1337. [Google Scholar] [CrossRef] [Scilit]
  45. Zacharias, S.; Bogena, H.; Samaniego, L.; Mauder, M.; Fuß, R.; Pütz, T.; Frenzel, M.; Schwank, M.; Baessler, C.; Butterbach-Bahl, K.; et al. A Network of Terrestrial Environmental Observatories in Germany. Vadose Zone J. 2011, 10, 955–973. [Google Scholar] [CrossRef] [Scilit]
  46. Caldwell, T.G.; Bongiovanni, T.; Cosh, M.H.; Jackson, T.J.; Colliander, A.; Abolt, C.J.; Casteel, R.; Larson, T.; Scanlon, B.R.; Young, M.H. The Texas Soil Observation Network:A Comprehensive Soil Moisture Dataset for Remote Sensing and Land Surface Model Validation. Vadose Zone J. 2019, 18, 1–20. [Google Scholar] [CrossRef] [Scilit]
  47. Bell, J.E.; Palecki, M.A.; Baker, C.B.; Collins, W.G.; Lawrimore, J.H.; Leeper, R.D.; Hall, M.E.; Kochendorfer, J.; Meyers, T.P.; Wilson, T.; et al. U.S. Climate Reference Network Soil Moisture and Temperature Observations. J. Hydrometeorol. 2013, 14, 977–988. [Google Scholar] [CrossRef] [Scilit]
  48. Wang, J.; Wang, Y.; Bai, X.; Shao, W. Machine Learning-Based Soil Moisture Retrieval from Sentinel-1A Observations over the International Soil Moisture Networks. Remote Sens. 2026, 18, 1914. [Google Scholar] [CrossRef] [Scilit]
  49. Štrumbelj, E.; Kononenko, I. Explaining prediction models and individual predictions with feature contributions. Knowl. Inf. Syst. 2014, 41, 647–665. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Global spatial distribution of the ISMN observation network stations used in this study. Triangles denote the eight networks with sufficient matched samples for independent local domain XGBoost and SHAP modelling. Circles represent the other 16 networks incorporated only into the global pooled dataset, for which standalone local domain models are not developed. Colors distinguish different observation networks, as listed in the legend.
Figure 1. Global spatial distribution of the ISMN observation network stations used in this study. Triangles denote the eight networks with sufficient matched samples for independent local domain XGBoost and SHAP modelling. Circles represent the other 16 networks incorporated only into the global pooled dataset, for which standalone local domain models are not developed. Colors distinguish different observation networks, as listed in the legend.
Remotesensing 18 03006 g001
Figure 2. Colored density scatter plots of measured and estimated SM for the merged global ISMN dataset: (a) training dataset, (b) validation dataset. The black dashed line represents the 1:1 reference, and the solid red line denotes the linear regression fit between estimated and measured SM. Color bars indicate point sample density.
Figure 2. Colored density scatter plots of measured and estimated SM for the merged global ISMN dataset: (a) training dataset, (b) validation dataset. The black dashed line represents the 1:1 reference, and the solid red line denotes the linear regression fit between estimated and measured SM. Color bars indicate point sample density.
Remotesensing 18 03006 g002
Figure 3. Colored density scatter plots of measured and estimated SM for four representative monitoring networks. Left panels (a,c,e,g) correspond to training datasets and right panels (b,d,f,h) correspond to validation datasets: (a,b) SNOTEL, (c,d) SCAN, (e,f) USCRN, (g,h) COSMOS-UK. The black dashed line represents the 1:1 reference, and the solid red line denotes the linear regression fit between estimated and measured SM. Color bars indicate point sample density.
Figure 3. Colored density scatter plots of measured and estimated SM for four representative monitoring networks. Left panels (a,c,e,g) correspond to training datasets and right panels (b,d,f,h) correspond to validation datasets: (a,b) SNOTEL, (c,d) SCAN, (e,f) USCRN, (g,h) COSMOS-UK. The black dashed line represents the 1:1 reference, and the solid red line denotes the linear regression fit between estimated and measured SM. Color bars indicate point sample density.
Remotesensing 18 03006 g003
Figure 4. Pearson correlation coefficient heatmap of all input predictor variables for the merged global ISMN dataset. The color bar represents correlation values ranging from −1 (dark blue) to 1 (dark red).
Figure 4. Pearson correlation coefficient heatmap of all input predictor variables for the merged global ISMN dataset. The color bar represents correlation values ranging from −1 (dark blue) to 1 (dark red).
Remotesensing 18 03006 g004
Figure 5. Pearson correlation coefficient heatmaps of input predictors for four individual monitoring networks: (a) SNOTEL, (b) SCAN, (c) USCRN, (d) COSMOS-UK. The color bar represents correlation values ranging from −1 (dark blue) to 1 (dark red).
Figure 5. Pearson correlation coefficient heatmaps of input predictors for four individual monitoring networks: (a) SNOTEL, (b) SCAN, (c) USCRN, (d) COSMOS-UK. The color bar represents correlation values ranging from −1 (dark blue) to 1 (dark red).
Remotesensing 18 03006 g005
Figure 6. SHAP-based feature interpretability results for the merged global ISMN dataset: (a) importance ranking bar chart showing raw mean |SHAP| values, with embedded donut sub-plot presenting normalized percentage contributions relative to the total sum of mean |SHAP|; (b) direction of driving factors. The horizontal axis denotes SHAP values, and color gradients represent raw feature magnitude (red = high feature value, blue = low feature value).
Figure 6. SHAP-based feature interpretability results for the merged global ISMN dataset: (a) importance ranking bar chart showing raw mean |SHAP| values, with embedded donut sub-plot presenting normalized percentage contributions relative to the total sum of mean |SHAP|; (b) direction of driving factors. The horizontal axis denotes SHAP values, and color gradients represent raw feature magnitude (red = high feature value, blue = low feature value).
Remotesensing 18 03006 g006
Figure 7. The SHAP-based importance ranking and direction of driving factors for four individual networks: (a,b) SNOTEL, (c,d) SCAN, (e,f) USCRN, (g,h) COSMOS-UK. The horizontal axis denotes SHAP values, and color gradients represent raw feature magnitude (red = high feature value, blue = low feature value).
Figure 7. The SHAP-based importance ranking and direction of driving factors for four individual networks: (a,b) SNOTEL, (c,d) SCAN, (e,f) USCRN, (g,h) COSMOS-UK. The horizontal axis denotes SHAP values, and color gradients represent raw feature magnitude (red = high feature value, blue = low feature value).
Remotesensing 18 03006 g007
Table 1. The basic information of the ISMN SM networks across the world.
Table 1. The basic information of the ISMN SM networks across the world.
NetworkNo. of
Station Used
No. of
Relative Orbit
No. of
Samples
IGBP
Land Cover *
Reference
AMMA-CATCH7128110,12[28]
ARM153410310,12[29]
BIEBRZA-S-1183256710[30]
COSMOS-UK211116,9824,5,9,10,12,14[31]
FMI17685738,9,10[32]
FR_Aqui3328678[33]
HOAL322543912[34]
HOBE28565251,5,9,12,14[35]
iRON835069,10[36]
OZNET6151210,12[37]
PBO_H2O4587032,6,7,8,9,10,12[38]
REMEDHUS20316,7747,10,12[39]
RISMA227134212[40]
RSMN1110797712-
Ru_CFR11845-
SCAN1755136,4011,2,4,5,7,8,9,10,12,14[41]
SD_DEM1110210[42]
SMN-SDR334123010[43]
SMOSMANIA-SWATMEX21916,5595,8,9,10,12,14[44]
SNOTEL3703256,3661,5,7,8,9,10-
TAHMO314868,9,10,12,14-
TERENO5451005,9,12[45]
TxSON40260339,10[46]
USCRN904424,9121,2,4,5,6,7,8,9,10,12,14[47]
* 1–14 refers to evergreen needleleaf forests, evergreen broadleaf forests, and deciduous needleleaf forests, deciduous broadleaf forests, mixed forests, closed shrublands, open shrublands, woody savannas, savannas, grasslands, permanent wetlands, croplands, urban and built-up lands, and cropland/natural vegetation mosaic.
Table 2. Overview of dynamics and static predictor variables adopted for XGBoost-based SM retrieval, including variable descriptions, physical units and corresponding data sources.
Table 2. Overview of dynamics and static predictor variables adopted for XGBoost-based SM retrieval, including variable descriptions, physical units and corresponding data sources.
CategoryInput VariablesDescriptionData Source
Dynamic predictors
Sentinel-1AVVVV co-polarized backscatter (dB)Sentinel-1A IW GRD
VHVH cross-polarized backscatter (dB)
ThetaLocal incidence angle (°)
DOYSentinel-1A day of year
VegetationLAILeaf area index (m2/m2)MODIS MCD15A3H
NDVINormalized difference vegetation index (-)MODIS MCD13A2
EVIEnhanced vegetation index (-)
MeteorologyERA5-Land topsoil SM0–7 cm SM (m3/m3)ERA5-Land
ERA5-Land topsoil ST0–7 cm ST (°C)
Static predictors
GeospatialLatitudeGeographic latitude (°)Geographic data
LongitudeGeographic longitude (°)
DEMDigital elevation model (m)Topographic data
Soil textureSandSand percent (%)HWSD data
ClayClay percent (%)
Bulk densityBulk density of the top soil (g/cm3)
Table 3. Training and validation performance of the XGBoost model for SM retrieval across different monitoring networks.
Table 3. Training and validation performance of the XGBoost model for SM retrieval across different monitoring networks.
DatasetsTrainingValidation
Bias (m3/m3)RMSE (m3/m3)R2 (-)Bias (m3/m3)RMSE (m3/m3)R2 (-)
SNOTEL−0.0010.0300.907−0.0010.0480.764
SCAN−0.0010.0250.9580.0010.0450.856
USCRN−0.0010.0160.978−0.0010.0380.868
COSMOS-UK−0.0010.0140.990−0.0010.0410.906
REMEDHUS−0.0010.0060.995−0.0010.0220.945
SMOSMANIA-SWATMEX0.0010.0090.9920.0010.0320.903
FMI0.0010.0090.992−0.0010.0250.930
RSMN0.0010.0020.998−0.0010.0190.894
ISMN SM−0.0010.0430.8650.0010.0490.820
Table 4. Global-domain model performance under different inputs combinations from ablation experiments.
Table 4. Global-domain model performance under different inputs combinations from ablation experiments.
DatasetsTrainingValidation
Bias (m3/m3)RMSE (m3/m3)R2 (-)Bias (m3/m3)RMSE (m3/m3)R2 (-)
ERA5-SM only−0.0010.0940.3430.0010.0940.337
All inputs without ERA5-SM0.0010.0470.8350.0010.0540.782
All inputs−0.0010.0430.8650.0010.0490.820
Table 5. Grouped feature mean |SHAP| importance values for SM retrieval across global and regional datasets.
Table 5. Grouped feature mean |SHAP| importance values for SM retrieval across global and regional datasets.
FeatureISMNSNOTELSCANUSCRNCOSMOS-UKREMEDHUSSMOSMANIA
-SWATMEX
FMIRSMN
LAT+LON+DEM0.04480.03930.03870.03950.07390.05590.03720.07960.0320
SAND0.00740.01000.00800.01070.00480.00150.02730.00170.0012
CLAY0.00310.00430.00230.00210.00140.00070.00110.00000.0009
DENSITY0.00460.00340.00470.00380.00940.00020.00070.00000.0019
VV+VH0.00670.00630.01060.01140.00450.00320.00590.00260.0035
THETA0.00390.00390.00530.00450.00220.00190.00350.00180.0015
LAI+NDVI+EVI0.01960.01260.01720.01260.03150.02000.01760.00880.0148
ERA5-SM0.04450.02950.05490.04740.02320.03240.04330.01280.0191
ERA5-ST0.00250.00240.00340.00310.00370.00310.00460.00380.0018
DOY0.01100.02080.00770.01080.02010.00910.00760.00660.0042
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhou, S.; Wang, Y.; Bai, X.; Shao, W. Spatial Domain Dependence Evolution of Input Parameter Importance in Soil Moisture Retrieval Under the XGBoost and SHAP Framework. Remote Sens. 2026, 18, 3006. https://doi.org/10.3390/rs18173006

AMA Style

Zhou S, Wang Y, Bai X, Shao W. Spatial Domain Dependence Evolution of Input Parameter Importance in Soil Moisture Retrieval Under the XGBoost and SHAP Framework. Remote Sensing. 2026; 18(17):3006. https://doi.org/10.3390/rs18173006

Chicago/Turabian Style

Zhou, Siyu, Yuzhu Wang, Xiaojing Bai, and Wei Shao. 2026. "Spatial Domain Dependence Evolution of Input Parameter Importance in Soil Moisture Retrieval Under the XGBoost and SHAP Framework" Remote Sensing 18, no. 17: 3006. https://doi.org/10.3390/rs18173006

APA Style

Zhou, S., Wang, Y., Bai, X., & Shao, W. (2026). Spatial Domain Dependence Evolution of Input Parameter Importance in Soil Moisture Retrieval Under the XGBoost and SHAP Framework. Remote Sensing, 18(17), 3006. https://doi.org/10.3390/rs18173006

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop