1. Introduction
Rice feeds more than half of the world’s population, making stable production closely linked to food security and rural livelihoods—particularly in Asian monsoon regions, where yield fluctuations can have broad socioeconomic effects [
1,
2]. Lodging, the permanent displacement of rice stems from their upright position, is one of the most damaging stresses in rice production. It reduces yield, lowers grain quality, and interferes with mechanical harvesting [
3,
4,
5]. Timely mapping of lodged areas is therefore critical for disaster assessment, crop insurance, phenotyping, and field management [
6,
7,
8]. Compared with labor-intensive field surveys and coarse-resolution satellite observations, unmanned aerial vehicle (UAV) remote sensing offers a flexible, high-resolution, and nondestructive solution for field-scale lodging monitoring [
9,
10,
11,
12]. Consequently, UAV-based lodging assessment has become an important topic in precision agriculture.
Deep-learning-based semantic segmentation has recently emerged as the dominant paradigm for extracting lodging information from UAV optical imagery. Su et al. [
13] proposed LodgeNet, which integrates DenseNet blocks, dilated convolutions, and channel attention into a U-Net backbone, achieving 97.30% pixel accuracy on UAV rice-lodging imagery. Sun et al. [
14] developed RL-DeepLabv3+, a lightweight DeepLabV3+ variant with a channel-wise feature-pyramid backbone and depth-wise separable convolutions for real-time lodging detection on unmanned rice harvesters. Kang et al. [
15] introduced GloAN, a plug-in global-attention module that can be embedded into various CNN backbones to improve rice-lodging segmentation with limited computational overhead. Zhang et al. [
16] presented AAUConvNeXt, which couples a U-Net–ConvNeXt encoder with an intelligent optimization algorithm for automatic hyperparameter selection. While these studies demonstrate the feasibility of deep-learning-based lodging segmentation from UAV RGB imagery, they leave two important limitations unaddressed.
The first limitation concerns input modality. Most existing lodging-mapping methods rely solely on RGB imagery, despite the fact that spectral and textural contrast between lodged and healthy rice is often weak. Lodged canopies may resemble healthy canopies under varying illumination and growth-stage conditions, while bare soil and field ridges introduce confounding textures. Consequently, color information alone is frequently insufficient for separating crop and background regions [
6,
7,
8]. Moreover, lodging is inherently a structural phenomenon: it reduces canopy height and weakens vertical structure—three-dimensional cues that cannot be fully captured by two-dimensional appearance features [
3,
17,
18,
19]. Digital surface models (DSMs), derived from photogrammetric reconstruction or LiDAR, encode canopy-height and surface-structure information [
20,
21]. Studies on wheat lodging have shown that RGB–DSM fusion improves accuracy over RGB alone [
20], and multimodal remote-sensing research broadly indicates that elevation or geometric cues can strengthen spectral features for land-cover classification [
22,
23,
24,
25,
26]. These findings suggest that DSM is a promising information source for rice-lodging segmentation, yet its integration into a dedicated framework remains underexplored.
The second limitation is annotation cost. Semantic segmentation demands dense pixel-level labels, which are expensive to produce for lodging scenes due to heterogeneous field conditions and gradual, often ambiguous boundaries [
7,
16]. Most existing lodging-segmentation studies operate in a fully supervised setting, assuming the availability of large annotated training sets [
13,
14,
16,
27]. Semi-supervised semantic segmentation addresses this issue by combining a small labeled subset with abundant unlabeled imagery. Teacher–student frameworks with weak-to-strong consistency, such as UniMatch [
28] and UniMatch V2 [
29], have substantially narrowed the gap between semi-supervised and fully supervised segmentation. Remote-sensing methods like RSProtoSemiSeg [
30] further adapt semi-supervised learning to RGB imagery through prototype-based regularization. However, these methods remain predominantly appearance-driven: their pseudo-label selection or weighting relies mainly on RGB semantic confidence, without leveraging DSM geometry to verify uncertain boundaries. In contrast, RGB–DSM segmentation networks such as CMGFNet [
22], LMFNet [
24], and GIMMNet [
26] focus on supervised cross-modal feature fusion and assume dense labels. They do not address pseudo-label noise in unlabeled multimodal data, nor do they explicitly assess whether local DSM evidence is reliable. Consequently, a direct combination of UniMatch V2 with a standard RGB–DSM fusion module is insufficient for limited-label rice-lodging segmentation. An effective method must exploit DSM height and boundary cues while simultaneously managing unreliable geometry and ambiguous pseudo-label supervision.
To address these limitations, we propose Geometry-Guided UniMatch (GUMatch), a semi-supervised segmentation framework for registered UAV RGB and DSM imagery. GUMatch is neither a simple DSM-input extension of UniMatch V2 nor a semi-supervised version of a fully supervised RGB–DSM fusion network. It retains the effective teacher–student optimization scaffold of UniMatch V2 but fundamentally changes how geometric information enters both representation learning and unlabeled supervision. The core idea is to treat DSM not as a uniformly fused auxiliary channel but as a reliability-aware geometric prior for height and boundary reasoning. This design is motivated by two observations: (1) DSM quality is spatially nonuniform because height estimates can be affected by canopy occlusion, reconstruction noise, surface discontinuities, and residual mis-registration; and (2) geometric cues are most valuable where RGB appearance is ambiguous, particularly near lodging boundaries and mixed-canopy areas.
Based on these observations, GUMatch integrates three geometry-guided components at different levels of the semi-supervised pipeline: Adaptive Geometric Prompting (AGP) for reliability-aware decoder prompting, Geometry-Calibrated Pseudo-Label Learning (GPL) for geometry-calibrated pseudo-label supervision, and Boundary-Aware Geometric Regularization (BGR) for conservative boundary refinement. This decomposition allows DSM evidence to guide representation learning and unlabeled supervision without treating every DSM response as equally trustworthy. During weak-to-strong training, strong appearance perturbations are applied only to RGB images, while DSM inputs follow synchronized geometric transformations to preserve height meaning and cross-modal alignment.
The main contributions of this study are summarized as follows:
We propose GUMatch, a geometry-guided semi-supervised multimodal segmentation framework for UAV rice-lodging mapping. Unlike single-modal semi-supervised methods, GUMatch leverages DSM to guide both feature decoding and pseudo-label reliability. Unlike supervised RGB–DSM fusion networks, it is designed for unlabeled RGB–DSM pairs and explicitly models unreliable local geometry.
We design AGP, a reliability-aware decoder prompting module that selectively injects DSM height and boundary evidence based on local geometric reliability and RGB–DSM compatibility.
We develop a coupled geometry-calibrated supervision strategy combining GPL and BGR to suppress boundary-risk pseudo-label noise and apply conservative geometric boundary constraints.
We construct a three-parcel UAV rice-lodging collection and evaluate GUMatch under a strict parcel-level split. HA-P2 serves for training and validation with 10%, 20%, and 40% labeled ratios; HA-P1 is the held-out Huai’an test parcel; and WX-P3 provides external cross-region testing. Under the DINOv2-B setting, GUMatch improves over UniMatch V2 by 5.39 mIoU points on WX-P3. The results demonstrate consistent gains over representative semi-supervised baselines, particularly in lodged-rice boundary recovery and heterogeneous-background suppression.
2. Dataset Acquisition
2.1. Study Area and Data Acquisition
The UAV rice-lodging data comprise three parcel-level subsets collected from two rice-production regions in Jiangsu Province, China. Two parcels were acquired at Baimahu Farm, a national seed-production base in Huai’an District, Huai’an City: the 4.2 ha parcel shown in
Figure 1a, denoted as
HA-P1, and the 4.9 ha parcel in
Figure 1c, denoted as
HA-P2. Located near
N and
E on the eastern Jianghuai Plain, the farm features flat terrain suitable for large-scale rice cultivation. The region has a temperate monsoon climate, with an annual mean temperature of 13.8–14.8 °C, annual precipitation of 906–1007 mm, and approximately 2160 h of sunshine annually. Paddy soil and fluvo-aquic soil are the dominant soil types.
A third 2.7 ha parcel from Binhu District, Wuxi City, denoted as
WX-P3 and shown in
Figure 1d, is also included. Located near
N and
E on the northern Taihu Plain, this region has a northern subtropical humid monsoon climate, with an annual mean temperature of 15.4–16.5 °C, annual precipitation of 1050–1150 mm, and approximately 1920 h of sunshine annually. Paddy soil and fluvo-aquic soil are again the main soil types. The Huai’an parcels were imaged during the growing stage, whereas
WX-P3 was imaged near harvest. Consequently,
WX-P3 contains more complex backgrounds due to stronger shadow effects and more exposed soil.
Figure 1 provides an overview of the three-parcel collection:
Figure 1a,c shows the global previews of
HA-P1 and
HA-P2,
Figure 1b marks the locations of Huai’an and Wuxi within Jiangsu Province,
Figure 1d shows the global preview of
WX-P3, and
Figure 1e presents representative lodged-rice image samples. In our evaluation protocol,
HA-P2 serves as the model-development parcel,
HA-P1 is the held-out Huai’an test parcel, and
WX-P3 provides the external cross-region test.
UAV observations were acquired with a DJI Matrice 300 RTK platform equipped with a Zenmuse L2 LiDAR payload. The system synchronously collected RGB imagery and point-cloud data, from which co-registered orthophoto mosaics and DSM products were generated.
Table 1 summarizes the main platform, sensor, flight, and image-product parameters.
2.2. RGB and DSM Preprocessing
The raw UAV observations were processed in DJI Terra to generate co-registered multimodal products. The workflow included image mosaicking, radiometric correction, color correction, spatial registration, point-cloud reconstruction, and the generation of digital orthophoto maps and DSMs. The resulting RGB images and DSMs were exported as paired .tif files to preserve georeferencing information and pixel-level alignment. This alignment is essential because the proposed framework uses DSM not only as an input modality but also as a source of geometric priors for decoder prompting, pseudo-label calibration, and boundary regularization.
Pixel-wise annotations were produced in ArcGIS Pro 3.1 on the high-resolution orthophoto mosaics. Each pixel was assigned to one of three semantic categories: lodged rice, healthy rice, or background. Lodged rice denotes areas with visible tilting, flattening, or canopy collapse following wind, heavy rainfall, or related disturbances; these regions typically exhibit disordered texture, darker tones, and irregular canopy structure. Healthy rice denotes upright rice with normal growth, appearing as brighter and more spatially homogeneous canopy cover. Field ridges, roads, bare soil, shadows, water surfaces, weeds, and other nontarget objects were assigned to the background class.
The annotation workflow comprised three steps. First, trained annotators delineated category polygons on the georeferenced orthophoto mosaics according to the visual interpretation rules above. Second, experts in UAV agricultural image interpretation reviewed the polygon layers and verified category consistency across parcels. Third, boundary correction was performed by jointly inspecting the RGB mosaic, DSM product, and annotation layer, particularly near gradual lodging transitions, narrow ridges, shadowed areas, and mixed crop–background boundaries. The final vector annotations were rasterized to the same grid as the RGB and DSM products and exported as pixel-level label masks.
Figure 2 shows representative examples.
2.3. Parcel-Level Split and Semi-Supervised Setting
To evaluate annotation-efficient rice-lodging mapping while reducing spatial leakage, we adopted a parcel-level semi-supervised split protocol with explicit cross-scene testing. In the main HA-P2/HA-P1 benchmark, HA-P2 was used for training and validation, while HA-P1 served as the held-out in-region test parcel. The external WX-P3 parcel was reserved for direct cross-region testing. No tile from HA-P1 or WX-P3 was used for parameter updates, model selection, or hyperparameter tuning. Random tile-level splitting was avoided because neighboring UAV tiles often share similar texture, illumination, canopy structure, and height patterns, which can lead to overly optimistic accuracy estimates.
Appendix A provides spatial partition maps with semi-transparent overlays for all three parcels, complementing
Figure 1 and allowing readers to inspect the spatial separation among the
HA-P2 model-development regions, the
HA-P1 in-region test parcel, and the
WX-P3 external test parcel.
For the main benchmark, HA-P2 was first divided into spatially contiguous training and validation regions. All parcels were then synchronously cropped into pixel RGB and DSM tiles. Tiles with large blank regions, insufficient valid content, or poor image quality were removed. The main benchmark contains approximately 1400 high-quality paired tiles, and the external WX-P3 set contributes 300 additional fully annotated paired tiles for direct cross-region testing. Under the semi-supervised setting, only 10%, 20%, or 40% of the HA-P2 training tiles retained pixel-level annotations, while the remaining training tiles were used as unlabeled data. The validation set was used for model selection and hyperparameter tuning but not for parameter updates. The HA-P1 test set was fixed across all labeled ratios and used only for final evaluation. Unless otherwise stated, each experiment was repeated three times with different labeled and unlabeled splits using seeds 0, 1, and 2, and the mean performance is reported.
To further examine cross-region generalization, we directly evaluated the 40% labeled
HA-P2-source models on
WX-P3. All training, validation, and model-selection operations remained confined to
HA-P2;
WX-P3 was used only for final testing and was not involved in parameter updates or hyperparameter tuning. The
WX-P3 results are reported in
Section 4.4.
This protocol defines the learning problem considered in the following
Section 3. The labeled set is denoted as
where
is the RGB tile,
is the aligned DSM tile, and
is the pixel-level annotation. The unlabeled set is denoted as
GUMatch uses paired RGB and DSM inputs in both sets. The labels in
provide supervised category learning, whereas the unlabeled pairs in
are used to construct geometry-guided pseudo-label supervision during teacher–student training.
3. Method
To improve readability, we use bold uppercase symbols for image-level maps or tensors and lower-case indexed symbols for their pixel values, e.g., and . Algorithms 1 and 2 use the same map-level symbols as the corresponding formulas. The main abbreviations and symbols are also summarized in the “Abbreviations and Symbols” section at the end of the paper.
3.1. Framework Overview
Figure 3 compares UniMatch V2 with the proposed GUMatch. UniMatch V2 [
29] employs an EMA teacher–student framework with weak-to-strong consistency. The teacher generates pseudo-labels from a weakly augmented view, while the student learns from two strong views derived from the same image. Complementary channel-wise dropout is applied to strong-view features before decoding, enabling the same weak-view supervision to regularize both streams. GUMatch retains this effective optimization scaffold but introduces a DSM encoder, a DSM edge prior, reliability-aware decoder prompting, geometry-calibrated pseudo-label weights, and a boundary regularizer. These components allow DSM evidence to guide both feature decoding and teacher–student supervision.
Unlike supervised RGB–DSM segmentation networks such as CMGFNet and LMFNet, which learn cross-modal feature interaction from densely labeled data, GUMatch focuses on unlabeled RGB–DSM pairs and explicitly assesses whether DSM evidence should be trusted before modifying decoder features or pseudo-label supervision. This distinction is critical for UAV rice-lodging scenes, where photogrammetric DSMs may contain local noise, voids, and residual mis-registration.
The data flow in GUMatch links AGP, GPL, and BGR through a shared DSM edge prior. For each RGB–DSM branch, the DSM encoder produces a prompt pyramid and an edge-prior map. AGP uses these signals to gate the RGB decoder features. On the weak view, the teacher prediction provides pseudo-labels, confidence scores, and a semantic boundary response. GPL combines these teacher outputs with the DSM edge prior to generate calibrated confidence weights, valid pseudo-label masks, and boundary bands in weak-view coordinates. The same geometric transformations then transfer these targets and the DSM edge prior to the two strong views. BGR is applied to the strong-view student predictions only within the transferred valid boundary regions. Thus, the edge prior flows from decoder prompting to pseudo-label calibration and boundary regularization, while GPL controls where the boundary loss is permitted to act.
During weak-to-strong training, strong appearance perturbations are applied only to RGB, whereas DSM follows synchronized geometric transformations to preserve height meaning and RGB–DSM alignment. GUMatch uses paired RGB–DSM inputs during both training and inference, so the DSM branch is not merely a training-only auxiliary module. AGP requires DSM to produce gated decoder features, and the geometry-guided losses use DSM to define boundary-risk and edge-prior signals during training.
3.2. Adaptive Geometric Prompting Module
AGP transforms DSM from an auxiliary input into a structured geometric prior through a selective three-stage design: multiscale DSM prompt construction, edge-prior extraction, and reliability-aware prompt gating. This design mitigates the risk of uniformly injecting DSM features into the decoder when DSMs contain reconstruction noise, local voids, or mis-registration.
Figure 4 provides the network-level view. For branch ★, the DSM branch produces the prompt pyramid
, the geometric branch derives the edge prior
, and scale matching supplies
to the gate. AGP then combines these terms with the RGB decoder feature
to compute the gated residual correction before DPT decoding.
3.2.1. DSM Encoder
RGB imagery and DSM provide complementary but fundamentally different evidence: RGB primarily contributes semantic appearance, whereas DSM captures local elevation changes and surface discontinuities. We therefore adopt an asymmetric dual-encoder design. Let
and
denote the RGB and DSM encoders, respectively. For branch ★, which denotes a labeled, weak, or strong view, multiscale features are extracted as
where
in our implementation. The RGB encoder
is instantiated by DINOv2 [
31], and
is a lightweight hierarchical convolutional encoder. The projection
maps DSM features to the decoder channel dimension, so
acts as a scale-aligned carrier of geometric evidence rather than as a second semantic stream.
3.2.2. DSM Edge Prior Construction
AGP also extracts a class-agnostic DSM edge prior as the shared geometric reference for prompt gating, GPL, and BGR. Because raw DSM gradients may contain large local outliers from reconstruction artifacts, we define a robust normalization operator for any nonnegative response map
:
where
is the pixel set and
prevents division by zero. Equation (
4) scales responses by the image-wide average magnitude and saturates smoothly, making it less sensitive to isolated DSM artifacts or spurious prediction gradients than max-based normalization. The class-agnostic DSM edge-magnitude map is then
where
are Sobel kernels,
is Gaussian smoothing, and ∗ denotes convolution. For the labeled and weak branches, Equation (
4) gives
and
from
and
. For the strong branches, we transform the raw weak-view edge magnitude before normalization to avoid CutMix-induced artificial edges:
When CutMix pastes a region from another unlabeled sample, the paired raw edge magnitude is pasted into the same region before normalization, ensuring that
remains tied to scene geometry rather than augmentation artifacts.
3.2.3. Reliability-Aware Prompt Gating and Decoder Injection
Given the prompt features and edge prior, AGP estimates where DSM should influence decoding. Let
be the RGB feature entering AGP; it equals the raw RGB feature on the labeled and weak branches and the complementary-dropout feature on the strong branches (
Section 3.3.1). The spatial gate is
where
is scale matching,
is feature normalization, and
are lightweight convolutional predictors. The first term measures intrinsic geometric reliability, and the second term measures RGB–DSM compatibility. Both terms are single-channel spatial maps broadcast along feature channels.
At scale
k, the DPT decoder consumes the fused feature
where
is a learnable scale-specific coefficient. The residual formulation preserves RGB as the primary semantic carrier and allows DSM to contribute only through a gated correction. We denote the complete segmentation network as
.
3.3. Geometry-Guided Semi-Supervised Optimization
After AGP establishes the shared geometric prior, unlabeled training must use it without corrupting the physical meaning of DSM. GUMatch achieves this through a modality-preserving weak-to-strong protocol, geometry-calibrated pseudo-label supervision, and a conservative boundary refinement loss.
3.3.1. Modality-Preserving Weak-to-Strong Training Protocol
Weak-to-strong semi-supervised learning requires strong student perturbations, but RGB–DSM training must preserve DSM height meaning. For each unlabeled sample
, we first build a weak teacher view
, where
comprises geometry-preserving resize, crop, and horizontal flip shared by RGB and DSM. Two student strong views are derived as
where
is RGB-only appearance perturbation and
is the synchronized spatial transform, primarily CutMix in our implementation. This asymmetric protocol perturbs RGB appearance while keeping DSM photometrically unchanged; we refer to it as ASV in the ablation study.
We retain complementary dropout only on the RGB strong-view features. To avoid conflict with the pseudo-label valid mask
used in GPL, the channel-wise dropout masks are denoted by
and
, with
and
. The same mask pair is shared across all selected RGB backbone scales:
For the weak teacher and labeled branches, complementary dropout is disabled, so
and
. DSM prompt features
are never masked.
3.3.2. Geometry-Calibrated Pseudo-Label Supervision
In lodging scenes, the teacher’s most harmful errors typically arise near ambiguous boundaries, where RGB appearance is weakly discriminative but supervision decisions carry high geometric consequence. GPL therefore does not alter the teacher’s semantic prediction itself; instead, it regulates how much each pseudo-label should be trusted, with geometry-guided down-weighting activated only within teacher-identified boundary-risk regions.
The student and teacher networks are parameterized by
and
, respectively, with
updated as the EMA of
with momentum
. On the weak view
, the teacher produces logits
and probabilities
. For each pixel
i, the pseudo-label is
and the raw confidence is
. The corresponding map-level symbols used in Algorithm 1 are
and
. To localize where calibration is needed, we derive the teacher boundary response
and activate calibration only within the teacher-derived boundary band
where
is max pooling with radius
and
controls the band width. The semantic-geometric agreement is
where
controls the tolerance to RGB–DSM boundary mismatch. As
outside the boundary band, nonboundary pixels are not penalized. The calibrated confidence is
where
controls the down-weighting strength. The factor in Equation (
14) lies in
, so geometry can suppress but never inflate teacher confidence. The valid pseudo-label mask is
.
The map-level targets
,
,
, and
are generated in weak-view coordinates and transferred to each strong branch by the same
used in Equation (
9);
follows Equation (
6). Discrete targets use nearest-neighbor transformation, whereas continuous confidence and edge maps use bilinear interpolation before CutMix pasting. Algorithm 1 summarizes this calibration-and-transfer procedure with symbols matched to Equations (
11)–(
14).
Given the transformed teacher targets, the student produces strong-view logits
and probabilities
for
. The resulting per-view unsupervised loss is a confidence-weighted cross-entropy:
where
is pixel-wise cross-entropy and
prevents division by zero. The final unlabeled objective averages the two strong views,
. Equation (
15) maintains the weak-to-strong consistency principle of UniMatch V2 while employing modality-preserving strong views and geometry-calibrated pseudo-label weights. RGB is optimized for appearance robustness, whereas DSM remains a stable geometric condition across the two strong branches.
| Algorithm 1 GPL target generation and transfer with formula-matched notation |
- Input:
weak view , weak-view DSM prior - Input:
EMA teacher , strong-view operators - Output:
transformed targets - 1:
and - 2:
and ▹ map forms of and - 3:
Compute teacher boundary response by Equation ( 11) - 4:
Compute boundary-band indicator by Equation ( 12) - 5:
Compute geometry agreement by Equation ( 13) - 6:
Compute calibrated confidence from , , and by Equation ( 14) - 7:
- 8:
for
do - 9:
and - 10:
and - 11:
return
|
3.3.3. Boundary-Aware Geometric Regularization
GPL improves pseudo-label reliability, but the supervision remains primarily category-level and does not directly constrain boundary shape. We therefore add BGR as a conservative refinement loss atop GPL-filtered supervision. It avoids global alignment to all DSM edges, as field ridges, tractor ruts, and drainage channels can create strong height discontinuities unrelated to lodging boundaries.
For each strong view, the student boundary response is extracted from the soft prediction analogously to Equation (
11):
Using the transformed DSM prior
from Equation (
6), BGR activates boundary supervision only at the intersection of valid pseudo-labels, boundary candidates, and nontrivial DSM edges:
where
is the DSM-edge threshold. The per-view boundary regularization loss is
and the overall boundary regularizer averages the two strong views:
Thus, Equation (
18) refines GPL-filtered supervision only where pseudo-labels and DSM evidence are jointly reliable.
Here,
is a semantic-boundary response, whereas
is a DSM height-discontinuity response. We use L1 amplitude matching because Equation (
4) puts the two responses into a comparable dynamic range under the mask in Equation (
17). Alternatives such as Dice on binarized maps or active-contour-style losses [
32] were less stable in our experiments, as they may lose useful gradients on weak boundaries or fit noisy DSM ridges too aggressively. reports the empirical comparison.
3.4. Supervised Loss and Overall Objective
For labeled samples, the student predicts logits
, and the supervised loss is the standard pixel-wise cross-entropy:
over the labeled pixel set
. The overall training objective combines the supervised loss with the two unsupervised components:
where
and
balance the two unsupervised terms. The teacher parameters are updated as the EMA of the student with momentum
. Algorithm 2 summarizes the full training routine and references the equations used in each step.
| Algorithm 2 Overall training procedure of GUMatch |
- Input:
labeled set , unlabeled set - Input:
student , EMA teacher , weak augmentation - Input:
strong-view operators , , , - Output:
trained student and teacher - 1:
Initialize - 2:
for each training iteration do - 3:
Sample and - 4:
- 5:
Compute by Equation ( 5) and by Equation ( 4) - 6:
for do - 7:
Generate by Equation ( 9) - 8:
Generate by Equation ( 6) - 9:
- 10:
for do - 11:
- 12:
Compute by Equation ( 15) and by Equation ( 18) - 13:
and compute - 14:
and - 15:
Compute by Equation ( 21) - 16:
Update by back-propagating - 17:
Update - 18:
return
|
4. Experimental Results and Analysis
4.1. Experimental Setting
All experiments were run on Ubuntu 20.04 using Python 3.11.9, PyTorch 2.5.0, and CUDA 12.4 under Miniconda. The workstation contained an Intel Xeon 8360Y CPU (2.40 GHz) and one NVIDIA RTX 3090 GPU (24 GB).
Following the parcel-level protocol in
Section 2, the main
HA-P2/
HA-P1 benchmark contained 1400 RGB–DSM tiles: 900 training tiles from
HA-P2, 200 validation tiles from
HA-P2, and 300 in-region test tiles from
HA-P1. The external
WX-P3 set contained 300 fully annotated tiles and was used only for direct cross-region testing. All models were trained and selected only on
HA-P2;
HA-P1 and
WX-P3 were used only as final test parcels.
The labeled ratios were applied only to the HA-P2 training set: 10%, 20%, and 40% correspond to 90, 180, and 360 labeled tiles, respectively, with the remaining training tiles used as unlabeled data. Inputs were cropped to , and all models were trained for 200 epochs. Each iteration drew one labeled mini-batch and one unlabeled mini-batch, each with batch size 2. Validation and testing used batch size 1. Weak augmentation applied synchronized random resize, random crop, and horizontal flip to RGB and DSM. Strong student views further applied RGB-only color jitter, grayscale conversion, Gaussian blur, and CutMix (probability 0.5), while DSM retained only the same geometric transforms to preserve height structure.
We used AdamW (
,
, weight decay 0.01). For DINOv2-B, the pretrained RGB backbone used an initial learning rate of
; the decoder, DSM encoder, AGP module, and other newly initialized layers used a
multiplier. We adopted a polynomial decay schedule,
, with
t and
T the current and total iterations. DINOv2-B was initialized from the official checkpoint, and RN-101 from ImageNet weights; the DSM encoder, AGP module, and task-specific decoder layers were randomly initialized. The pseudo-label confidence threshold was
. The unsupervised consistency and boundary regularization weights were
and
. The EMA teacher momentum was updated as
. For GPL and BGR, we set
,
,
,
, and
. Each setting was repeated with seeds 0, 1, and 2, and we report the mean. For the seed-matched method comparisons in
Table 2, we used a paired nonparametric sign test.
4.2. Evaluation Metrics
We adopt mean intersection over union (mIoU) as the primary metric and report class-wise IoU for per-category analysis. All metrics are computed on the corresponding fixed test parcel over three classes: lodged rice, healthy rice, and background. As lodging boundaries are often gradual and spatially ambiguous, we further report the boundary F1-score (BF1) in the ablation and robustness analyses. We extract class-agnostic boundaries from the predicted and ground-truth maps and count a boundary pixel as matched if it lies within a small tolerance band of the other boundary. All metrics are reported as percentages.
4.3. Comparison with State of the Art
Table 2 is organized into three comparison groups. Panel A reports the RN-101 benchmark, including single-modal semi-supervised baselines, supervised RGB–DSM references trained on the same labeled subsets, simple RGB–DSM UniMatch V2 variants, and GUMatch with RN-101. Panel B reports the multimodal semi-supervised baselines M3L and DepMatch with their corresponding encoders. Panel C focuses on the DINOv2-B variants of UniMatch V2 and GUMatch, isolating the effect of geometry-guided learning from the backbone upgrade. The 100% fully supervised RGB and RGB–DSM references are reported separately in
Table 3.
Under the RN-101 setting, GUMatch consistently outperforms the best single-modal baseline (RSProtoSemiSeg) and the strongest simple RGB–DSM UniMatch V2 variant across all three labeled ratios. These results indicate that the gain does not come from adding DSM as a fixed auxiliary input alone. M3L and DepMatch provide a more stringent multimodal semi-supervised comparison. Against these directly related baselines, GUMatch with DINOv2-B achieves the highest mIoU under all three labeled ratios. Panel C further confirms that the improvement is not simply a backbone effect: with the same DINOv2-B backbone, GUMatch consistently outperforms UniMatch V2 and its decoder-add RGB–DSM extension.
Table 3 provides a direct annotation-efficiency comparison. With all 900 labeled tiles, the RGB–DSM full-supervision version of GUMatch outperforms the RGB-only UniMatch V2 reference by 1.72 mIoU points, with larger gains in lodged rice and background IoU than in healthy rice IoU. The semi-supervised rows show that as the labeled ratio increases from 10% to 40%, GUMatch improves from 82.64% to 86.72% mIoU. Notably, the 40% semi-supervised model surpasses the pure RGB 100% full-supervision reference in mIoU, lodged rice IoU, precision, and F1, though it remains below the RGB–DSM 100% full-supervision model.
To provide a more complete view of segmentation quality beyond mIoU,
Table 4 summarizes the class-wise IoU and mean precision/recall/F1 results across all three labeled ratios.
The detailed metrics in
Table 4 reveal complementary behavior among the multimodal semi-supervised baselines. M3L obtains competitive lodged-rice IoU at 10% labels and the highest background IoU at 40% labels, but its healthy-rice IoU is consistently lower than that of GUMatch with DINOv2-B. DepMatch gives the highest recall in all three settings, but its lodged-rice IoU, background IoU, and F1-score are generally lower than those of GUMatch. GUMatch with DINOv2-B achieves the highest F1-score, lodged-rice IoU, healthy-rice IoU, and precision across all three settings.
The comparison with UniMatch V2 using DINOv2-B further pinpoints where the gains occur. GUMatch improves lodged-rice IoU by 2.23–2.53 points, healthy-rice IoU by 1.27–1.46 points, background IoU by 3.15–3.25 points, and F1-score by 1.36–1.57 points across the three labeled ratios. The largest class-wise gain is observed for background IoU, which is consistent with the qualitative examples where geometry-guided calibration reduces appearance-driven background confusion.
Figure 5 provides local zoom-in comparisons on two representative difficult regions containing mixed crop–background boundaries, fragmented lodged-rice areas, and confusing background patterns. M3L and DepMatch reduce some errors compared with weaker baselines, but still show local fragmentation or boundary discontinuities. GUMatch recovers more complete lodged-rice areas, suppresses background false positives, and preserves clearer local boundaries. The DINOv2-B version is visually closest to the ground truth in both examples.
4.4. Source-Parcel Generalization from HA-P2 to WX-P3
To evaluate cross-region transfer, we conducted an external source-parcel generalization experiment on
WX-P3. All models were trained and selected under the 40% labeled setting using only
HA-P2. The
WX-P3 parcel was then used only for direct testing, with no tile used for parameter updates, model selection, or hyperparameter tuning.
Table 5 reports the results.
A clear cross-region domain gap exists. With DINOv2-B, GUMatch drops from 86.72% mIoU on HA-P1 to 68.34% mIoU on WX-P3, confirming that the Wuxi scene is substantially more challenging. Nevertheless, GUMatch remains the strongest method on WX-P3, improving over UniMatch V2 with the same backbone by 5.39 mIoU points. The class-wise results indicate that geometry-guided prompting and pseudo-label calibration help suppress background-driven errors under distribution shift, though lodged-rice IoU is not the highest among all methods.
Figure 6 provides a visual companion: several baselines produce fragmented masks or large background-confusion regions, while GUMatch gives more spatially coherent predictions.
4.5. Ablation Studies
Unless otherwise stated, all ablation experiments use the 20% labeled split with DINOv2-B. We focus on the three main components of GUMatch: AGP for geometry-guided feature prompting, GPL for pseudo-label calibration, and BGR for boundary refinement. We also evaluate the edge-prior extraction and response-normalization functions.
4.5.1. Contribution of the Core Design Elements
Table 6 shows the progression from a simple RGB–DSM extension to the full geometry-guided design. Fixed decoder summation of DSM features gives only a modest improvement, suggesting that simple multimodal stacking is insufficient. AGP improves the result through reliability gating, GPL further enhances pseudo-label reliability, and BGR gives the best BF1 and mIoU.
4.5.2. Fusion Strategy and Injection Position
We introduce DSM as an adaptive geometric prompt after RGB complementary dropout and before DPT decoding.
Table 7 supports this choice. Early concatenation and backbone fusion are weaker than decoder-side integration. Injecting DSM after RGB complementary dropout outperforms injection before dropout.
4.5.3. Teacher–Student Modality Configuration
Table 8 shows the best performance when both teacher and student receive RGB–DSM inputs. This symmetric setting allows geometry to contribute to both pseudo-label generation and student learning.
4.5.4. Pseudo-Label Quality Analysis
We evaluate pseudo-label quality on the unlabeled training set using hidden annotations only for post hoc evaluation.
Table 9 shows that geometry-calibrated confidence retains slightly fewer pixels than raw semantic confidence from the same teacher, but improves precision, recall, and BF1.
4.5.5. Edge-Prior Extraction and Normalization
The DSM edge prior is shared by AGP, GPL, and BGR. We conduct two controlled ablations: fixing normalization and varying edge extraction, and fixing edge extraction and varying normalization.
Table 10 shows that Gaussian-smoothed Sobel gives the best overall result. Removing smoothing weakens both mIoU and BF1. Scharr yields close results but offers no clear advantage. LoG and Canny perform worse, especially in BF1. The normalization results support the proposed mean-exponential form, which preserves relative edge strength while smoothly saturating large outliers.
Figure 7 provides mechanism-oriented qualitative analysis. Panel A shows progressive improvement from UniMatch V2 to the full model. Panel B visualizes GPL’s effect on pseudo-label quality. Panel C demonstrates BGR’s boundary refinement.
4.6. Robustness, Sensitivity, and Efficiency Analysis
Remote-sensing DSMs may contain spatially correlated reconstruction failures. We therefore focus on structured DSM degradation that better reflects photogrammetric failure patterns, with simpler controlled perturbations reported in
Appendix B. Unless otherwise stated, this analysis uses the 20% labeled split with DINOv2-B and compares GUMatch with DepMatch.
4.6.1. Structured DSM Degradation Protocol
We evaluate three structured degradation modes: (1) spatially correlated height distortion, (2) regional topological collapse, and (3) interpolated voids and smoothed holes. Their physical interpretations and severity-specific parameter settings are summarized in
Table 11. Degradation is applied to unlabeled DSMs during semi-supervised training; labeled samples, validation, and test sets remain clean.
Table 12 shows that GUMatch maintains higher absolute mIoU and BF1 than DepMatch under all settings, with smaller
mIoU values. Regional topological collapse remains the most challenging degradation.
4.6.2. Hyperparameter Sensitivity
We sweep the four most influential parameters around their defaults.
Table 13 shows that mIoU varies by less than 0.7 points across the swept ranges, and GUMatch’s superiority over the strongest non-GUMatch baseline is preserved.
We additionally evaluated three boundary objectives within BGR.
Table 14 shows that the L1 formulation yields the best BF1.
4.6.3. Computational Cost
Table 15 shows that DINOv2-B GUMatch adds only moderate parameters and FLOPs compared with UniMatch V2, with the extra cost mainly from the lightweight DSM encoder and adaptive prompting layers.