1. Introduction
Urban forest parks are important spatial systems that connect ecological processes with public needs in built environments. Through canopy shading, evapotranspiration, pollutant deposition, and carbon storage, their vegetation contributes to the regulation of urban heat, water, carbon, and pollution cycles and affects environmental quality and human health [
1,
2,
3,
4,
5,
6]. Tree species composition and spatial arrangement are fundamental descriptors of structural and functional heterogeneity in urban forests and provide a basis for ecosystem-service assessment, biodiversity conservation, stand-health monitoring, and pest and disease risk management [
7]. Compared with natural forests, urban forests commonly exhibit complex species composition, intensive anthropogenic disturbance, interlaced crowns, and diverse background land covers. Rapid and accurate acquisition of tree species information is therefore a foundational technical requirement for urban forest inventories and refined management [
8,
9,
10].
Conventional tree species surveys primarily rely on plot establishment, individual-tree measurements, and visual interpretation. Although these approaches can provide accurate local species information, their application to large-area and repeated monitoring is constrained by high labor costs and limited updating efficiency. Remote sensing has therefore become an important approach for tree species identification and distribution mapping because of its spatial continuity, repeatability, and multiscale representation [
11,
12]. Early studies commonly used medium- and low-spatial-resolution data, such as MODIS and Landsat, to identify forest types or dominant species groups. Although these data are suitable for regional monitoring, mixed pixels limit crown-boundary delineation and fine-scale species separation in mixed stands [
13]. Multi-temporal imagery records spectral responses during green-up, leaf expansion, peak growth, and senescence, providing temporal cues for distinguishing species with similar appearances but different phenological rhythms [
14]. High-spatial-resolution data from WorldView, Pleiades, GF-2, and unmanned aerial vehicles can represent crown outlines, texture, and local spatial organization more clearly, extending tree species classification from the stand scale to patch- and individual-tree-scale applications [
15,
16,
17]. Nevertheless, a single optical image mainly captures the spectral and spatial properties of the crown surface and remains limited for species with similar crown forms or spectral responses, particularly under shadow occlusion. Consequently, LiDAR-derived canopy height and three-dimensional structure, continuous hyperspectral responses, and phenological differences from multi-temporal imagery have been increasingly incorporated to provide complementary observations of tree species [
18,
19,
20,
21,
22].
Early remote sensing-based tree species classification methods typically extracted spectral, textural, vegetation-index, geometric, and topographic features from imagery, combined them with LiDAR height, intensity, and canopy structural features, and then used classifiers such as support vector machines, random forests, and XGBoost to establish nonlinear relationships between features and species classes [
23,
24]. These methods can operate with relatively small sample sets and organize diverse statistical features effectively. However, their performance depends strongly on manual feature design, scale selection, and feature combination strategies [
25]. When data sources differ substantially in spatial resolution, numerical distribution, and semantic level, input-level stacking or static feature concatenation cannot adequately characterize the class- and location-dependent contributions of different observations. Such strategies may also introduce registration errors and observation noise into a common feature space [
26,
27]. These limitations highlight the need for representation learning methods that can preserve the characteristics of each data source while modeling their complementary relationships [
16,
28,
29].
Deep learning offers a technical pathway for hierarchical representation learning and end-to-end modeling of tree species characteristics. Convolutional neural networks can progressively extract edges, textures, crown morphology, and high-level semantic features from high-resolution imagery and have been applied to individual-tree species classification and vegetation semantic segmentation [
29,
30,
31]. For hyperspectral data with high dimensionality and strong inter-band correlation, models such as 3D-CNNs enhance class representation by jointly learning spatial–spectral features [
21,
32]. Multi-modal deep learning has further introduced branch-specific encoding, intermediate-layer interaction, attention modulation, and cross-modal relationship modeling to learn interactions among heterogeneous observations. However, the effectiveness of these approaches depends on how differences in spatial support, semantic level, and information reliability are reconciled [
18,
33,
34]. In fine-scale urban forest classification, sub-meter optical imagery and coarser hyperspectral and LiDAR grids describe the scene at different spatial supports, while the discriminative value of each modality varies across tree species and local environments [
35,
36]. Crown overlap, shadow occlusion, and the visual similarity of closely related broadleaf species further increase uncertainty near class boundaries. The resulting challenge is to learn spatially precise representations while aligning cross-modal features, adaptively weighting modality contributions according to spatial and semantic context, and limiting the influence of unreliable auxiliary information [
22,
37,
38].
The high cost of field surveys and crown-level annotation has also motivated increasing interest in label-efficient learning. Recent studies have explored several complementary strategies. Weakly supervised methods can exploit accessible low-resolution historical land-cover maps as inexact guidance; for example, Li et al. developed Paraformer with a pseudo-label-assisted training module to refine spatially mismatched coarse labels for high-resolution land-cover mapping [
39]. Semi-supervised remote sensing segmentation combines limited labeled samples with abundant unlabeled imagery, while decoupled learning and confidence-ranking weights reduce the influence of erroneous and class-imbalanced pseudo-labels [
40]. Difference-complementary learning and label reassignment further improve pseudo-label utilization in multi-modal optical-SAR segmentation [
41]. Under extremely limited supervision, Confucius tri-learning jointly trains two classifiers and one generator, enabling the classifiers to exchange reliable “good” quasi-labels and reject generated “bad” examples [
42]. Subsequent theoretical analysis connected mutual-guide learning with Confucius tri-learning and provided a theoretical justification for this mutual-guidance process under limited supervision [
43]. However, even when pseudo-labels provide additional supervisory information, reliable fine-grained tree species mapping requires the labels to remain consistent with crown boundaries and the associated spatial, phenological, structural, and spectral representations to be properly aligned. This requirement motivates the present study, which uses field surveys and visual interpretation to construct reference labels, takes high-resolution imagery as the spatial basis, and integrates multi-temporal NDVI, LiDAR canopy structural features, and hyperspectral–semantic priors through branch-specific encoding, dynamic gated fusion, and reliability-aware semantic modulation.
This study focuses on the northern section of Olympic Forest Park, Beijing. High-resolution imagery was used as the spatial information source and was integrated with NDVI phenological features, LiDAR canopy structural features, and hyperspectral–semantic priors. Field surveys and visual interpretation were combined to construct the tree species classification sample set. To address differences in information form and reliability across data sources, we developed a phenology–structure synergistic HybridFusion model: (1) the high-resolution branch uses branch encoding and dynamic gating to regulate the contributions of spatial, phenological, and structural features through pixel- and channel-level weights; and (2) the hyperspectral branch introduces a reliability-aware semantic prior that allows coarse-resolution spectral information to contribute to fine-grained species discrimination in a controlled manner. A set of comparative experiments evaluates how each information source and its interaction mechanisms affect classification accuracy, class confusion, and spatial mapping stability, providing a methodological reference for urban forest species mapping and coordinated use of multi-source remote sensing data.
2. Study Area and Multi-Source Remote Sensing Data
2.1. Study Area
This study was conducted in the North Area of Beijing Olympic Forest Park (
Figure 1), an approximately 3.0 km
2 urban planted forest ecosystem located on Beijing’s central axis. Shaped by long-term artificial tending and natural succession, the moderately undulating terrain supports a complex stand structure comprising coniferous, deciduous broadleaved, and mixed forests. The field inventory recorded 25 tree taxa, whose common names, scientific names, broad vegetation categories, and assignments to the model classes are summarized in
Table 1. Representative taxa include Chinese pine (
Pinus tabuliformis), deodar cedar (
Cedrus deodara), Chinese arborvitae (
Platycladus orientalis), Chinese willow (
Salix matsudana), weeping willow (
Salix babylonica), Siberian elm (
Ulmus pumila), Japanese pagoda tree (
Styphnolobium japonicum), green ash (
Fraxinus pennsylvanica), Shantung maple (
Acer truncatum), and white poplar (
Populus alba).
Table 1 complements
Figure 1 by clarifying both the taxonomic composition of the study area and the aggregation of the recorded taxa into the model classes. Spatially, the landscape is highly heterogeneous. Contiguous pure stands resulting from planned planting coexist with succession-driven mixed patches characterized by overlapping crowns and intermixed species. Furthermore, artificial features such as roads, water bodies, and buildings are interspersed throughout the forested landscape. This high degree of structural and landscape complexity provides an appropriate test site for evaluating fine-scale tree species identification using multi-source remote sensing data across pure stands, mixed stands, and forest–non-forest transition zones.
2.2. Multi-Source Remote Sensing Data
2.2.1. Multispectral Data
The primary dataset comprised multi-temporal Beijing-3 satellite imagery (2024–2025). Following radiometric calibration, atmospheric correction, orthorectification, and pan-sharpening, 0.5 m resolution multispectral images were generated to capture crown boundaries and textural details. Additionally, optical aerial imagery covering the North Area of Beijing Olympic Forest Park was acquired by an unmanned aerial vehicle (UAV) in 2025 at a spatial resolution of 0.1 m. This imagery was used solely as auxiliary data for visual interpretation, sample identification, and label verification.
2.2.2. Hyperspectral Data
Hyperspectral data were acquired in August 2024 using the GF-5B Advanced Hyperspectral Imager (AHSI). The imagery has a native spatial resolution of approximately 30 m and contains 330 contiguous bands covering 400–2500 nm, with a spectral resolution of 5–10 nm. Atmospheric correction was performed using the Fast Line-of-sight Atmospheric Analysis of Spectral Hypercubes (FLAASH), which is based on the MODTRAN radiative-transfer model, to obtain surface reflectance. The hyperspectral data were georeferenced to the common map reference and extracted over the same geographic bounds as each classification patch. For network processing, each hyperspectral window was represented as a 330 × 5 × 5 input patch. Because of its coarse native resolution, the hyperspectral imagery was not used for individual crown-boundary delineation; instead, it was used to generate a patch-level spectral–semantic prior for species classification.
2.2.3. LiDAR Data
Forest three-dimensional structural data were derived from airborne LiDAR point clouds acquired in 2022 using a Leica CityMapper-2 system mounted on a piloted P-750 aircraft. The flight altitude was 2100 m, the pulse repetition frequency was 2000 Hz, the average point density was approximately 16 points m−2, and up to five returns were recorded per pulse. The flight was conducted under clear-sky conditions. The point clouds were processed in LiDAR360 through strip adjustment, noise removal, classification, and visual inspection. Ground, vegetation, building, vehicle, and other points were identified. A digital elevation model was generated from ground points, and vegetation-point heights were normalized relative to the ground surface. The normalized vegetation points were aggregated into 2 m × 2 m grid cells, where maximum height, P95 height, mean height, height standard deviation, height coefficient of variation, canopy relief ratio, canopy closure, gap fraction, and leaf area index were calculated. The resulting structural rasters were transformed to the common coordinate reference system. For each label patch, the corresponding geographic extent was extracted, aligned with the 2 m classification grid, and resampled to the spatial dimensions required by the model.
The relationships among the multi-source remote sensing data, the derived features, and their roles in the tree species classification framework are summarized in
Figure 2. High-resolution imagery provides spatial and textural information, multi-temporal NDVI represents phenological variation, LiDAR provides canopy structural features, and hyperspectral imagery supplies spectral–semantic priors. These complementary data sources are subsequently integrated by the HybridFusion model for fine-grained tree species classification.
2.3. Ground Survey Data
Field surveys of tree species were conducted in the northern section of Olympic Forest Park, Beijing, in autumn 2025. According to the area’s species composition, stand types, and spatial distribution, sample locations covered typical settings, including extensive pure stands, mixed stands of Conifer and broadleaf species, forest edges, roadside green belts, forests surrounding water bodies, and landscape planting areas, thereby improving coverage across stand structures and background conditions. During field surveys, high-precision GPS equipment was used to record sample coordinates, tree species names, class attributes, and surrounding environmental information. Photographs of trunks, bark, leaves, branches, canopy morphology, and the local environment were also collected.
Areas with conspicuous crown interlacing or trees whose species could not be confirmed directly in the field were rechecked using field observations, 0.1 m aerial imagery, individual-tree detail photographs, and information from nearby samples to reduce the effects of crown mismatches and species misidentification on subsequent label production. The compiled survey comprised 1960 field tree species samples, covering the principal species in the study area, including Willow, Elm, Sophora, Poplar, and various conifers. Together with aerial image-assisted interpretation, these samples were used to prepare and validate the tree species labels.
2.4. Sample Production and Dataset Partitioning
This study conducted tree species sample interpretation and label production based on field survey points, field photographs, 0.1 m ultra-high-resolution aerial imagery, and 0.5 m high-resolution remote sensing imagery. After overlaying the field survey points onto the high-resolution imagery, polygon labels were created for areas with clearly identified tree species attributes and relatively clear canopy boundaries, using information such as tree crown boundaries, canopy texture, spatial location, and field photographs. Areas with severe crown overlap, obvious shadow occlusion, or unreliable tree species identification were excluded from the main training samples to reduce the influence of label uncertainty on model learning. Considering the complex tree species composition of the North Area of Beijing Olympic Forest Park and the limited sample size of some classes, this study did not establish a separate class for every original tree species. Instead, a classification system was constructed by considering sample size, ecological dominance, and remote sensing separability. Willow tree, elm tree, sophora tree, and poplar tree were treated as separate classes because they are major dominant broadleaf species in the study area. Chinese pine and Chinese arborvitae were merged into the “conifer” class. Species with fewer samples or scattered spatial distributions, such as ash, ginkgo, trident maple, paper mulberry, London plane, purple-leaf plum, apricot, and flowering crabapple, were merged into the “other tree species” class. Roads, water bodies, grassland, bare land, buildings, shadows, and similar areas were assigned to the “non-tree” class. This produced a seven-class classification system: Non-tree, Willow tree, Conifer, Elm tree, Sophora tree, Poplar tree, and Other tree species.
Label production was based on manually delineated tree crown polygons. To reconcile the spatial-resolution differences among the multi-source datasets, the original labels and the 0.5 m high-resolution imagery were transformed to a common 2 m classification grid. Subsequent model training, accuracy assessment, and full-area tree species mapping were conducted at this 2 m spatial resolution. Each label patch contains 20 × 20 cells and covers a geographic extent of 40 m × 40 m. The high-resolution imagery, NDVI features, and LiDAR structural features were extracted over this extent as 80 × 80 arrays, whereas the hyperspectral data were extracted over the same geographic extent and represented as a 330 × 5 × 5 coarse input patch. Thus, the hyperspectral input provides patch-level spectral–semantic context rather than one-to-one information for each 2 m label cell. During preprocessing, all source windows were generated through CRS-aware reprojection to the label-grid bounds. Patches with no source-label overlap were rejected. The study area contained 1960 field tree species samples, 889 tree species label polygons, and 26,381 valid tree species label pixels. A training set and validation set were then constructed, including 676 training samples and 64 validation samples, for a total of 740 patch samples used for model training, parameter selection, and classification accuracy assessment. The classification system and sample quantities are shown in
Table 2.
4. Experimental Design and Results
4.1. Experimental Setup and Ablation Design
The model and all additional baseline models were implemented in PyTorch version 2.2.1 and trained on the same NVIDIA RTX 4070 GPU using 20 × 20 pixel label patches. The AdamW optimizer was employed with an initial learning rate of 0.001 and a weight decay of 0.0001. The learning rate was adjusted using cosine annealing over 100 epochs, with a batch size of 128 [
52]. The total loss combined class-weighted cross-entropy and an auxiliary boundary loss, for which the loss weight and dilation radius were set to 0.2 and 1, respectively. Overall accuracy (OA), tree species OA (Tree OA), the Kappa coefficient, mean Intersection over Union (mIoU), and Macro-F1 were used as evaluation metrics, and the checkpoint with the highest validation Tree OA was selected. As summarized in
Table 3,
Table 4 and
Table 5, three groups of experiments were designed to systematically evaluate data source contributions and fusion strategies:
(1) Data source contributions (Exps. 1–6): Evaluated single-source versus multi-source classification. Exp. 1 served as the baseline (August high-resolution image only). Exps. 2–5 progressively introduced NDVI, LiDAR, and HSI priors to assess their independent and joint contributions, culminating in Exp. 6, which represented the complete HybridFusion model with all modalities.
(2) High-resolution fusion strategies (Exps. 7, 8, and 4): Compared input-level stacking, feature-level fusion, and dynamic gated fusion (the no-HSI baseline), respectively, to determine the optimal joint representation mechanism for spatial, phenological, and structural features.
(3) Hyperspectral prior ablations (Exps. 4, 9–12, and 6): Using Exp. 4 as a baseline, these experiments compared spatial bias, FiLM channel-modulation, class prior, ungated-reliability, and reliability-aware configurations to validate the effectiveness of hyperspectral–semantic guidance and reliability control.
To provide an architecture-level comparison under identical information availability, four external baselines were additionally implemented: FullModal-U-Net, FullModal-DeepLabV3+, FullModal-FPN, and MultiModal-Concat. Each model received four-channel high-resolution imagery, two-channel multi-temporal NDVI, nine-channel LiDAR structural features, and a 330-band HSI patch. For FullModal-U-Net, FullModal-DeepLabV3+, and FullModal-FPN, an identical lightweight HSI adapter projected the hyperspectral input from 330 to 32 and subsequently to 8 channels using two 1 × 1 convolutional layers. The resulting features were bilinearly upsampled and concatenated with the high-resolution imagery, NDVI, and LiDAR features before being processed by the corresponding segmentation backbone. MultiModal-Concat independently encoded the four modalities and concatenated their features at a common spatial resolution before classification. Except for architecture-specific fusion and regularization components, all models used the same data partition, preprocessing procedure, optimizer, learning-rate schedule, number of epochs, batch size, loss functions, checkpoint-selection criterion, and evaluation metrics.
4.2. Results
4.2.1. Separability of Multi-Source Remote Sensing Features
To examine how multi-source remote sensing information affects tree species discrimination, we evaluated feature separability at two levels: original input features and deep fused features learned by the model. Using valid pixels from the training and validation sets, four original feature combinations were constructed: high-resolution imagery alone, high-resolution imagery + NDVI, high-resolution imagery + LiDAR, and high-resolution imagery + NDVI + LiDAR. After standardization and preliminary dimensionality reduction with PCA, each feature set was mapped to two-dimensional space using t-SNE to represent class clustering and overlap. The silhouette coefficient, Davies–Bouldin (DB) index, and Calinski–Harabasz (CH) index were also used for quantitative evaluation. Higher silhouette and CH values and a lower DB value generally indicate stronger class separability. The t-SNE visualizations of the original feature space are shown in
Figure 8, and the quantitative results are listed in
Table 6.
Figure 8 and
Table 6 show that the use of high-resolution imagery alone leads to substantial overlap among tree species in feature space. The boundaries of broadleaf classes, including Willow, Elm, Sophora, and Other trees, are particularly indistinct. The silhouette coefficient is −0.1260 and the DB index is 5.4952, indicating that a single-date high-resolution image can represent crown texture and local spatial morphology but has limited ability to separate fine-grained tree species classes. After multi-temporal NDVI is introduced, the organization of the feature space improves markedly: the silhouette coefficient increases to −0.0492, the DB index decreases to 3.2514, and the CH index increases from 186.96 to 248.30. These changes indicate that phenological information enhances interspecies feature differences. LiDAR structural features complement optical imagery primarily through canopy height and vertical structure. Although the high-resolution imagery + LiDAR combination does not achieve the best global clustering indices in the original feature space, the class-mean heat map of LiDAR structural features (
Figure 9) shows that Poplar differs substantially in maximum height, mean height, canopy closure, and leaf area index. Thus, structural information helps distinguish species with pronounced crown morphology. When NDVI and LiDAR are jointly introduced, the silhouette coefficient further increases to −0.0464, demonstrating complementarity between phenological and structural features in representing interspecies differences.
To further assess feature organization after model learning, we extracted deep encoding features from the best model of each experiment and visualized them using t-SNE (
Figure 10 and
Table 7). Compared with the original features, deep features learned under supervision are more closely aligned with the classification objective and no longer depend on a single spectral, phenological, or structural difference. Adding NDVI clarifies the class distribution relative to the baseline. Feature-level fusion and dynamic fusion can adjust interclass distances to some extent. After the hyperspectral prior is introduced, the global clustering indices do not improve monotonically, indicating that its principal role is semantic correction for complex samples and locally confused regions. Overall, multi-source remote sensing features enhance tree species representation through phenological variation, canopy structure, and spectral semantics, establishing a feature basis for improved classification accuracy.
4.2.2. Accuracy of the Full-Modality Model
To further evaluate class-level performance of the full-modality HybridFusion model, the best result from Experiment 6 is summarized in
Table 8. Overall, the full-modality model achieved stable classification performance for the Non-tree class and the major tree species. The Non-tree class achieved an F1 score of 98.58% and an IoU of 97.21%, indicating effective separation of crown areas from roads, water bodies, grassland, and buildings. Among tree species, Poplar achieved the best result, with an F1 score of 95.31% and PA of 97.60%. Its pronounced canopy structure and spatial morphology provide high separability after multi-source fusion. F1 scores for Willow, Conifer, and Other trees were 86.27%, 84.90%, and 85.35%, respectively, indicating that phenological, structural, and spectral–semantic information effectively supports recognition of the principal tree species.
Differences in UA and PA across species reveal distinct error patterns. Willow attained a UA of 92.34%, exceeding its PA of 80.96%, indicating high reliability of Willow predictions but some omission errors. Conifer achieved balanced performance, with UA and PA of 84.66% and 85.14%, respectively. Elm had the lowest F1 score among tree species (78.58%) and a UA of only 73.89%, indicating residual confusion with other broadleaf species in crown texture and phenological response. Sophora achieved an F1 score of 81.05% and a PA of 76.02%, suggesting some omission errors that may be related to crown interlacing in mixed stands and within-class sample variation. Other trees achieved a PA of 91.93% but a UA of 79.65%. This class combines multiple species with small sample sizes and scattered distributions, resulting in strong within-class heterogeneity and a tendency to absorb predictions from similar broadleaf species. Overall, HybridFusion performs well for major tree species, but Elm, Sophora, and Other trees remain priority classes for further improvement in fine-grained classification.
4.2.3. Comparison with Representative Full-Modality Baselines
Under the unified full-modality setting, FullModal-FPN was the strongest of the three conventional semantic segmentation backbones, achieving an OA of 89.64%, a Tree OA of 82.35%, a Kappa coefficient of 0.8593, an mIoU of 74.19%, and a Macro-F1 of 84.73%. Compared with FullModal-U-Net, FullModal-FPN improved Tree OA, mIoU, and Macro-F1 by 4.10, 4.44, and 3.54 percentage points, respectively (
Table 9). Compared with FullModal-DeepLabV3+, the corresponding improvements were 5.76, 5.32, and 4.33 percentage points. MultiModal-Concat achieved a Tree OA of 80.39%, indicating that direct feature concatenation can exploit complete multi-modal inputs but provides limited adaptive intermodal interaction.
HybridFusion achieved the highest values for all evaluation metrics, with an OA of 90.44%, a Tree OA of 85.06%, a Kappa coefficient of 0.8707, an mIoU of 77.88%, and a Macro-F1 of 87.15%. Relative to FullModal-FPN, HybridFusion improved Tree OA, mIoU, and Macro-F1 by 2.71, 3.69, and 2.42 percentage points, respectively. Because all models received the complete set of remote sensing sources, these results indicate that the improvement of HybridFusion is not attributable solely to the inclusion of HSI data. Instead, the results support the effectiveness of dynamic multi-modal fusion and reliability-aware hyperspectral–semantic prior modeling.
4.2.4. Computational Complexity and Inference Efficiency
The computational complexity and inference efficiency of the five full-modality models were further evaluated. Parameter count and FLOPs were calculated for the complete inference graph, including the HSI-processing components. All inference measurements were performed on the same NVIDIA RTX 4070 GPU using FP32 inference, a batch size of 1, five warm-up iterations, and 30 timed iterations.
MultiModal-Concat had the smallest parameter count and lowest latency, with 0.507 M parameters and a mean latency of 1.349 ms (
Table 10). FullModal-DeepLabV3+ had the lowest theoretical computational cost among the three conventional segmentation backbones, requiring 0.854 GFLOPs. FullModal-FPN provided the most favorable accuracy–efficiency balance, requiring 1.258 M parameters and 1.262 GFLOPs while achieving a mean latency of 1.468 ms and a throughput of 681.23 patches/s. HybridFusion required 40.061 M parameters and 3.915 GFLOPs, with a mean latency of 14.800 ms. Compared with FullModal-FPN, HybridFusion used approximately 31.8 times as many parameters, 3.1 times the FLOPs, and 10.1 times the measured latency while improving Tree OA by 2.71 percentage points and mIoU by 3.69 percentage points. HybridFusion therefore trades greater computational cost for improved fine-grained classification accuracy.
4.3. Spatial Distribution Patterns of Tree Species
To transfer the tree species classification model trained at the sample scale to the full study area, we performed study area-wide mapping using overlapping sliding-window inference. High-resolution imagery served as the common spatial reference, and the corresponding high-resolution imagery, multi-temporal NDVI, LiDAR structural features, and HSI semantic priors were read for each window. Each experiment activated or masked input branches according to its modality configuration, ensuring consistency in spatial extent, class scheme, and inference workflow across the 12 classification results. The final outputs comprised seven classes: Non-tree, Willow, Conifer, Elm, Sophora, Poplar, and Other trees. To reduce the effect of unstable predictions near sliding-window edges on the full-image classification result, Gaussian-weighted fusion was applied across overlapping windows. Let the output-window side length be L, the relative pixel coordinate within a window be (i, j), and the window-center coordinate be
. The two-dimensional Gaussian weight is defined as:
To combine predictions from different windows, the Gaussian weights are normalized as:
For any pixel x in the full image that is covered by multiple sliding windows, let K(x) be the set of windows covering x and let p
k,
m(x) be the probability for class m predicted by the kth window. The fused class probability at that pixel is then:
The final class is determined by the maximum fused probability:
This strategy improves prediction continuity in overlapping regions, yielding more stable spatial representation at stand boundaries and within patches in the study area-wide classification map (
Figure 11).
Using the best-performing HybridFusion model in the integrated evaluation, we further analyzed tree species spatial distribution in the northern section of Olympic Forest Park (
Figure 12). Based on the 2 m classification grid, the effective study area is approximately 310.85 ha, including 175.12 ha of tree-covered area (56.34%) and 135.73 ha of Non-tree area (43.66%). The percentages of tree-covered and Non-tree areas were calculated relative to the total effective study area, whereas the species-level percentages were calculated relative to the total tree-covered area. Other trees occupy the largest tree-covered area, covering 67.69 ha or 38.65% of the total tree-covered area. Elm, Willow, and Conifer cover 27.77, 25.81, and 23.56 ha, accounting for 15.86%, 14.74%, and 13.45%, respectively. Sophora and Poplar cover 17.74 and 12.55 ha, accounting for 10.13% and 7.17%, respectively.
Spatially, the tree species distribution reflects characteristic patterns of urban mixed forests and landscape planting. The “Other trees” class dominates the areal coverage, highlighting the prevalence of diverse ornamental species and complex mixed patches beyond the primary targets. This inherent compositional complexity makes it a highly heterogeneous target for remote sensing classification. Willows predominantly form belts along water bodies, aligning with typical riparian planting strategies. Conifers cluster to form localized, contiguous evergreen stands. In contrast, poplars cover a minor area and are largely confined to distinct linear patterns along roadsides and forest edges. Elms and Sophoras are spatially dispersed, frequently occurring within mixed interior stands or roadside greenspaces, thereby epitomizing the highly intermixed broadleaf structures typical of urban forests.
4.4. Comparison of Classification Results in Representative Areas
To visually evaluate the impact of multi-source data fusion on fine-scale classification, three representative local scenes were selected (
Figure 13): a contiguous Conifer patch (Scene A), a linear roadside Poplar belt (Scene B), and a complex broadleaf mixed forest (Scene C).
Scene A is a Conifer forest patch. When only high-resolution imagery is used, Conifer predictions are fragmented and some locations are misclassified as Elm, Willow, or Other trees. This result indicates that a single-date high-resolution image can represent crown texture and color differences but cannot reliably separate the Conifer class from neighboring broadleaf species. After NDVI is introduced, predicted Conifer areas correspond more closely to Conifer canopy locations in the imagery; mixed predictions of broadleaf species within patches decrease, and both spatial continuity and class purity improve. After LiDAR structural features and the HSI semantic prior are incorporated, the main Conifer area remains spatially continuous, indicating that phenological, canopy structural, and spectral–semantic information jointly improves the reliability of evergreen forest patch recognition.
Scene B is a roadside Poplar stand and tall-tree belt, where crowns are distributed linearly along roads and forest edges. With high-resolution imagery alone or combined with NDVI, the target area still contains mixed predictions of Elm, Sophora, and Other trees, and the Poplar map lacks continuity. After LiDAR structural features are introduced, Poplar’s roadside linear distribution becomes clearer and more consistent with the actual location of the tall-tree belt. This finding indicates that structural information such as tree height, canopy relief, and canopy closure provides effective constraints for recognizing species with distinctive canopy structure. The full-modality HybridFusion model preserves the linear spatial pattern while reducing locally fragmented misclassification, producing a more complete classification of roadside tree belts.
Scene C is a complex broadleaf mixed forest containing Willow, Elm, Sophora, and Other trees. Crown overlap, forest-edge shadows, and mosaic species distributions are pronounced. With high-resolution imagery alone, predictions contain numerous scattered patches and some class boundaries correspond inconsistently to actual crown structure. Adding NDVI and LiDAR improves local spatial continuity, but confusion persists among Elm, Sophora, and Other trees. By jointly leveraging phenological variation, canopy structure, and hyperspectral–semantic information, the full-modality model improves class continuity and reduces local errors across spatial contexts; nevertheless, fine-grained discrimination within complex broadleaf mixed forests remains highly uncertain.
6. Conclusions
This study developed an advanced multi-source data fusion framework implemented in the northern section of Beijing’s Olympic Forest Park. By systematically integrating 0.5 m high-resolution imagery, multi-temporal NDVI, LiDAR structural metrics, and hyperspectral–semantic priors on a 2 m grid, comprehensive experiments, class-level assessments, and full-area mapping yielded the following key conclusions:
(1) Superior performance of multi-source data fusion. The full-modality integration achieved exceptional classification accuracy, with an OA of 90.44%, Tree OA of 85.06%, Kappa of 0.8707, mIoU of 77.88%, and Macro-F1 of 87.15%, substantially outperforming the single-source baseline. Poplar attained the highest F1 score (95.31%), followed by Willow (86.27%), Other trees (85.35%), and Conifer (84.90%).
(2) Phenological, canopy structural, and spectral–semantic data offer complementary contributions to tree species classification. Using high-resolution imagery alone, the overall accuracy (OA) was 67.29%. The addition of NDVI increased the OA to 76.06%. Incorporating LiDAR data raised the OA to 76.63%, indicating that metrics such as canopy height, relief, and closure effectively address the vertical spatial limitations of two-dimensional imagery. Furthermore, the joint application of NDVI and LiDAR boosted the OA to 81.90%. Finally, integrating a hyperspectral–semantic prior achieved a full-modality OA of 85.06%, highlighting that hyperspectral data serves as a powerful deep semantic complement to spatial texture, phenological, and structural features.
(3) Dynamic gated fusion and a reliability-aware hyperspectral prior improve multi-source fusion. High-resolution-side experiments showed that input-level stacking cannot adequately address differences in scale and noise among modalities, whereas dynamic gated fusion achieved the highest OA and Tree OA. LiDAR spatial modulation and joint pixel-channel gating therefore support adaptive allocation of modal contributions. Hyperspectral prior ablations showed that the reliability-aware semantic prior outperformed spatial bias, FiLM modulation alone, a class prior, and fixed reliability.
(4) Reliable large-scale mapping of urban forest structure. Study area-wide mapping revealed an effective mapping area of ~310.85 ha, with tree-covered zones accounting for ~175.12 ha (56.34%). Visual comparisons across representative scenes further corroborated that NDVI stabilizes conifer identification, LiDAR reinforces the spatial continuity of roadside tall-tree belts, and hyperspectral priors suppress local classification errors in complex mixed forests.
Several limitations should be acknowledged. First, the LiDAR observations were acquired in 2022, whereas the multispectral and hyperspectral data were collected during 2024–2025. Although no large-scale tree removal occurred in the protected study area, crown growth, pruning, branch loss, disease, storm damage, and routine maintenance may have altered local canopy structure, particularly in mixed stands and near crown boundaries. Second, the “Other trees” category combines multiple taxa with heterogeneous spectral, structural, and phenological characteristics, while the available samples may not represent each constituent species evenly. This limits class-specific interpretation and transferability to areas with different species compositions. Third, the full-modality HybridFusion model improves classification accuracy at the cost of increased computational complexity, memory usage, and inference time compared with simpler baseline models, which may constrain direct edge deployment. Future work should emphasize multi-site and multi-year validation, temporal calibration, and model-compression techniques.