Next Article in Journal
Sea Surface Current Vector Reconstruction from Multitemporal Sentinel-1 Doppler Observations: A Trajectory-Crossing Approach with Spatial Registration
Previous Article in Journal
Low-Cost Geological Reconnaissance for Artisanal and Small-Scale Mining: RGB–HSV Analysis of Rendered Google Earth Imagery in Arid Copper-Prospective Terrains of Chile and Balochistan
Previous Article in Special Issue
MambaHSINet: A Dual-Branch Bidirectional State Space Network for Hyperspectral Tree Species Classification
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Unsupervised Scale-Conditioned Hyperspectral and Multispectral Image Fusion via a Frequency–Spatial Dual-Domain Network

1
College of Geography and Environment, Liaocheng University, Liaocheng 252000, China
2
School of Computer Science and Technology (National Exemplary Software School), Chongqing University of Posts and Telecommunications, Chongqing 400065, China
3
Information Science and Technology College, Dalian Maritime University, Dalian 116026, China
4
Key Laboratory of Computational Optical Imaging Technology, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(18), 3187; https://doi.org/10.3390/rs18183187
Submission received: 7 August 2026 / Revised: 12 September 2026 / Accepted: 14 September 2026 / Published: 16 September 2026

Highlights

What are the main findings?
  • A novel unsupervised framework, SCDF-Net, was developed for HSI–MSI fusion under both integer and fractional resolution ratios, operating directly on the native input grids without any ground-truth HR-HSI labels.
  • The two-stage design freezes self-supervised PSF/SRF estimates as physical priors and drives a frequency–spatial dual-domain network with a continuous scale embedding. SCDF-Net attains the lowest SAM on all four datasets and improves PSNR/RMSE in most baseline comparisons at ×1.5, ×2.0, ×2.4, and ×4.0, while the multi-scale self-evaluation further covers ×5.0 and ×6.0 to assess robustness under severe spatial degradation.
What are the implications of the main findings?
  • Treating the resolution ratio as an explicit continuous condition establishes a general principle for scale-flexible fusion, removing the interpolation error incurred when fractional cases are forced onto a nearby integer ratio.
  • Coupling learned degradation priors with scale-conditioned dual-domain detail modeling under observation-domain supervision provides a robust, label-free foundation for real cross-sensor Earth-observation tasks where HR-HSI references are unavailable.

Abstract

Hyperspectral–multispectral image fusion reconstructs high-spatial-resolution hyperspectral images (HR-HSIs) by combining low-resolution hyperspectral images (LR-HSIs) with high-resolution multispectral images (HR-MSIs). Many methods only support integer resolution ratios; when the LR-HSI/HR-MSI ratio is fractional, inputs are often resampled to a nearby integer ratio, altering observations and introducing interpolation error. We propose SCDF-Net, an unsupervised Scale-Conditioned Dual-domain Fusion Network that treats the spatial ratio as an explicit conditioning variable and fuses directly on native grids without HR-HSI labels. A degradation network first estimates the point spread function and spectral response function in a self-supervised manner; the learned operators are then frozen as physical priors. Conditioned on a continuous scale embedding, SCDF-Net integrates HSI spectral features and MSI spatial features through coupled frequency- and spatial-domain branches, trained with dual observation-domain consistency and spectral/frequency regularizations. On four benchmarks, baseline comparisons at scale factors ×1.5, ×2.0, ×2.4, and ×4.0 show that SCDF-Net obtains the lowest SAM on all datasets and improves PSNR/RMSE in most reported settings, with clear advantages under fractional scale factors. In addition, the multi-scale self-evaluation is extended to ×5.0 and ×6.0 to assess robustness under more severe spatial degradation.

1. Introduction

Hyperspectral remote sensing plays an important role in Earth observation because hyperspectral images (HSIs) contain rich and continuous spectral information. By recording dozens to hundreds of narrow spectral bands, HSIs can capture detailed material-specific spectral signatures and have been widely used in land-cover classification, vegetation monitoring, mineral mapping, environmental assessment, and disaster analysis [1]. However, due to sensor design, signal-to-noise ratio requirements, and payload constraints, HSIs usually have limited spatial resolution, which makes it difficult to distinguish fine spatial structures and mixed land-cover patterns in complex remote sensing scenes [2,3,4]. In contrast, multispectral images (MSIs), such as those acquired by Landsat OLI and Sentinel-2 MSIs, generally provide higher spatial resolution but contain fewer and broader spectral bands [3,4]. Therefore, HSI–MSI fusion, also known as spatio-spectral fusion, aims to integrate the rich spectral information of low-resolution HSIs (LR-HSIs) with the fine spatial details of high-resolution MSIs (HR-MSIs) to reconstruct high-spatial-resolution HSIs (HR-HSIs) [2,3,4,5,6,7].
Traditional HSI–MSI fusion methods mainly rely on explicit observation models and handcrafted priors. Representative methods include component substitution, multi-resolution analysis, spectral unmixing, matrix factorization, sparse representation, and tensor-based models [2,3,4,5,6,7,8,9,10,11,12,13]. These methods are usually physically interpretable and can achieve reasonable reconstruction performance under well-defined imaging assumptions. However, they often depend on accurate prior knowledge of degradation operators, such as the point spread function (PSF) and spectral response function (SRF) [3,5,7,11]. In practical remote sensing scenarios, these degradation operators are usually unknown or difficult to estimate accurately, which limits the robustness and generalization ability of traditional methods.
With the development of deep learning, supervised HSI–MSI fusion methods have shown strong capability in learning nonlinear spatial–spectral mappings from data. CNN-based methods usually extract spectral features from LR-HSIs and spatial features from HR-MSIs [14,15,16], while model-guided networks further embed observation models or image priors into deep architectures to improve interpretability [17,18]. More recently, Transformer-based methods have been introduced to enhance long-range spatial–spectral interaction [19,20,21]. Although supervised deep learning methods achieve promising performance, they generally require paired training samples or simulated HR-HSI references [14,15,16,17,18,19,20,21,22,23,24,25]. Such reference data are difficult or expensive to obtain in real remote sensing applications. As a result, the dependence on supervised signals limits their practical applicability and cross-sensor generalization.
To reduce the dependence on ground-truth HR-HSIs, unsupervised and self-supervised HSI–MSI fusion methods have attracted increasing attention. These methods usually train fusion networks through observation-domain reconstruction, degradation consistency, cycle consistency, latent-space constraints, or modality alignment, using only the available LR-HSI/HR-MSI pair as supervision [26,27,28,29,30,31,32,33]. Compared with supervised fusion methods, unsupervised approaches are more practical for real remote sensing scenarios where HR-HSI labels are unavailable; meanwhile, recent graph-based HSI clustering studies have also demonstrated the potential of learning robust spatial–spectral representations from unlabeled hyperspectral data [34,35,36]. However, many existing unsupervised fusion methods are still designed under fixed degradation settings or integer spatial resolution ratios and cannot directly handle non-integer (fractional) scale factors without first resampling the inputs to an integer ratio [26,27,28,29,32,33]. For example, fusing a 30 m hyperspectral image (e.g., EnMAP/PRISMA) with the 20 m bands of Sentinel-2 yields a spatial ratio of 1.5. Integer-only methods cannot handle this ratio directly and must first resample the inputs to an integer ratio (e.g., 1× or 2×). This resampling changes the original LR-HSI/HR-MSI observations and adds interpolation error before fusion even begins, which motivates a method that operates directly under fractional ratios on the native input grids. Beyond HSI–MSI fusion, unsupervised learning has also been explored in related remote sensing tasks. For instance, tensor-based learning and graph-fusion strategies have shown promise in handling complex multi-modal data with missing information, such as incomplete multi-view clustering [37].
Despite recent progress in continuous- and implicit-representation-based fusion [30,31], scale-conditioned unsupervised HSI–MSI fusion under fractional resolution ratios still faces several challenges. First, integer-ratio methods assume that spatial downsampling produces an exact integer-size reduction. Under a fractional ratio, such a degradation operator cannot generate an LR grid of the correct size, so aligned observation-domain (self-supervised) losses cannot be constructed unless the inputs are first resampled to an integer ratio. This resampling step modifies the observations that the losses are computed on, leading to inaccurate supervision. Second, existing fusion networks often use fixed-scale architectures or fixed detail injection strategies, which may not adapt well to different resolution gaps. Third, HR-MSIs contain abundant high-frequency spatial details, but directly injecting these details may cause spectral distortion, especially in heterogeneous regions with complex land-cover mixtures [2,3,4,10,22,31]. Therefore, an effective framework should simultaneously support non-integer scale degradation alignment, scale-conditioned detail injection, and spectral fidelity preservation.
To address these challenges, we propose SCDF-Net, an unsupervised Scale-Conditioned Dual-domain Fusion Network that couples scale-controlled degradation alignment with frequency–spatial detail injection, supporting both integer and non-integer spatial resolution ratios. Here, dual-domain refers to the frequency-domain and spatial-domain feature pathways, which is distinct from the observation-domain consistency used for self-supervised training. The proposed method treats the spatial scale factor as a continuous conditioning variable. In the degradation estimation stage, the deterministic scale factor is used to define scale-controlled target-grid degradation, so that degraded observations are spatially aligned and comparable. In the fusion stage, the same scale factor is mapped into a learnable scale embedding vector to modulate frequency-domain and spatial-domain detail extraction. In this way, SCDF-Net can be applied to a given spatial resolution ratio—integer or fractional—directly on the native input grids, without resampling the inputs to an integer ratio and without requiring ground-truth HR-HSIs.
The main contributions of this work are summarized as follows:
(1)
We propose an unsupervised scale-conditioned HSI–MSI fusion framework that supports both integer and non-integer spatial resolution ratios. By introducing scale-controlled target-grid degradation and a scale embedding, the proposed method allows observation-domain constraints to be constructed on aligned grids under fractional scale factors.
(2)
We design a two-stage degradation-aware training strategy. The PSF and SRF are first estimated from the observed LR-HSI/HR-MSI pair in a self-supervised manner. The learned degradation operators are then frozen as physical priors to construct observation-domain constraints for unsupervised fusion training.
(3)
We develop a scale-conditioned frequency–spatial dual-domain fusion network. The frequency-domain branch enhances high-frequency spatial details from HR-MSIs, while the spatial-domain branch provides complementary structural information. Both branches are modulated by the scale embedding, enabling adaptive detail injection while preserving spectral fidelity.

2. Related Work

2.1. Traditional HSI–MSI Fusion Methods

Traditional HSI–MSI fusion methods usually reconstruct HR-HSIs by combining explicit observation models with handcrafted priors. According to existing surveys, these methods can be broadly grouped into component-substitution/detail-injection methods, spectral-unmixing or matrix-factorization methods, and optimization-based reconstruction methods [2,3,4].
In component-substitution and detail-injection methods, spatial information from high-resolution observations is injected into the upsampled hyperspectral image. Aiazzi et al. proposed the Gram–Schmidt Adaptive (GSA) method, which improves component substitution by using multivariate regression between multispectral and panchromatic-like components [8]. This type of method is efficient and easy to implement, but it may introduce spectral distortion when the assumed relationship between spatial and spectral components is inaccurate.
Spectral-unmixing and matrix-factorization methods describe the target HR-HSI using low-dimensional spectral and spatial factors. Yokoya et al. proposed coupled nonnegative matrix factorization (CNMF), which jointly factorizes LR-HSIs and HR-MSIs into endmembers and abundance maps [5]. Lanaras et al. further developed coupled spectral unmixing to estimate the shared spectral and spatial factors from the two modalities [9]. These methods provide clear physical interpretability, but their performance depends strongly on endmember estimation, linear mixing assumptions, and degradation-model accuracy.
Optimization-based methods formulate HSI–MSI fusion as an inverse problem with explicit priors and regularization. Simões et al. proposed HySure, which combines subspace representation, spatial-spectral degradation modeling, and total variation regularization [7]. Wei et al. introduced a sparse-representation-based fusion method that exploits spectral redundancy and spatial correlation through learned dictionaries [10]. These methods are more flexible than simple detail-injection approaches, but they often require reliable PSF/SRF information and careful parameter tuning.
Traditional methods are interpretable and do not require large-scale training data. However, they usually rely on fixed degradation assumptions or accurately known observation models, which limits their robustness under unknown degradation or non-integer spatial resolution ratios.

2.2. Supervised Deep Learning-Based HSI–MSI Fusion Methods

Supervised deep learning methods learn nonlinear spatial-spectral mappings from paired or simulated training data. Compared with traditional optimization-based methods, deep networks usually achieve stronger representation ability and better reconstruction performance when reliable HR-HSI references are available [14,15,16,17,18,19,20,21,22,23,24,25].
CNN-based methods first became a common solution for extracting complementary modality features. Yang et al. proposed a two-branch CNN that extracts spectral features from LR-HSIs and spatial features from HR-MSIs before fusion [15]. Palsson et al. introduced a 3D CNN-based fusion method to exploit local spectral–spatial correlations [14]. These CNN-based methods improve reconstruction quality, but they usually require paired HR-HSI labels and are trained under fixed degradation or scale settings.
Model-guided deep networks further embed physical observation models into trainable architectures. Xie et al. proposed MS/HS Fusion Net, which unfolds an optimization process into a deep network for HSI–MSI fusion [17]. They further developed MHF-Net, an interpretable proximal-gradient-unfolding network based on the linear observation model [18]. These methods improve interpretability and observation consistency, but they still depend on supervised training samples and predefined degradation assumptions.
Transformer-based methods have recently been introduced to model long-range spatial-spectral dependencies and cross-modal interactions. Jia et al. proposed MSST-Net, which uses multiscale spatial-spectral Transformer modeling to enhance feature representation [19]. Ma et al. proposed a reciprocal Transformer to strengthen bidirectional interaction between LR-HSI and HR-MSI features [20]. Although Transformer-based methods increase representation capacity, they often require large training sets, high computation, and fixed training configurations.
Supervised deep learning methods have strong reconstruction ability, but their dependence on paired HR-HSI references limits their applicability in real remote sensing scenarios. Their generalization ability is also restricted when the testing scale or degradation differs from the training setting.

2.3. Unsupervised Deep Learning-Based HSI–MSI Fusion Methods

Unsupervised and self-supervised HSI–MSI fusion methods reduce the dependence on ground-truth HR-HSIs by constructing supervision from the observed LR-HSI/HR-MSI pair. These methods usually degrade the reconstructed HR-HSI back to the observation domains and optimize the network through observation-domain consistency [26,27,28,29,30,31,32,33].
Degradation-aware unsupervised methods focus on learning or using spatial and spectral degradation models during training. Li et al. proposed UDALN, which estimates degradation processes with dedicated modules in a multi-stage unsupervised framework [26]. Li et al. further proposed UMC2FF, a model-guided coarse-to-fine fusion network that progressively refines the fused image through physical observation constraints [28]. These methods improve practical usability, but they still mainly work under predefined degradation and scale settings.
Other unsupervised methods address practical cross-sensor issues such as misregistration and modality alignment. Qu et al. proposed u2MDN, which introduces Mutual Dirichlet-Net to improve robustness under unregistered HSI–MSI inputs [27]. This method is useful for real-scene fusion, but its main focus is modality alignment rather than reconstruction under varying (including non-integer) scale ratios.
Recent continuous- and implicit-representation methods improve scale flexibility. Wang et al. proposed CLoRF, which uses continuous low-rank factorization to reconstruct fused images at user-specified resolutions [30]. Liang et al. proposed FeINFN, a Fourier-enhanced implicit neural fusion network that improves high-frequency reconstruction through spatial-frequency implicit modeling [31]. These methods show the importance of continuous scale modeling and frequency-domain information, but the joint modeling of scale-controlled degradation alignment, learned PSF/SRF priors, and scale-conditioned frequency–spatial detail injection remains insufficiently explored.
Overall, degradation-aware and alignment-oriented unsupervised methods mainly improve training practicality and robustness, whereas continuous- and implicit-representation-based methods further enhance resolution flexibility. However, degradation alignment under non-integer scales, adaptive detail injection, and spectral fidelity preservation remain insufficiently coupled. This motivates the proposed SCDF-Net framework, which uses learned degradation priors, continuous scale embedding, and scale-conditioned frequency- and spatial-domain fusion for unsupervised HSI–MSI reconstruction under integer and fractional scale factors.

3. Method

3.1. Overall Training Framework

Let X LR R B × h × w denote the observed low-spatial-resolution hyperspectral image (LR-HSI), and let Y HR R C × H × W denote the observed high-spatial-resolution multispectral image (HR-MSI), where B and C represent the numbers of hyperspectral and multispectral bands, respectively, and C B . The objective of HSI–MSI fusion is to reconstruct a high-spatial-resolution hyperspectral image X ^ HR R B × H × W [2,3,4].
The spatial scale factor is determined by the input spatial sizes:
s   =   [ s h , s w ]   =   [ H / h ,   W / w ]
This scale factor is deterministic and is directly computed from the observed image sizes. Specifically, s describes the vertical and horizontal resolution ratios between the HR-MSI and LR-HSI, and therefore provides an explicit scale condition for fusion under integer and fractional target ratios. In Stage 1, s is used to define the scale-controlled spatial degradation and target-grid resampling process. In Stage 2, the same scale factor is further mapped into a learnable scale embedding vector to condition the fusion on the target ratio.
The proposed framework is trained in two consecutive stages, as shown in Figure 1. In Stage 1, a degradation model estimation module estimates the spatial point spread function (PSF) P and spectral response function (SRF) R from the available LR-HSI/HR-MSI pair. In Stage 2, the learned degradation operators are frozen and used as physical degradation priors to train the proposed SCDF-Net fusion network through self-supervised observation-domain consistency.

3.2. Stage 1: Degradation Model Estimation

The purpose of Stage 1 is to estimate the spatial and spectral degradation relationships between the two observed modalities. The SRF R maps the HSI spectral space to the MSI spectral space [3,5,7,11]. For the c -th multispectral band, the spectral degradation is written as
Y c =   R c ( X ) ,   c   = 1 , , C
where R c ( · ) denotes the spectral response mapping from the HSI bands to the c -th MSI band. The SRF matrix ( R   R C × B ) maps B hyperspectral bands to C multispectral bands, and each row corresponds to the spectral response of one MSI band. The SRF weights are constrained to be nonnegative and normalized to preserve physical interpretability. In this work, the SRF is assumed to be spatially invariant, nonnegative, and row-normalized. The learned SRF is therefore used as an effective spectral degradation operator for self-supervised training.
The spatial degradation is modeled by a PSF-based blur-and-resampling operation [3,7,11,26]. The PSF describes local spatial blur, while the spatial size reduction is achieved by resampling the blurred image to the LR-HSI grid. The PSF is not assumed to perform downsampling alone; instead, the spatial size reduction is explicitly determined by the target-grid resampling operation. For an input image Z , the spatial degradation is denoted as
D P , s ( Z )   = S s   ( P * Z )
where * denotes band-wise two-dimensional convolution, and S s ( · ) denotes scale-controlled resampling to the LR-HSI grid. The same blur-and-resampling formulation is used for both integer and non-integer scale factors. Because the same scale-controlled blur-and-resampling operator is used in both degradation estimation and fusion training, any interpolation effect under non-integer scale factors is modeled consistently by this shared operator, so the self-supervised observation-domain losses remain self-consistent across the two stages.
Based on the two degradation operators, the degradation model estimation module constructs two LR-MSI estimates on the LR-HSI grid:
Z 1   =   R ( X LR ) ,   Z 2   =   D P , s ( Y HR )
The first path spectrally projects the LR-HSI into the MSI spectral space. The second path blurs the HR-MSI with the learned PSF and resamples it onto the LR-HSI grid. Therefore, both Z 1 and Z 2 are defined on the LR-HSI reference grid with the same size C × h × w , ensuring that the degradation alignment loss compares spatially aligned LR-MSI observations pixel by pixel. The degradation alignment loss is defined as:
L deg =   Z 1     Z 2 1
where L deg denotes the degradation alignment loss used to train the degradation model estimation module in Stage 1. During optimization, the PSF and SRF are constrained to be nonnegative and normalized. After Stage 1, the learned PSF P and SRF R are fixed and transferred to Stage 2 as frozen degradation operators. The degradation estimation stage plays two important roles in the proposed framework. First, it provides a physically meaningful initialization of the spatial and spectral degradation relationships between the observed LR-HSI and HR-MSI, which alleviates the ill-posedness of the unsupervised fusion problem. Second, the estimated PSF and SRF serve as frozen degradation priors in Stage 2, enabling the construction of self-supervised observation-domain constraints that connect the reconstructed HR-HSI with the observed LR-HSI and HR-MSI. Without these priors, the network would lack explicit guidance on how the observed images are related to the target HR-HSI, making the unsupervised fusion task more challenging.

3.3. Stage 2: Unsupervised Fusion Network Training

In Stage 2, the PSF P and SRF R estimated in Stage 1 are fixed and used as physical degradation priors [26,28,29]. The reconstructed HR-HSI is constrained by two complementary observation-domain constraints, together with HSI-space spectral consistency and MSI-space frequency consistency. The two observation-domain constraints shown in Figure 1—main-reconstruction consistency L HV and low-resolution cycle consistency L LV —provide complementary self-supervision: L HV degrades the predicted HR-HSI to reproduce the observed LR-HSI and HR-MSI, whereas L LV reconstructs the original LR-HSI from a further degraded LR-HSI/LR-MSI pair constructed at cycle ratios randomly sampled from a continuous range. We refer to these two terms jointly as dual observation-domain consistency. It is important to clarify that the term “dual” is used in two different contexts in this paper. The first refers to the two complementary observation-domain constraints ( L HV and L LV ), which we collectively term “dual observation-domain consistency”. The second refers to the two feature-processing pathways in the network architecture, i.e., the frequency-domain and spatial-domain branches, which we term the “dual-domain network”. These two uses of “dual” denote different aspects of the framework and should not be confused. In parallel, the frequency- and spatial-domain branches form the dual-domain detail-modeling pathway; thus, dual observation-domain consistency and dual-domain detail modeling denote different components of the framework.

3.3.1. SCDF-Net

The overall architecture of SCDF-Net is illustrated in Figure 2. SCDF-Net takes the observed LR-HSI X LR and HR-MSI Y HR as inputs. The continuous scale vector s defined in Equation (1) is first mapped into a learnable scale embedding v s . Rather than being concatenated with the input images or encoder features, the scale embedding is projected into the frequency- and spatial-domain gates: the projection is reshaped and broadcast as a scale-conditioned bias map before gate generation, so that the gates are conditioned on the resolution gap of the current input. Although each model is trained for a single target ratio, the low-resolution cycle (Section 3.3.4) reconstructs X LR at cycle ratios randomly sampled from a continuous range, so the scale embedding is exposed to a continuous range of resolution gaps during training.
The scale embedding serves as a bridge between the target resolution ratio and the feature extraction process. Since different scale factors correspond to different degrees of spatial information loss, the network needs to know the current resolution gap to adjust its feature extraction accordingly. The scale embedding encodes this information and modulates the frequency- and spatial-domain gates, allowing the network to adapt its behavior to the specific scale setting. This design is particularly important for fractional scale factors, where the resolution gap cannot be represented by simple integer interpolation.
The dual-domain feature learning strategy is motivated by the complementary nature of frequency-domain and spatial-domain information. The frequency-domain branch is sensitive to global periodic patterns and high-frequency variations, which correspond to texture and edge structures in the image, while the spatial-domain branch captures local structural details that may be less prominent in the frequency representation. By combining these two pathways, the network can recover both global and local spatial structures, leading to improved reconstruction quality.
v s = ϕ ( s ) ,   F X = U ( E X ( X LR ) ; H , W ) ,   F Y =   E Y ( Y HR )
Here, ϕ ( · ) denotes the scale-embedding MLP, and E X and E Y denote the HSI encoder and MSI encoder, respectively. U ( · ) denotes bicubic feature upsampling, which aligns the LR-HSI feature with the HR-MSI spatial grid before dual-domain fusion. The aligned HSI feature F X provides the spectral foundation, while the MSI feature F Y provides high-resolution spatial structures. In the following network description and figures, bold italic symbols denote the aligned encoder features supplied to the two branches, whereas bold upright symbols denote the intermediate and output tensors generated within the frequency- and spatial-domain branches.
After feature encoding and spatial alignment, F X is transformed into the HSI spectrum F X . In parallel, F Y is processed by the scale-conditioned frequency-domain and spatial-domain branches. The frequency-domain branch generates the gated MSI spectrum F Y g , which is fused with F X and converted back to the spatial domain to obtain F freq . The spatial-domain branch produces the scale-conditioned spatial detail feature F Y s . As illustrated in Figure 2, the frequency-fused feature and the spatial-detail feature are added and then passed through the refinement and output heads:
X ^ HR   =   H out H ref F freq + F Y s
Here, F freq denotes the frequency-fused spatial feature obtained after complex Fourier fusion and inverse Fourier transformation, and F Y s denotes the scale-conditioned spatial detail feature produced by the spatial-domain branch. H ref ( · ) denotes the lightweight refinement module, and H out ( · ) maps the refined feature to the hyperspectral output space.

3.3.2. Frequency-Domain Branch

As illustrated in Figure 3, the frequency-domain branch transforms the MSI feature into the Fourier domain and generates a scale-conditioned gate from its magnitude response. The gate modulates the complete complex spectrum, thereby preserving both amplitude and phase information for subsequent fusion.
F X = FFT ( F X ) ,   F Y = FFT ( F Y ) ,   A Y = mean c ( log ( 1 + F Y ) )
where FFT ( · ) denotes the two-dimensional Fourier transform, and mean c ( · ) denotes channel-wise averaging. Here, A Y is the channel-averaged magnitude descriptor of the MSI spectrum. It is computed only from the magnitude spectrum and is used for gate generation. The magnitude descriptor is not used to reconstruct the spectrum directly; it only provides a scale-conditioned gating cue, while the complete complex spectrum F Y is retained for subsequent modulation and fusion. The Fourier transform is adopted in the frequency-domain branch for two main reasons. First, it provides a global frequency representation that directly aligns with the frequency consistency loss defined in Equation (17), enabling efficient comparison of high-frequency components between the reconstructed and observed HR-MSI without introducing additional hyperparameters such as decomposition levels or wavelet bases. Second, the Fourier transform preserves phase information, which is critical for maintaining spatial structures when the modulated complex spectrum is transformed back to the spatial domain. Although wavelet transform offers multi-scale localization, it would increase the complexity of the branch and is not directly required for the proposed frequency consistency constraint. The magnitude descriptor and the scale embedding are jointly fed into a frequency gate generator to produce a scale-conditioned frequency gate:
G f = σ ( g f   ( A Y , v s ) ) ,   F Y g = G f     F Y
where g f ( · , · ) denotes the scale-conditioned frequency gate mapping before sigmoid activation, σ ( · ) denotes the sigmoid function, G f denotes the resulting single-channel frequency gate, denotes element-wise multiplication, and F Y denotes the complete complex MSI spectrum. The gate G f is broadcast across feature channels before modulation, so the complete spectrum, including its phase information, is retained.
After gate modulation, the scale-gated MSI spectrum F Y g is combined with the spectrum F X of the spatially aligned HSI feature through a learnable complex Fourier fusion operator:
F Z   =   F X   +   α C f ( F X ,   F Y g )   +   ( 1 α ) F Y g ,   F freq   =   Re ( IFFT ( F Z ) )
Here, F Z denotes the fused complex spectrum produced by the complex Fourier fusion operator. C f ( · ) predicts a complex residual correction from F X and F Y g , while the complete weighted combination in Equation (10) forms the complex Fourier fusion operation shown in Figure 3. IFFT(·) denotes the inverse Fourier transform, and Re(·) extracts the real component of the inverse-transformed feature. In practice, C f ( · ) is implemented by applying real-valued convolutions to the concatenated real and imaginary components of the two complex spectra. By operating on both real and imaginary components, C f ( · ) exploits amplitude- and phase-related information, so the fusion is not limited to magnitude enhancement. The coefficient α is a learnable scalar bounded in [0, 1], and F freq denotes the frequency-fused spatial feature after inverse Fourier transformation.

3.3.3. Spatial-Domain Branch

Figure 4 illustrates the scale-conditioned spatial-domain branch, which provides a complementary real-domain pathway for extracting edge- and texture-related local structures from the encoded MSI feature. It extracts a spatial-detail proxy from the MSI feature using depthwise and pointwise convolutions, and then uses the scale embedding to control the spatial gate. This branch extracts spatial details only from the MSI feature F Y , while the HSI branch mainly provides the spectral foundation through the main fusion pathway. The projected scale embedding is reshaped and broadcast to the spatial size of the MSI feature before being added to the spatial-detail proxy:
Q Y = W p δ ( W d ( F Y ) ) ,   G s = σ ( g s ( Q Y + Π s ( v s ) ) ) ,   F Y s = G s   F Y
where Q Y denotes the spatial-detail proxy extracted from the encoded MSI feature. W d and W p denote the depthwise and pointwise convolution operators, respectively, and δ ( · ) denotes the GELU activation. Π s ( · ) denotes the scale projection in the spatial-domain branch, while g s ( · ) denotes the spatial gate mapping before sigmoid activation. The resulting single-channel gate G s is broadcast across feature channels and applied to the encoded MSI feature F Y . The output F Y s therefore represents the scale-conditioned spatial-detail feature selected from the MSI feature.
The resulting spatial-detail feature F Y s is added to the frequency-fused spatial feature F freq before being passed to the refinement module, as defined in Equation (7).

3.3.4. Loss Functions

Since HR-HSI ground truth is unavailable in Stage 2, SCDF-Net is trained using self-supervised losses constructed from the observed LR-HSI and HR-MSI [26,27,28,29]. The second-stage objective consists of four terms: low-resolution cycle loss, high-resolution observation-domain loss, HSI-space SAM loss, and MSI-space frequency loss.
For the main reconstruction branch, the network first estimates the HR-HSI:
X ^ HR = f θ ( X LR , Y HR ; ϕ ( s ) )
where f θ ( · ) denotes the SCDF-Net fusion network with learnable parameters θ. The reconstructed HR-HSI is then projected back to the observed LR-HSI and HR-MSI domains:
X ^ LR = D P , s ( X ^ HR ) ,   Y ^ HR = R ( X ^ HR )
Here, D P , s ( · ) denotes the frozen spatial degradation operator, including PSF blur and target-grid resampling, and R ( · ) denotes the frozen SRF-based spectral projection. The resulting X ^ LR and Y ^ HR are the simulated observations generated from the reconstructed HR-HSI.
The high-resolution observation-domain loss is defined as
L HV = X ^ LR X LR 1 + Y ^ HR Y HR 1
where L HV denotes the high-resolution observation-domain loss, including the LR-HSI observation term and the HR-MSI observation term. To provide a direct self-supervised target, a low-resolution cycle is further introduced. At each training iteration, a cycle ratio s LV is randomly sampled from a continuous scale range, with the upper bound adjusted according to the maximum scale factor used in the corresponding experiment. The observed LR-HSI is spatially downsampled by s LV to obtain a lower-resolution HSI X LLR , while the HR-MSI is degraded to the LR-HSI grid (a factor of s) to obtain Y L R . Since Y L R lies on the grid of size [h, w] and X LLR on the grid of size [ h / s LV ,   w / s LV ], the spatial ratio between the two cycle inputs is exactly s LV , which is the value passed to the scale embedding. The same SCDF-Net with shared parameters reconstructs X LR from this degraded pair:
L LV = f θ ( X LLR , Y LR ; ϕ ( s LV ) ) X LR 1
where L L V denotes the low-resolution cycle loss, and s LV is the randomly sampled cycle scale drawn from U ( 1 , s max ) , where s max is adjusted according to the maximum scale factor considered in the corresponding experiment. The original X LR serves as an available reconstruction target, providing direct self-supervision without requiring HR-HSI ground truth. Because s LV varies continuously across iterations, the scale embedding is exposed to different resolution gaps around the target setting, even though each model is optimized for a specific target scale.
The HSI-space SAM loss is used to preserve spectral consistency in the LR-HSI observation space:
L SAM = 1 hw q = 1 hw SAM ( x ^ q ,   x q )
where SAM ( · ) denotes the spectral angle between two spectral vectors, and x ^ q and x q denote the reconstructed and observed LR-HSI spectral vectors at pixel q , respectively.
To enhance high-frequency spatial details, a frequency consistency loss is imposed in the MSI observation space [31]:
L FREQ = M ( | FFT ( Y ^ HR ) | FFT ( Y HR ) ) 1
where M is a binary high-frequency mask that suppresses the central low-frequency region, and Y ^ HR and Y HR denote the simulated and observed HR-MSI, respectively. The binary high-frequency mask M is constructed in the Fourier domain after centering the spectrum via fftshift. The central low-frequency region is masked out, while the surrounding high-frequency region is preserved. In our experiments, the size of the preserved high-frequency region is determined by a keep ratio of 0.15, as described in Section 4.2. The final training objective is
L   = L LV + L HV +   λ SAM L SAM +   λ FREQ L FREQ
where λ SAM and λ FREQ are the weights of the HSI-space SAM loss and MSI-space frequency loss, respectively. The LV and HV losses are used with unit weights because they provide the main reconstruction constraints in the low-resolution cycle and observation domains.
The final objective combines the LV, HV, HSI-space SAM, and FREQ losses. These terms respectively constrain low-resolution self-supervision, high-resolution observation-domain reconstruction, HSI-space spectral preservation, and MSI-space high-frequency consistency. The LV and HV losses are assigned unit weights because they provide the primary observation-domain reconstruction constraints and directly connect the reconstructed HR-HSI with the available LR-HSI and HR-MSI observations. The HSI-space SAM loss and MSI-space frequency loss are used as auxiliary regularization terms; therefore, relatively small weights are adopted to improve spectral-angle consistency and high-frequency spatial consistency without dominating the main observation-domain objective. In our experiments, λSAM and λFREQ were both set to 0.01 according to preliminary convergence observations and the relative numerical magnitudes of different loss terms, which maintained stable optimization while keeping the auxiliary constraints effective.
The four loss terms are designed to address different aspects of the unsupervised fusion problem. L LV and L HV provide the main reconstruction constraints by enforcing consistency between the reconstructed HR-HSI and the observed LR-HSI/HR-MSI through the frozen degradation operators. L SAM is introduced to preserve spectral fidelity, as high PSNR or low RMSE does not guarantee spectral consistency. By enforcing angular similarity between the reconstructed and observed LR-HSI spectra, L SAM encourages the network to maintain the spectral relationships learned from the input. L FREQ targets the recovery of high-frequency spatial details, which are often lost during reconstruction. By aligning the high-frequency components of the MSI spectra, L FREQ encourages the network to preserve fine spatial structures. Together, these four losses jointly guide the network toward a solution that is both spectrally accurate and spatially detailed.

4. Experiments and Results

4.1. Datasets

Experiments were conducted on four widely used hyperspectral remote sensing datasets, including Pavia University (PaviaU) [38], Washington D.C. Mall (WADC) [39], University of Houston 2018 (Houston18) [40], and Chikusei [41]. These datasets contain diverse land-cover types, spatial structures, and spectral characteristics, and therefore allow us to evaluate the fusion performance and robustness of HSI–MSI fusion methods across different scenes and spatial resolution ratios.
Since real paired LR-HSI, HR-MSI, and HR-HSI observations are difficult to obtain, the experiments were conducted following Wald’s protocol [42]. Specifically, the original HSI was regarded as the reference HR-HSI. The LR-HSI was generated by applying spatial degradation to the reference HR-HSI, including PSF-based spatial blurring and scale-controlled downsampling. The HR-MSI was generated by applying spectral degradation to the reference HR-HSI using the SRF. In this way, the reconstructed HR-HSI can be quantitatively compared with the reference image. Since both the LR-HSI and HR-MSI were simulated from the same reference HR-HSI, their spatial correspondence was maintained under the predefined target grids.
The PaviaU dataset captures an urban scene and contains 103 spectral bands ranging from 430 nm to 860 nm. The WADC dataset was acquired by the HYDICE sensor and contains 191 spectral bands covering approximately 400 nm to 2400 nm. The Houston18 dataset used in this study was derived from the 2018 IEEE GRSS Data Fusion Contest. The original hyperspectral data contain 48 spectral bands covering 380–1050 nm with a 1 m ground sampling distance. After removing unusable or noisy bands during preprocessing, 46 bands ranging from 389 nm to 1033 nm were retained in our experiments. The Chikusei dataset covers urban and agricultural areas and contains 128 spectral bands ranging from 363 nm to 1018 nm.
All images were normalized to the range [0, 1] band by band before training and testing. The input LR-HSI/HR-MSI pairs were spatially aligned before fusion. Unlike patch-based training strategies, the proposed method directly used the whole LR-HSI/HR-MSI image pair as input, without random cropping, sliding-window inference, or patch stitching. This whole-image setting avoids additional patch boundary effects and preserves the spatial correspondence between the input observations and the reconstructed HR-HSI.

4.2. Implementation Details

All experiments were implemented in PyTorch 2.0.0, with Python 3.8 and CUDA 11.8, and conducted on a workstation equipped with an NVIDIA GeForce RTX 4090D GPU. The proposed SCDF-Net was trained in an unsupervised manner without using ground-truth HR-HSI supervision. For each dataset and scale factor, the LR-HSI and HR-MSI inputs were generated following the degradation protocol described in Section 3. The same degradation strategy was applied to all compared methods to ensure consistent input conditions and fair evaluation under different scale settings.
The training process followed the two-stage optimization framework introduced in Section 3. In Stage 1, the degradation operators, including the PSF and SRF, were estimated from the observed LR-HSI/HR-MSI pairs. In Stage 2, the learned degradation operators were fixed and used as physical priors to construct self-supervised constraints for optimizing SCDF-Net. For different target scales, the degradation settings were adjusted according to the predefined scale factors. The learned PSF kernels were constrained to be nonnegative and normalized to unit sum before being used for spatial degradation, ensuring the physical consistency of the simulated observations.
To achieve arbitrary-scale fusion, the scale factor was represented as a continuous two-dimensional vector s   =   [ H / h , W / w ] , where H and W denote the spatial dimensions of the HR-MSI, and h and w represent those of the LR-HSI. The scale vector was mapped into a learnable scale embedding, which was subsequently used to condition the frequency-domain and spatial-domain gates. Before feature fusion, the LR-HSI feature was upsampled to the HR-MSI resolution to achieve spatial alignment between heterogeneous observations.
For optimization, the Adam optimizer was adopted with β 1   =   0.5 and β 2   =   0.999 [43]. The initial learning rate was set to 1   ×   1 0 4 , and the number of training epochs was set to 10,000. Since the LR-HSI and HR-MSI observations were spatially aligned through the degradation-based simulation process, whole-image training was adopted instead of patch-based training. The batch size was set to 1 to preserve the complete spatial correspondence between the input observations and the self-supervised degradation constraints.
Following the loss definitions in Section 3, the low-resolution cycle constraint and the high-resolution observation-domain constraint were implemented using L1 reconstruction losses. During training, the low-resolution cycle ratio s LV was randomly sampled from a continuous scale range, with the upper bound adjusted according to the maximum scale factor used in the corresponding experiment. The HSI-space SAM loss and MSI-space frequency consistency loss were weighted by λ SAM   =   0.01 and λ FREQ   =   0.01 , respectively. These weights were empirically determined based on convergence behavior and validation performance. Meanwhile, the relative magnitudes of different loss terms were considered to maintain comparable gradient contributions during optimization and avoid domination by any single constraint. The high-frequency preservation ratio in the frequency consistency loss was set to 0.15.
For a fair comparison, all baseline methods were evaluated using the same simulated LR-HSI and HR-MSI inputs, normalization strategy, degradation settings, and evaluation metrics. For methods originally designed for fixed integer scale factors, the input observations were adapted to the target scale through interpolation-based resampling when necessary, and the models were retrained or re-executed under the corresponding scale settings. For UMC2FF, the available implementation was adopted and trained following the recommended settings as closely as possible. The input generation protocol, normalization strategy, and evaluation criteria were kept consistent with those used for SCDF-Net and other compared methods.
We further evaluated the computational complexity and efficiency of SCDF-Net in terms of the number of trainable parameters, training time, and inference time. SCDF-Net contains approximately 3.1 M trainable parameters. On an NVIDIA GeForce RTX 4090D GPU, the average training time is approximately 0.5 h for 10,000 epochs. Since the benchmark datasets have different spatial resolutions and image sizes, the average inference time over all four datasets is approximately 1.0 s per image (batch size = 1). For the efficiency comparison reported in Table 1, the inference time of all deep learning-based methods was measured on the Houston18 dataset under the same hardware environment and identical input setting, yielding an inference time of 0.0071 s per image for SCDF-Net. The frequency-domain branch mainly uses FFT-based transformation and lightweight gate generation, while the spatial-domain branch is implemented with depthwise and pointwise convolutions. Therefore, the additional computational cost introduced by the dual-domain branches remains moderate.

4.3. Evaluation Metrics

Quantitative performance was evaluated using five metrics: peak signal-to-noise ratio (PSNR), root mean square error (RMSE), spectral angle mapper (SAM), relative dimensionless global error in synthesis (ERGAS), and correlation coefficient (CC) [2,3,4,27,30]. PSNR and RMSE measure absolute reconstruction accuracy, ERGAS evaluates band-normalized global relative error, CC measures band-wise correlation, whereas SAM evaluates spectral fidelity.
For the reconstructed HR-HSI X ^ HR and reference HR-HSI X HR , PSNR, RMSE, and ERGAS are computed as
PSNR   = 10 log 10 ( BHWL 2 X ^ HR X HR F 2 )
RMSE = X ^ HR X HR F BHW
where B is the number of spectral bands, H × W is the spatial size, L   =   1 is the maximum pixel value after normalization, and   F denotes the Frobenius norm.
ERGAS = 100 r 1 B b = 1 B RMSE b μ b 2
where r denotes the spatial scale factor, RMSE b is the RMSE of the b-th spectral band, and μ b is the mean value of the b-th band in the reference HR-HSI. ERGAS measures the global relative reconstruction error after band-wise normalization, and lower ERGAS values indicate better overall fidelity.
CC = 1 B b = 1 B i = 1 HW x ^ I , b x ^ ¯ b x I , b x ¯ b i = 1 HW x ^ I , b x ^ ¯ b 2 i = 1 HW x I , b x ¯ b 2
where x ¯ I , b and x I , b denote the reconstructed and reference pixel values at spatial pixel i and spectral band b, respectively. x ^ ¯ b   and   x ¯ b are the mean values of the reconstructed and reference images in the b -th spectral band. CC measures the band-wise linear correlation between the reconstructed and reference HR-HSI, and the final value is averaged over all spectral bands. Higher CC values indicate better reconstruction performance.
SAM is obtained by averaging the spectral angle over all spatial pixels:
SAM = 1 N i = 1 N arccos x ^ i T x i x ^ i 2 x i 2 + ε ,   N = HW
where x ^ I and x I are the reconstructed and reference spectral vectors at pixel i , respectively, and ε is a small constant for numerical stability. The resulting mean spectral angle is converted from radians to degrees for reporting. Higher PSNR, CC and lower RMSE, ERGAS, and SAM values indicate better reconstruction performance.
Residual, MRAE, and SAM maps were additionally used for visual evaluation [27,30]. The pixel-wise MRAE map is defined as
MRAE i = 1 B b = 1 B x ^ I , b x I , b | x I , b | + ε
where i and b index the spatial pixels and spectral bands, respectively. The residual map shows absolute reconstruction errors, whereas the MRAE and SAM maps visualize the spatial distributions of relative errors and spectral distortion, respectively. These maps are used only for qualitative comparison and are not reported as scalar metrics.

4.4. Performance of Different Datasets at Different Scales

To evaluate the performance of SCDF-Net under different target scale factors, experiments were conducted under multiple spatial scale factors, including both integer and fractional ratios. Table 2 and Table 3 report the quantitative results on PaviaU, WADC, Houston18, and Chikusei.
As shown in Table 2 and Table 3, the reconstruction performance generally decreases as the scale factor increases. This is reasonable because larger scale factors lead to more severe spatial information loss in LR-HSIs and make HR-HSI reconstruction more challenging. Nevertheless, SCDF-Net maintains stable performance across all datasets. In particular, the results at fractional scale factors, such as ×1.5, ×2.4, and ×3.6, demonstrate that the proposed method can effectively handle non-integer spatial resolution ratios.
The performance differences among datasets are related to their scene complexity, spatial structure, and spectral variability. For example, WADC and Chikusei obtain relatively high PSNR values, while Houston18 presents a more challenging urban scene with complex spatial patterns. Overall, the multi-scale results verify that the proposed scale-controlled degradation and continuous scale embedding improve the adaptability of SCDF-Net under both integer and fractional scale factors. Notably, SCDF-Net is also evaluated at larger scale factors, including ×5.0 and ×6.0, which further examines its stability under more severe spatial information loss.
Table 2. Quantitative evaluation of SCDF-Net on PaviaU and WADC under different scale factors.
Table 2. Quantitative evaluation of SCDF-Net on PaviaU and WADC under different scale factors.
ScalesPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
1.542.85952.42030.00724.48950.992651.19571.35980.00284.04840.9791
2.042.07872.47650.00793.52090.992148.95871.73070.00362.47970.9732
2.440.43732.72020.00952.06290.991247.56501.94840.00422.56640.9724
3.039.77332.84470.01032.62840.989947.51592.06850.00421.89840.9705
3.639.27492.94150.01092.23330.989446.22452.47590.00491.84350.9677
4.038.95113.01690.01131.64210.988945.41892.71390.00541.53190.9623
5.038.70183.19300.01161.47740.987644.97163.24940.00561.68520.9567
6.038.58853.27360.01181.24100.986444.75583.32910.00581.42760.9538
Table 3. Quantitative evaluation of SCDF-Net on Houston18 and Chikusei under different scale factors.
Table 3. Quantitative evaluation of SCDF-Net on Houston18 and Chikusei under different scale factors.
ScalesHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
1.546.47201.25040.00472.36330.998350.90841.10250.00287.02010.9853
2.045.24811.30930.00551.82890.998950.14001.12150.00314.98870.9851
2.443.48041.38110.00671.77080.997949.15001.14990.00354.17060.9850
3.043.44721.46950.00671.59830.997748.72001.23450.00373.44550.9842
3.641.53001.68920.00841.63430.997148.65001.28470.00372.91320.9842
4.041.35501.69570.00861.55550.995747.65001.38700.00412.82900.9813
5.039.67082.25640.01041.38310.993745.13741.81660.00552.29330.9792
6.037.87632.87000.01281.66600.991044.03391.94140.00632.03690.9758

4.5. Baseline Comparison

For a fair comparison under different scale factors, all baseline methods were independently retrained or re-run for each target scale factor using the same simulated LR-HSI and HR-MSI inputs. Specifically, the LR-HSI was generated by the same PSF-based spatial degradation and target-grid resampling protocol, while the HR-MSI was generated using the same SRF-based spectral degradation. The training and testing data, normalization strategy, degradation setting, and evaluation metrics were kept identical for all methods. For methods originally designed for fixed integer scales, the input image pairs were regenerated on the corresponding target grid, and the model was retrained or re-executed under each scale setting. Fixed integer-ratio baselines cannot operate under ×2.4 directly, so they were adapted by interpolation-based resampling before training/testing, following the standard workaround for such methods. This adaptation may introduce additional approximation error; therefore, the ×2.4 comparison is interpreted as a practical comparison between fixed-scale adaptation and the proposed native fractional-scale fusion. All outputs were evaluated on the same HR-HSI grid with identical metrics. For learning-based baseline methods, we used the publicly available implementations when available and followed their recommended training settings. All baselines were trained or re-run under the same degradation setting and evaluation protocol, and the reported results were obtained after convergence.
Before comparing different fusion methods, we further analyze the degradation operators estimated in Stage 1, because the learned PSF and SRF are frozen and used as physical priors in Stage 2. Under Wald’s protocol, the spatial degradation kernel is predefined during data generation, which allows us to compare the estimated PSF with the reference PSF. Specifically, both kernels are normalized to unit sum, and the mean squared error (MSE) between them is computed. Under the ×2.4 scale setting, the PSF MSE values are 0.000059, 0.000041, 0.000123, and 0.000109 on PaviaU, WADC, Houston18, and Chikusei, respectively. These small errors indicate that Stage 1 can reliably recover the spatial degradation patterns across different scenes.
For the SRF, direct element-wise comparison is less straightforward because the learned spectral mapping and the simulated spectral response may use different parameterizations. Therefore, we evaluate its effect through observation-domain reconstruction consistency: the LR-HSI projected by the learned SRF should be consistent with the MSI observation after spectral degradation. Together, the low PSF estimation error and the observation-domain reconstruction consistency support the use of the learned PSF/SRF as frozen degradation priors in Stage 2.
We compared SCDF-Net with representative traditional and deep learning-based HSI–MSI fusion methods. The traditional methods include GSA [8], CNMF [5], and HySure [7], while the deep learning-based methods include CLoRF [30], UDALN [26], u2MDN [27], and UMC2FF [28]. The quantitative comparison results are reported in Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9 and Table 10. In all quantitative comparison tables, bold and underlined values indicate the best and second-best results, respectively. In addition to PSNR, SAM, and RMSE, ERGAS and CC are also considered to evaluate band-normalized global error and spectral correlation, respectively. Overall, traditional methods generally show lower performance than deep learning-based methods because they rely on handcrafted priors, linear mixture assumptions, or fixed degradation models. These assumptions limit their flexibility when the spatial resolution ratio changes, especially under fractional and large-scale settings.

4.5.1. Comparison at Scale Factor ×1.5

Table 4 and Table 5 present the quantitative results at scale factor ×1.5, which is a fractional scale setting where fixed-integer methods require interpolation-based input adaptation. Traditional methods show limited performance, as their handcrafted priors and fixed degradation assumptions are insufficient to handle complex spatial–spectral relationships under non-integer scales. Deep learning-based methods generally achieve better reconstruction quality. Among them, UMC2FF achieves relatively strong results, particularly on Houston18 and Chikusei. SCDF-Net reports the highest PSNR and lowest SAM on all four datasets. Compared with the strongest PSNR baseline in each dataset, SCDF-Net improves PSNR by 1.2369 dB on PaviaU, 9.3838 dB on WADC, 1.4069 dB on Houston18, and 2.1754 dB on Chikusei. For ERGAS and CC, SCDF-Net remains close to the strongest results on most datasets, while some baselines obtain lower ERGAS or higher CC on individual datasets. The consistent improvements indicate that operating directly on the native fractional grid, rather than resampling to an integer ratio, is beneficial for fractional-scale fusion.
Table 4. Quantitative comparison on PaviaU and WADC at scale factor ×1.5.
Table 4. Quantitative comparison on PaviaU and WADC at scale factor ×1.5.
MethodsPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA33.90904.05290.02027.73900.978936.84553.44890.01435.54950.9748
HySure28.02795.56870.039713.68390.915331.43055.24430.026810.94010.8260
CNMF37.31743.51950.01366.32280.974735.69972.38950.01644.36820.9699
CLoRF37.11713.50060.01395.91110.987936.64133.78880.01478.15870.8884
UDALN37.50613.39820.01336.11410.986335.98712.03850.01593.89890.9640
u2MDN39.16972.93620.01104.75510.991141.81193.08190.00818.51110.9479
UMC2FF41.62262.95700.00834.58470.990640.08021.55390.00992.55250.9605
Proposed42.85952.42030.00724.48950.992651.19571.35980.00284.04840.9791
Note: Bold and underlined values denote the best and second-best results, respectively.
Table 5. Quantitative comparison on Houston18 and Chikusei at scale factor ×1.5.
Table 5. Quantitative comparison on Houston18 and Chikusei at scale factor ×1.5.
MethodsHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA40.78311.68270.00913.70980.996040.35231.59810.00965.73610.9829
HySure33.71992.62600.02068.39850.978548.73302.72490.003712.70470.9199
CNMF37.80033.26140.01296.41630.981143.56131.65380.00667.12340.9493
CLoRF41.29732.18420.00863.70400.995445.20461.69040.00558.46970.9747
UDALN39.42202.00050.01073.29570.997736.37323.29180.015213.06140.9677
u2MDN43.30151.85520.00682.92990.997746.23343.48490.00497.13000.9831
UMC2FF45.06511.77570.00562.35100.998046.91081.67270.00446.69780.9827
Proposed46.47201.25040.00472.36330.998350.90841.10250.00287.02010.9853
Note: Bold and underlined values denote the best and second-best results, respectively.
Table 6. Quantitative comparison on PaviaU and WADC at scale factor ×2.0.
Table 6. Quantitative comparison on PaviaU and WADC at scale factor ×2.0.
MethodsPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA30.41305.40540.03026.40310.963933.08595.72830.02225.46690.9541
HySure35.94323.36600.016017.67470.750437.90162.12390.012714.20710.6433
CNMF41.56043.57160.00845.02490.967236.16012.86940.01563.67450.9537
CLoRF38.31163.15020.01213.87470.990337.36622.52050.01355.81010.9022
UDALN40.47692.90710.00954.33860.987834.12943.93340.01973.42450.9650
u2MDN39.20492.91970.01103.98470.988845.26541.78970.00553.13330.9303
UMC2FF41.84442.82340.00813.27040.991340.15232.06610.00983.28900.9616
Proposed42.07872.47650.00793.52090.992148.95871.73070.00362.47970.9732
Note: Bold and underlined values denote the best and second-best results, respectively.
Table 7. Quantitative comparison on Houston18 and Chikusei at scale factor ×2.0.
Table 7. Quantitative comparison on Houston18 and Chikusei at scale factor ×2.0.
MethodsHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA34.03462.97170.01993.30350.992535.25892.51510.01734.99850.9716
HySure40.76272.34440.009213.12230.908442.87721.51690.007215.39220.8104
CNMF43.78201.73810.00655.40080.976541.54331.91600.00845.84580.9505
CLoRF41.23862.19530.00872.28440.996845.86971.57110.00516.14750.9778
UDALN41.03172.34490.00893.01340.996540.18091.70970.00988.30840.9695
u2MDN44.22631.76540.00612.91040.995748.62401.48740.00375.25560.9840
UMC2FF44.95431.70590.00571.67400.998049.18341.39810.00356.87370.9747
Proposed45.24811.30930.00551.82890.998950.14001.12150.00314.98870.9851
Note: Bold and underlined values denote the best and second-best results, respectively.
Table 8. Quantitative comparison on PaviaU and WADC at scale factor ×2.4.
Table 8. Quantitative comparison on PaviaU and WADC at scale factor ×2.4.
MethodsPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA32.52425.00660.02365.92090.962535.46594.75110.01694.69520.9525
HySure25.88286.57620.050811.16330.840628.22827.91160.03889.63390.7327
CNMF36.73303.82890.01464.19050.971033.74113.44580.02063.79100.9506
CLoRF37.09883.49030.01403.80930.987736.18143.90010.01555.30410.8810
UDALN39.50643.09200.01064.03620.985634.01283.04270.01992.57010.9626
u2MDN38.72933.12100.01163.11670.990542.46232.12620.00752.47170.9587
UMC2FF40.14513.00140.00982.76900.991039.72983.25160.01032.78470.9599
Proposed40.43732.72020.00952.06290.991247.56501.94840.00422.56640.9724
Note: Bold and underlined values denote the best and second-best results, respectively.
Table 9. Quantitative comparison on Houston18 and Chikusei at scale factor ×2.4.
Table 9. Quantitative comparison on Houston18 and Chikusei at scale factor ×2.4.
MethodsHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA38.12822.44740.01243.15560.991937.19092.10160.01384.55620.9678
HySure29.24574.50270.03458.31530.941834.76333.14080.01839.30210.8877
CNMF39.78022.23660.01033.61080.984839.73622.06390.01034.79890.9522
CLoRF40.85852.03370.00912.41020.995043.83361.92610.00645.24710.9728
UDALN40.59761.76650.00931.92490.997736.53542.03660.01497.74540.9665
u2MDN40.09441.84180.00991.80610.997845.47291.89820.00535.18630.9805
UMC2FF41.73541.46690.00821.40640.998145.55991.82260.00534.05820.9824
Proposed43.48041.38110.00671.77080.997949.15001.14990.00354.17060.9850
Note: Bold and underlined values denote the best and second-best results, respectively.
Table 10. Quantitative comparison on PaviaU and WADC at scale factor ×4.0.
Table 10. Quantitative comparison on PaviaU and WADC at scale factor ×4.0.
MethodsPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA28.99038.48440.03553.87000.937131.27497.75870.02733.65570.9186
HySure31.03324.91110.02817.60830.761033.38325.39820.02147.06040.6490
CNMF38.25594.10090.01222.77600.958833.33514.70700.02153.29450.9413
CLoRF37.71583.32990.01302.04090.989035.26524.30430.01723.48470.8690
UDALN37.29713.91980.01372.33420.986733.52384.81350.02111.72800.9637
u2MDN38.07203.48080.01251.99030.988843.37253.05220.00681.56630.9303
UMC2FF38.82723.78700.01141.68110.990839.77993.45870.01031.53250.9601
Proposed38.95113.01690.01131.64210.988945.41892.71390.00541.53190.9623
Note: Bold and underlined values denote the best and second-best results, respectively.

4.5.2. Comparison at Scale Factor ×2.0

Table 6 and Table 7 report the quantitative comparison results at scale factor ×2.0. Under this relatively moderate integer-scale setting, traditional methods provide reasonable results on some datasets, but they are generally inferior to deep learning-based methods because fixed observation assumptions are insufficient to model complex spatial–spectral relationships. SCDF-Net reports the highest PSNR and lowest SAM on all four datasets. Compared with the strongest PSNR baseline in each dataset, SCDF-Net improves PSNR by 0.2343 dB on PaviaU, 3.6933 dB on WADC, 0.2938 dB on Houston18, and 0.9566 dB on Chikusei. The ERGAS and CC results show a similar trend on most datasets, indicating that SCDF-Net improves reconstruction accuracy while maintaining spectral correlation. These results show that the proposed degradation alignment and frequency–spatial detail fusion improve reconstruction quality under the integer scale setting.

4.5.3. Comparison at Scale Factor ×2.4

Table 8 and Table 9 present the results at the fractional scale factor ×2.4. Compared with integer-scale settings, this setting additionally evaluates fractional-ratio fusion, where fixed-scale baselines require interpolation-based adaptation and SCDF-Net incorporates the scale factor as an explicit continuous condition. Under this condition, traditional methods show more obvious limitations, while deep learning-based methods generally achieve better reconstruction performance. SCDF-Net reports the highest PSNR and lowest RMSE on all four datasets, improving over the strongest PSNR baseline by 0.2922 dB, 5.1027 dB, 1.7450 dB, and 3.5901 dB, respectively. In addition, SCDF-Net obtains the lowest SAM on all four datasets. The ERGAS and CC values further show that SCDF-Net maintains stable band-normalized reconstruction quality and spectral correlation under the fractional scale setting. This result suggests that operating directly on the native fractional grid, rather than resampling to an integer ratio, is beneficial under non-integer scale factors.

4.5.4. Comparison at Scale Factor ×4.0

Table 10 and Table 11 show the comparison results at scale factor ×4.0. As the scale factor increases, more spatial information is missing from LR-HSIs, and the fusion task becomes more difficult. At this larger scale factor, the gap between traditional methods and learning-based methods becomes more evident, since handcrafted priors and fixed degradation assumptions are less effective when spatial details are severely degraded. SCDF-Net still reports the highest PSNR and lowest SAM on all four datasets. Compared with the strongest PSNR baseline in each dataset, SCDF-Net improves PSNR by 0.1239 dB on PaviaU, 2.0464 dB on WADC, 0.0162 dB on Houston18, and 1.0710 dB on Chikusei. For RMSE, SCDF-Net obtains the lowest values on PaviaU, WADC, and Chikusei, and remains close to the lowest reported value on Houston18. The ERGAS and CC results indicate that SCDF-Net remains stable under larger spatial resolution gaps, although the correlation metric is affected by dataset-specific spectral variability. These results suggest that the proposed scale-conditioned frequency–spatial detail fusion can better adapt to larger spatial resolution gaps and preserve spatial details under more challenging reconstruction conditions.

4.6. Visual Evaluation

Figure 5, Figure 6, Figure 7 and Figure 8 show the visual comparison results of different fusion methods on PaviaU, Chikusei, Houston18, and WADC at scale factor ×2.4, respectively. For each dataset, the reconstructed images are shown together with residual, MRAE, and SAM maps to visualize pixel-wise reconstruction errors, relative reconstruction errors, and spectral distortions.
Compared with traditional methods, SCDF-Net produces clearer spatial structures and lower error responses. GSA and HySure tend to generate larger residual values around edges and high-contrast regions, indicating insufficient detail preservation or spectral distortion. CNMF preserves some spectral characteristics but may lose fine spatial structures. Deep learning-based methods improve visual quality compared with traditional methods, but some of them still exhibit noticeable errors around building boundaries, roads, vegetation regions, and field edges.
Across the four datasets, SCDF-Net generally produces lower residual and MRAE responses in structural regions and fewer high-error regions in the SAM maps. On urban scenes such as PaviaU, Houston18, and WADC, the proposed method better preserves building boundaries and road structures. On Chikusei, it maintains clearer agricultural field boundaries and reduces spectral distortion in heterogeneous regions. These visual results are consistent with the quantitative metrics and further demonstrate the effectiveness of the scale-conditioned frequency-domain branch and spatial-domain branch.
To further evaluate spectral fidelity, we compare the spectral reflectance curves on several representative pixels. Figure 9, Figure 10, Figure 11 and Figure 12 show the results on four datasets at ×2.4 scale. The proposed method generally follows the reference spectral curves more closely than the compared methods at representative pixels, especially in bands where several baselines show visible deviations. This comparison indicates better spectral agreement between the reconstructed and reference spectra. Additional spectral curve comparisons at scale factors ×2.0 and ×4.0 are provided in Appendix A.2 (Figure A9, Figure A10, Figure A11, Figure A12, Figure A13, Figure A14, Figure A15 and Figure A16).

4.7. Ablation Study

To evaluate the contribution of the main components in SCDF-Net, ablation experiments were conducted under scale factors ×2.0, ×2.4, and ×4.0. The ablation study contains two groups of variants. The first group removes network components, including removing both the frequency-domain and spatial-domain branches, removing only the frequency-domain branch, removing only the spatial-domain branch, and removing the scale-embedding module. In the variant without scale embedding, the scale-projection terms are removed from the frequency- and spatial-domain gates, while the remaining branches, losses, and training settings are kept unchanged. The second group removes the two auxiliary loss terms, including W/O LFREQ and W/O LSAM. In W/O LFREQ, the MSI-space frequency consistency loss is removed; in W/O LSAM, the HSI-space SAM consistency loss is removed. The LV and HV losses are kept in all variants because they provide the main self-supervised observation-domain reconstruction constraints. The results are reported in Table 12, Table 13, Table 14, Table 15, Table 16 and Table 17.
Across different scale factors, removing both the frequency-domain and spatial-domain branches leads to the most obvious performance degradation in most cases. The consistent drop shows that direct fusion without explicit frequency- and spatial-domain detail enhancement is insufficient for high-quality HR-HSI reconstruction. When only the frequency-domain branch is removed, the model loses part of its ability to recover high-frequency spatial details. When only the spatial-domain branch is removed, the performance also decreases, showing that local spatial structural guidance is complementary to frequency-domain detail modeling.
The comparison among W/O FFT, W/O spatial, and W/O FFT and spatial further reveals the interaction between the two branches. Removing either branch alone causes a moderate performance drop, whereas removing both branches simultaneously leads to a much larger degradation. This indicates that the frequency-domain and spatial-domain branches are not simply independent add-ons. Instead, they provide complementary information: the frequency-domain branch emphasizes high-frequency spectral–spatial responses, while the spatial-domain branch preserves local structural details in the real domain.
Removing the scale-embedding module also causes consistent performance degradation across different datasets and scale factors. Compared with the full model, the W/O scale-embedding variant reduces the average PSNR by about 3.05 dB over the three scale settings. Although each model targets a single ratio, the low-resolution cycle exposes the scale embedding to a continuous range of cycle ratios s LV during training; the embedding is therefore the signal that informs the frequency- and spatial-domain gates of the current resolution gap.
For example, on the PaviaU dataset at ×2.4 scale, removing the frequency-domain branch alone causes a PSNR drop of 1.6419 dB, and removing the spatial-domain branch alone causes a drop of 1.6726 dB. In contrast, removing both branches simultaneously results in a much larger drop of 8.5344 dB. This larger degradation further supports the complementary interaction between the two branches.
We further evaluate the contribution of the two auxiliary loss terms by removing L FREQ and L SAM individually. As shown in Table 12, Table 13, Table 14, Table 15, Table 16 and Table 17, removing either auxiliary loss term leads to performance degradation across most datasets and scale factors. Removing L SAM mainly weakens spectral-angle consistency and usually increases SAM, while removing L FREQ reduces the frequency-domain constraint on high-frequency spatial details and leads to lower reconstruction accuracy. These results indicate that L SAM and L FREQ provide complementary regularization to the main LV and HV observation-domain losses.
The full SCDF-Net model generally gives the strongest overall results across the ablation tables. The comparison shows that the frequency-domain branch, spatial-domain branch, and scale-embedding module all contribute to the final reconstruction quality. The frequency-domain branch helps recover high-frequency information, the spatial-domain branch provides local structural guidance, and the scale embedding further improves adaptation to different scale factors. Their combination enables SCDF-Net to better preserve spatial details while maintaining spectral fidelity.

5. Discussion

The experimental results on four benchmark datasets demonstrate that SCDF-Net improves both spatial reconstruction quality and spectral fidelity under integer and fractional scale factors. These gains mainly come from the scale-conditioned degradation alignment and the dual-domain detail modeling. The degradation estimation stage provides physically meaningful PSF/SRF priors, which alleviates the ill-posedness of unsupervised fusion. The continuous scale embedding adapts the network to different resolution gaps without resampling inputs to an integer ratio, which is particularly important for fractional ratios. The frequency-domain and spatial-domain branches then recover global high-frequency and local structural details in a complementary manner, while the observation-domain and auxiliary regularizations jointly preserve spectral fidelity. Ablation results further confirm that the frequency-domain branch, the spatial-domain branch, and the scale embedding are mutually reinforcing rather than independent add-ons.
Nevertheless, SCDF-Net still has two main limitations. First, the model contains approximately 3.1 M trainable parameters. Although its inference speed is competitive compared with representative deep learning-based methods, the training and storage costs may limit deployment in resource-constrained scenarios. Second, although the model supports fractional scale factors, its robustness under extremely large resolution gaps and severe spatial degradation has not been systematically validated, especially on real cross-sensor image pairs. Future work will therefore focus on developing more lightweight architectures for efficient training and deployment, and extending the framework to broader scale settings, severe degradation conditions, and real cross-sensor acquisition scenarios.

6. Conclusions

This paper proposes SCDF-Net, an unsupervised scale-conditioned framework for hyperspectral–multispectral image fusion under integer and fractional spatial resolution ratios. The proposed method addresses the fusion of LR-HSIs and HR-MSIs under both integer and non-integer spatial resolution ratios without requiring ground-truth HR-HSIs for training. To improve degradation consistency under fractional scale factors, a two-stage training strategy was developed. In the first stage, the PSF and SRF were estimated from the observed image pair. In the second stage, the learned degradation operators were frozen as physical priors and used to construct scale-conditioned observation-domain constraints for unsupervised fusion training.
SCDF-Net further introduces a continuous scale embedding module and a scale-conditioned frequency–spatial dual-domain fusion network. The scale embedding conditions the fusion process according to the actual resolution gap between the LR-HSI and HR-MSI. The frequency-domain branch enhances high-frequency spatial details, while the spatial-domain branch provides complementary local structural information. Combined with observation-domain consistency, low-resolution self-supervision, HSI-space SAM consistency, and MSI-space frequency consistency, the proposed framework helps preserve spectral fidelity and spatial details during reconstruction under different target scale factors.
Experiments on four benchmark datasets demonstrate that SCDF-Net maintains stable reconstruction performance under both integer and fractional scale factors. Quantitative comparisons show that SCDF-Net achieves the strongest overall performance in most cases and consistently provides low spectral distortion. For example, at ×2.4 scale on the PaviaU dataset, SCDF-Net achieves a PSNR of 40.4373 dB, which is 0.2922 dB higher than that of UMC2FF, and reduces the SAM to 2.7202°, compared with 3.0014° obtained by UMC2FF. On the Chikusei dataset, SCDF-Net improves PSNR by 3.5901 dB over the strongest PSNR baseline and achieves the lowest SAM of 1.1499°. These quantitative gains, together with the visual results, support the effectiveness of the proposed scale-conditioned dual-domain fusion framework. Visual results further confirm its ability to reduce spatial artifacts and spectral distortions, while ablation studies verify the complementary contributions of the frequency-domain and spatial-domain branches.
Future work will focus on validating the framework on real cross-sensor image pairs, improving computational efficiency, and enhancing robustness to severe misregistration, noise, and multi-temporal differences.

Author Contributions

Conceptualization, P.T. and K.Z.; methodology, P.T. and K.Z.; software, P.T.; validation, P.T., J.L. and H.Y.; formal analysis, P.T.; investigation, P.T.; resources, K.Z. and X.S.; data curation, P.T. and J.L.; writing—original draft preparation, P.T.; writing—review and editing, K.Z., J.L., H.Y. and X.S.; visualization, P.T.; supervision, K.Z.; project administration, K.Z.; funding acquisition, K.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Ministry of Science and Technology of the People’s Republic of China under Grants 2025ZD1008600 and 2025ZD1008604; the National Natural Science Foundation of China under Grants 42671613, 42471380 and 42201362; the Shandong Provincial Natural Science Foundation of China under Grant ZR2024QD042; the Science and Technology Research Program of Chongqing Municipal Education Commission under Grant KJQN202620606; the Postdoctoral Fellowship Program (Grade B) of China Postdoctoral Science Foundation under Grant GZB20260443; and the Chongqing Postdoctoral Science Foundation under Grant 2025CQBSHTB2012.

Data Availability Statement

The datasets used in this study are publicly available benchmark hyperspectral datasets, including Pavia University, Washington D.C. Mall (WADC), the University of Houston 2018 (2018 IEEE GRSS Data Fusion Contest), and Chikusei. Further information regarding the processed data and experimental code is available from the corresponding author upon reasonable request.

Acknowledgments

The authors would like to thank the providers of the public benchmark hyperspectral datasets used in this work.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Appendix A.1. Visual Comparison of Fusion Results

This appendix presents qualitative results of SCDF-Net and baseline methods on representative regions from the four benchmark datasets (PaviaU, Houston18, Chikusei, and WADC), as shown in Figure A1, Figure A2, Figure A3, Figure A4, Figure A5, Figure A6, Figure A7 and Figure A8.
Figure A1. Visual comparison of different fusion methods on PaviaU at scale factor ×2.0.
Figure A1. Visual comparison of different fusion methods on PaviaU at scale factor ×2.0.
Remotesensing 18 03187 g0a1
Figure A2. Visual comparison of different fusion methods on Houston18 at scale factor ×2.0.
Figure A2. Visual comparison of different fusion methods on Houston18 at scale factor ×2.0.
Remotesensing 18 03187 g0a2
Figure A3. Visual comparison of different fusion methods on Chikusei at scale factor ×2.0.
Figure A3. Visual comparison of different fusion methods on Chikusei at scale factor ×2.0.
Remotesensing 18 03187 g0a3
Figure A4. Visual comparison of different fusion methods on WADC at scale factor ×2.0.
Figure A4. Visual comparison of different fusion methods on WADC at scale factor ×2.0.
Remotesensing 18 03187 g0a4
Figure A5. Visual comparison of different fusion methods on PaviaU at scale factor ×4.0.
Figure A5. Visual comparison of different fusion methods on PaviaU at scale factor ×4.0.
Remotesensing 18 03187 g0a5
Figure A6. Visual comparison of different fusion methods on Houston18 at scale factor ×4.0.
Figure A6. Visual comparison of different fusion methods on Houston18 at scale factor ×4.0.
Remotesensing 18 03187 g0a6
Figure A7. Visual comparison of different fusion methods on Chikusei at scale factor ×4.0.
Figure A7. Visual comparison of different fusion methods on Chikusei at scale factor ×4.0.
Remotesensing 18 03187 g0a7
Figure A8. Visual comparison of different fusion methods on WADC at scale factor ×4.0.
Figure A8. Visual comparison of different fusion methods on WADC at scale factor ×4.0.
Remotesensing 18 03187 g0a8

Appendix A.2. Spectral Curve Comparison at Different Scales

This appendix presents spectral curve comparisons of different fusion methods on representative pixels from the four benchmark datasets (PaviaU, Houston18, Chikusei, and WADC) at scale factors ×2.0 and ×4.0, as shown in Figure A9, Figure A10, Figure A11, Figure A12, Figure A13, Figure A14, Figure A15 and Figure A16.
Figure A9. Spectral curve comparison on Chikusei at scale factor ×2.0.
Figure A9. Spectral curve comparison on Chikusei at scale factor ×2.0.
Remotesensing 18 03187 g0a9
Figure A10. Spectral curve comparison on Houston18 at scale factor ×2.0.
Figure A10. Spectral curve comparison on Houston18 at scale factor ×2.0.
Remotesensing 18 03187 g0a10
Figure A11. Spectral curve comparison on PaviaU at scale factor ×2.0.
Figure A11. Spectral curve comparison on PaviaU at scale factor ×2.0.
Remotesensing 18 03187 g0a11
Figure A12. Spectral curve comparison on WADC at scale factor ×2.0.
Figure A12. Spectral curve comparison on WADC at scale factor ×2.0.
Remotesensing 18 03187 g0a12
Figure A13. Spectral curve comparison on Chikusei at scale factor ×4.0.
Figure A13. Spectral curve comparison on Chikusei at scale factor ×4.0.
Remotesensing 18 03187 g0a13
Figure A14. Spectral curve comparison on Houston18 at scale factor ×4.0.
Figure A14. Spectral curve comparison on Houston18 at scale factor ×4.0.
Remotesensing 18 03187 g0a14
Figure A15. Spectral curve comparison on PaviaU at scale factor ×4.0.
Figure A15. Spectral curve comparison on PaviaU at scale factor ×4.0.
Remotesensing 18 03187 g0a15
Figure A16. Spectral curve comparison on WADC at scale factor ×4.0.
Figure A16. Spectral curve comparison on WADC at scale factor ×4.0.
Remotesensing 18 03187 g0a16

References

  1. Bioucas-Dias, J.M.; Plaza, A.; Dobigeon, N.; Parente, M.; Du, Q.; Gader, P.; Chanussot, J. Hyperspectral unmixing overview: Geometrical, statistical, and sparse regression-based approaches. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2012, 5, 354–379. [Google Scholar] [CrossRef] [Scilit]
  2. Loncan, L.; Almeida, L.B.; Bioucas-Dias, J.M.; Briottet, X.; Chanussot, J.; Dobigeon, N.; Fabre, S.; Liao, W.; Licciardi, G.A.; Simões, M.; et al. Hyperspectral pansharpening: A review. IEEE Geosci. Remote Sens. Mag. 2015, 3, 27–46. [Google Scholar] [CrossRef] [Scilit]
  3. Yokoya, N.; Grohnfeldt, C.; Chanussot, J. Hyperspectral and multispectral data fusion: A comparative review of the recent literature. IEEE Geosci. Remote Sens. Mag. 2017, 5, 29–56. [Google Scholar] [CrossRef] [Scilit]
  4. Vivone, G. Multispectral and hyperspectral image fusion in remote sensing: A survey. Inf. Fusion 2023, 89, 405–417. [Google Scholar] [CrossRef] [Scilit]
  5. Yokoya, N.; Yairi, T.; Iwasaki, A. Coupled nonnegative matrix factorization unmixing for hyperspectral and multispectral data fusion. IEEE Trans. Geosci. Remote Sens. 2012, 50, 528–537. [Google Scholar] [CrossRef] [Scilit]
  6. Bendoumi, M.A.; He, M.; Mei, S. Hyperspectral Image Resolution Enhancement Using High-Resolution Multispectral Image Based on Spectral Unmixing. IEEE Trans. Geosci. Remote Sens. 2014, 52, 6574–6583. [Google Scholar] [CrossRef] [Scilit]
  7. Simões, M.; Bioucas-Dias, J.; Almeida, L.B.; Chanussot, J. A convex formulation for hyperspectral image superresolution via subspace-based regularization. IEEE Trans. Geosci. Remote Sens. 2015, 53, 3373–3388. [Google Scholar] [CrossRef] [Scilit]
  8. Aiazzi, B.; Baronti, S.; Selva, M. Improving component substitution pansharpening through multivariate regression of MS + Pan data. IEEE Trans. Geosci. Remote Sens. 2007, 45, 3230–3239. [Google Scholar] [CrossRef] [Scilit]
  9. Lanaras, C.; Baltsavias, E.; Schindler, K. Hyperspectral super-resolution by coupled spectral unmixing. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; IEEE: New York, NY, USA, 2015; pp. 3586–3594. [Google Scholar] [CrossRef] [Scilit]
  10. Wei, Q.; Bioucas-Dias, J.; Dobigeon, N.; Tourneret, J.-Y. Hyperspectral and multispectral image fusion based on a sparse representation. IEEE Trans. Geosci. Remote Sens. 2015, 53, 3658–3668. [Google Scholar] [CrossRef] [Scilit]
  11. Wei, Q.; Dobigeon, N.; Tourneret, J.-Y. Fast fusion of multi-band images based on solving a Sylvester equation. IEEE Trans. Image Process. 2015, 24, 4109–4121. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Dian, R.; Li, S.; Kang, X. Hyperspectral image super-resolution: A coupled tensor factorization approach. IEEE Trans. Signal Process. 2018, 66, 6503–6517. [Google Scholar] [CrossRef] [Scilit]
  13. Zhou, Y.; Feng, L.; Hou, C.; Kung, S.Y. Hyperspectral and multispectral image fusion based on local low rank and coupled spectral unmixing. IEEE Trans. Geosci. Remote Sens. 2017, 55, 5997–6009. [Google Scholar] [CrossRef]
  14. Palsson, F.; Sveinsson, J.R.; Ulfarsson, M.O. Multispectral and hyperspectral image fusion using a 3-D-convolutional neural network. IEEE Geosci. Remote Sens. Lett. 2017, 14, 639–643. [Google Scholar] [CrossRef] [Scilit]
  15. Yang, J.; Zhao, Y.-Q.; Chan, J.C.-W. Hyperspectral and multispectral image fusion via deep two-branches convolutional neural network. Remote Sens. 2018, 10, 800. [Google Scholar] [CrossRef] [Scilit]
  16. Lanaras, C.; Bioucas-Dias, J.M.; Galliani, S.; Baltsavias, E.; Schindler, K. Super-resolution of Sentinel-2 images: Learning a globally applicable deep neural network. ISPRS J. Photogramm. Remote Sens. 2018, 146, 305–319. [Google Scholar] [CrossRef] [Scilit]
  17. Xie, Q.; Zhou, M.; Zhao, Q.; Meng, D.; Zuo, W.; Xu, Z. Multispectral and hyperspectral image fusion by MS/HS Fusion Net. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; IEEE: New York, NY, USA, 2019; pp. 1585–1594. [Google Scholar] [CrossRef] [Scilit]
  18. Xie, Q.; Zhou, M.; Zhao, Q.; Xu, Z.; Meng, D. MHF-Net: An interpretable deep network for multispectral and hyperspectral image fusion. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 1457–1473. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Jia, S.; Min, Z.; Fu, X. Multiscale spatial-spectral transformer network for hyperspectral and multispectral image fusion. Inf. Fusion 2023, 96, 117–129. [Google Scholar] [CrossRef] [Scilit]
  20. Ma, Q.; Jiang, J.; Liu, X.; Ma, J. Reciprocal transformer for hyperspectral and multispectral image fusion. Inf. Fusion 2024, 104, 102148. [Google Scholar] [CrossRef] [Scilit]
  21. Wang, X.; Wang, X.; Song, R.; Zhao, X.; Zhao, K. MCT-Net: Multi-hierarchical cross transformer for hyperspectral and multispectral image fusion. Knowl.-Based Syst. 2023, 264, 110362. [Google Scholar] [CrossRef] [Scilit]
  22. Zhang, X.; Huang, W.; Wang, Q.; Li, X. SSR-NET: Spatial-spectral reconstruction network for hyperspectral and multispectral image fusion. IEEE Trans. Geosci. Remote Sens. 2021, 59, 5953–5965. [Google Scholar] [CrossRef] [Scilit]
  23. Li, J.; Cui, R.; Li, B.; Song, R.; Li, Y.; Dai, Y.; Du, Q. Hyperspectral image super-resolution by band attention through adversarial learning. IEEE Trans. Geosci. Remote Sens. 2020, 58, 4304–4318. [Google Scholar] [CrossRef] [Scilit]
  24. Hu, J.-F.; Huang, T.-Z.; Deng, L.-J.; Jiang, T.-X.; Vivone, G.; Chanussot, J. Hyperspectral image super-resolution via deep spatiospectral attention convolutional neural networks. IEEE Trans. Neural Netw. Learn. Syst. 2022, 33, 7251–7265. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Hu, J.-F.; Huang, T.-Z.; Deng, L.-J.; Dou, H.-X.; Hong, D.; Vivone, G. Fusformer: A transformer-based fusion network for hyperspectral image super-resolution. IEEE Geosci. Remote Sens. Lett. 2022, 19, 1–5. [Google Scholar] [CrossRef] [Scilit]
  26. Li, J.; Zheng, K.; Yao, J.; Gao, L.; Hong, D. Deep unsupervised blind hyperspectral and multispectral data fusion. IEEE Geosci. Remote Sens. Lett. 2022, 19, 1–5. [Google Scholar] [CrossRef] [Scilit]
  27. Qu, Y.; Qi, H.; Kwan, C.; Yokoya, N.; Chanussot, J. Unsupervised and unregistered hyperspectral image super-resolution with mutual Dirichlet-Net. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–18. [Google Scholar] [CrossRef] [Scilit]
  28. Li, J.; Zheng, K.; Liu, W.; Li, Z.; Yu, H.; Ni, L. Model-guided coarse-to-fine fusion network for unsupervised hyperspectral image super-resolution. IEEE Geosci. Remote Sens. Lett. 2023, 20, 1–5. [Google Scholar] [CrossRef] [Scilit]
  29. Hong, D.; Yao, J.; Li, C.; Meng, D.; Yokoya, N.; Chanussot, J. Decoupled-and-coupled networks: Self-supervised hyperspectral image super-resolution with subpixel fusion. IEEE Trans. Geosci. Remote Sens. 2023, 61, 1–12. [Google Scholar] [CrossRef] [Scilit]
  30. Wang, T.; Yan, Z.; Li, J.; Zhao, X.; Wang, C.; Ng, M.K.-P. Hyperspectral and multispectral image fusion with arbitrary resolution through self-supervised representations. Int. J. Comput. Vis. 2025, 133, 7515–7535. [Google Scholar] [CrossRef] [Scilit]
  31. Liang, Y.-J.; Cao, Z.; Deng, S.; Dou, H.-X.; Deng, L.-J. Fourier-enhanced implicit neural fusion network for multispectral and hyperspectral image fusion. In Proceedings of the Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
  32. Jiang, Z.; Chen, M.; Wang, W. UMMFF: Unsupervised multimodal multilevel feature fusion network for hyperspectral image super-resolution. Remote Sens. 2024, 16, 3282. [Google Scholar] [CrossRef] [Scilit]
  33. Su, Y.; Li, S.; Zhou, Y.; Gao, L.; Jiang, M.; Sun, X.; Li, H.; Hou, E. Dilated Transformation-Guided Unsupervised Multimodal Learning for Hyperspectral and Multispectral Image Fusion. IEEE Trans. Geosci. Remote Sens. 2025, 63, 1–15. [Google Scholar] [CrossRef] [Scilit]
  34. Liang, F.; Ding, Y.; Zhang, Z.; Cai, Y.; Feng, J.; Liang, L.; Cheng, S. DSGC: Dynamic Sparse Graph Constrained Autoencoder for Hyperspectral Image Clustering. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5519012. [Google Scholar] [CrossRef] [Scilit]
  35. Ding, Y.; Zhang, Z.; Yang, A.; Cai, Y.; Xiao, X.; Hong, D.; Yuan, J. SLCGC: A lightweight Self-supervised Low-Pass Contrastive Graph Clustering Network for Hyperspectral Images. IEEE Trans. Multimed. 2025, 27, 8251–8262. [Google Scholar] [CrossRef] [Scilit]
  36. Ding, Y.; Zhang, Z.; Kang, W.; Yang, A.; Zhao, J.; Feng, J.; Hong, D.; Zheng, Q. Adaptive homophily clustering: Structure homophily graph learning with adaptive filter for hyperspectral image. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5510113. [Google Scholar] [CrossRef] [Scilit]
  37. Dong, R.; Xue, J.; Zhao, Y.; Wu, T.; Liu, Y.; Chan, J.C.-W. Robust geo-sparse and graph-fusion tensor learning for incomplete multi-view clustering. Inf. Fusion 2027, 138, 104716. [Google Scholar] [CrossRef] [Scilit]
  38. Gamba, P. A collection of data for urban area characterization. In Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Anchorage, AK, USA, 20–24 September 2004; IEEE: New York, NY, USA, 2004; Volume 1, pp. 69–72. [Google Scholar] [CrossRef] [Scilit]
  39. Landgrebe, D.A. Signal Theory Methods in Multispectral Remote Sensing; John Wiley & Sons: Hoboken, NJ, USA, 2003. [Google Scholar]
  40. Xu, Y.; Du, B.; Zhang, L.; Cerra, D.; Pato, M.; Carmona, E.; Prasad, S.; Yokoya, N.; Hänsch, R.; Le Saux, B. Advanced multi-sensor optical remote sensing for urban land use and land cover classification: Outcome of the 2018 IEEE GRSS Data Fusion Contest. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2019, 12, 1709–1724. [Google Scholar] [CrossRef] [Scilit]
  41. Yokoya, N.; Iwasaki, A. Airborne Hyperspectral Data over Chikusei; Technical Report SAL-2016-05-27; Space Application Laboratory, University of Tokyo: Tokyo, Japan, 2016. [Google Scholar]
  42. Wald, L.; Ranchin, T.; Mangolini, M. Fusion of satellite images of different spatial resolutions: Assessing the quality of resulting images. Photogramm. Eng. Remote Sens. 1997, 63, 691–699. [Google Scholar]
  43. Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
Figure 1. Overall training framework of the proposed SCDF-Net.In this figure, solid arrows represent data forward propagation, dashed arrows denote loss calculation paths; dashed boxes indicate different training stages. Different colors are used to distinguish various spectral curves and feature maps.
Figure 1. Overall training framework of the proposed SCDF-Net.In this figure, solid arrows represent data forward propagation, dashed arrows denote loss calculation paths; dashed boxes indicate different training stages. Different colors are used to distinguish various spectral curves and feature maps.
Remotesensing 18 03187 g001
Figure 2. Network architecture of SCDF-Net.Solid arrows represent data forward propagation, and different colored modules denote distinct functional sub-networks. The purple dashed box marks the scale-conditioned dual-domain feature learning module.
Figure 2. Network architecture of SCDF-Net.Solid arrows represent data forward propagation, and different colored modules denote distinct functional sub-networks. The purple dashed box marks the scale-conditioned dual-domain feature learning module.
Remotesensing 18 03187 g002
Figure 3. Scale-conditioned frequency-domain branch. Solid arrows represent data forward propagation. Modules with different background colors denote distinct functional blocks. The blue dashed box marks the magnitude descriptor calculation module, and the orange solid box indicates the complex Fourier fusion module.
Figure 3. Scale-conditioned frequency-domain branch. Solid arrows represent data forward propagation. Modules with different background colors denote distinct functional blocks. The blue dashed box marks the magnitude descriptor calculation module, and the orange solid box indicates the complex Fourier fusion module.
Remotesensing 18 03187 g003
Figure 4. Scale-conditioned spatial-domain branch. The green dashed box represents the feature extraction module; colored backgrounds distinguish different functional submodules. Arrows indicate the data flow.
Figure 4. Scale-conditioned spatial-domain branch. The green dashed box represents the feature extraction module; colored backgrounds distinguish different functional submodules. Arrows indicate the data flow.
Remotesensing 18 03187 g004
Figure 5. Visual comparison of different fusion methods on PaviaU at scale factor ×2.4.
Figure 5. Visual comparison of different fusion methods on PaviaU at scale factor ×2.4.
Remotesensing 18 03187 g005
Figure 6. Visual comparison of different fusion methods on Chikusei at scale factor ×2.4.
Figure 6. Visual comparison of different fusion methods on Chikusei at scale factor ×2.4.
Remotesensing 18 03187 g006
Figure 7. Visual comparison of different fusion methods on Houston18 at scale factor ×2.4.
Figure 7. Visual comparison of different fusion methods on Houston18 at scale factor ×2.4.
Remotesensing 18 03187 g007
Figure 8. Visual comparison of different fusion methods on WADC at scale factor ×2.4.
Figure 8. Visual comparison of different fusion methods on WADC at scale factor ×2.4.
Remotesensing 18 03187 g008
Figure 9. Spectral curve comparison on Chikusei at scale factor ×2.4.
Figure 9. Spectral curve comparison on Chikusei at scale factor ×2.4.
Remotesensing 18 03187 g009
Figure 10. Spectral curve comparison on Houston18 at scale factor ×2.4.
Figure 10. Spectral curve comparison on Houston18 at scale factor ×2.4.
Remotesensing 18 03187 g010
Figure 11. Spectral curve comparison on PaviaU at scale factor ×2.4.
Figure 11. Spectral curve comparison on PaviaU at scale factor ×2.4.
Remotesensing 18 03187 g011
Figure 12. Spectral curve comparison on WADC at scale factor ×2.4.
Figure 12. Spectral curve comparison on WADC at scale factor ×2.4.
Remotesensing 18 03187 g012
Table 1. Efficiency comparison of different deep learning-based methods on the same hardware.
Table 1. Efficiency comparison of different deep learning-based methods on the same hardware.
MethodParameters (M)Inference Time (s)
UDALN0.00230.0004
UMC2FF0.00560.0029
u2MDN0.00660.1383
CLoRF1.420.0075
SCDF-Net3.10.0071
Table 11. Quantitative comparison on Houston18 and Chikusei at scale factor ×4.0.
Table 11. Quantitative comparison on Houston18 and Chikusei at scale factor ×4.0.
MethodsHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
GSA31.49516.48120.02662.35640.983732.68373.27980.02323.28090.9438
HySure35.60074.18210.01666.57660.877937.71212.20380.01307.19660.7783
CNMF41.33881.71340.00862.02440.986640.41281.95990.00952.89490.9532
CLoRF40.28802.35920.00971.27930.996045.56031.61300.00533.41150.9745
UDALN40.95141.88410.00901.60410.995238.84043.43840.011410.13090.9575
u2MDN40.95981.82390.00901.55560.995746.57901.42640.00472.63130.9840
UMC2FF40.49492.44510.00940.91880.994846.33631.51930.00483.95380.9689
Proposed41.35501.69570.00861.55550.995747.65001.38700.00412.82900.9813
Note: Bold and underlined values denote the best and second-best results, respectively.
Table 12. Ablation results on PaviaU and WADC at scale factor ×2.0.
Table 12. Ablation results on PaviaU and WADC at scale factor ×2.0.
VariantPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
W/O FFT and spatial32.49424.22340.02376.82390.970142.55483.47930.00755.40020.8913
W/O FFT39.38752.95100.01073.85420.990444.91173.05550.00574.15080.9052
W/O spatial39.46292.94760.01063.85710.990644.97123.09880.00564.15730.9070
W/O scale-embedding39.05582.97460.01114.23120.988444.09683.17570.00624.63030.8848
W/O LFREQ38.96823.01390.01133.99700.989844.78103.05890.00584.20930.9002
W/O LSAM38.53423.11370.01184.15290.989044.27643.16030.00614.47540.8886
FULL42.07872.47650.00793.52090.992148.95871.73070.00362.47970.9732
Table 13. Ablation results on Houston18 and Chikusei at scale factor ×2.0.
Table 13. Ablation results on Houston18 and Chikusei at scale factor ×2.0.
VariantHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
W/O FFT and spatial39.31331.99730.01083.39270.994442.64451.92700.00747.20280.9676
W/O FFT42.36711.68070.00762.51510.997045.56221.62860.00535.93580.9792
W/O spatial42.48081.69490.00752.47520.997146.26331.61180.00494.90840.9804
W/O scale-embedding42.00221.73620.00792.61130.996844.98251.68410.00566.63270.9760
W/O LFREQ42.69981.62980.00732.42080.997345.25391.69300.00556.15240.9781
W/O LSAM41.88541.77350.00802.64970.996744.61961.78140.00596.42500.9761
FULL45.24811.30930.00551.82890.998950.14001.12150.00314.98870.9851
Table 14. Ablation results on PaviaU and WADC at scale factor ×2.4.
Table 14. Ablation results on PaviaU and WADC at scale factor ×2.4.
VariantPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
W/O FFT and spatial31.90294.06760.02546.67000.958441.68763.92360.00825.15400.8815
W/O FFT38.79543.03250.01153.36860.989444.42273.35380.00602.75200.8868
W/O spatial38.76473.04390.01153.38020.989444.38443.46680.00603.80600.9058
W/O scale-embedding38.79473.02050.01153.53840.988344.06483.37140.00634.07220.8846
W/O LFREQ38.21163.19510.01233.55250.988344.58313.30790.00593.66520.9081
W/O LSAM38.22393.19210.01233.54860.988343.67203.49920.00664.06830.8867
FULL40.43732.72020.00952.06290.991247.56501.94840.00422.56640.9724
Table 15. Ablation results on Houston18 and Chikusei at scale factor ×2.4.
Table 15. Ablation results on Houston18 and Chikusei at scale factor ×2.4.
VariantHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
W/O FFT and spatial38.49171.88560.01193.29240.992541.31662.15190.00866.19970.9643
W/O FFT42.38101.68000.00762.36970.996345.80481.59340.00514.87520.9805
W/O spatial41.60031.73840.00832.27970.996646.55331.54850.00472.97460.9795
W/O scale-embedding41.78031.67420.00812.32880.996444.77561.77430.00583.31580.9760
W/O LFREQ41.77721.68050.00822.23400.996844.97461.73980.00565.15380.9782
W/O LSAM41.30711.80440.00862.35290.996445.03091.74730.00565.19040.9780
FULL43.48041.38110.00671.77080.997949.15001.14990.00354.17060.9850
Table 16. Ablation results on PaviaU and WADC at scale factor ×4.0.
Table 16. Ablation results on PaviaU and WADC at scale factor ×4.0.
VariantPaviaUWADC
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
W/O FFT and spatial29.04905.12780.03535.58110.913739.27865.55970.01094.13330.8213
W/O FFT37.74623.25780.01302.19350.987342.93484.36300.00713.73200.9052
W/O spatial37.66143.31220.01312.21030.987143.44784.07000.00672.58960.8939
W/O scale-embedding37.05593.45960.01402.34950.985543.09364.09800.00704.07220.8846
W/O LFREQ37.57153.32030.01322.23070.986943.30813.88850.00682.49650.8894
W/O LSAM37.20263.40980.01382.30550.986043.02883.95890.00712.67490.8815
FULL38.95113.01690.01131.64210.988945.41892.71390.00541.53190.9623
Table 17. Ablation results on Houston18 and Chikusei at scale factor ×4.0.
Table 17. Ablation results on Houston18 and Chikusei at scale factor ×4.0.
VariantHouston18Chikusei
PSNRSAMRMSEERGASCCPSNRSAMRMSEERGASCC
W/O FFT and spatial39.72461.86280.01033.03380.981142.63521.92770.00745.05790.9311
W/O FFT39.66462.05930.01041.70820.994445.29771.67440.00542.97880.9791
W/O spatial39.82392.02820.01021.67730.994645.41551.67730.00542.97460.9795
W/O scale-embedding39.64612.07710.01041.70500.994344.46031.80080.00605.21860.9774
W/O LFREQ40.08901.96850.00991.64800.995144.99651.74220.00563.09340.9782
W/O LSAM39.95502.01000.01011.66120.994944.62281.82220.00593.20100.9768
FULL41.35501.69570.00861.55550.995747.65001.38700.00412.82900.9813
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Tang, P.; Zheng, K.; Li, J.; Yu, H.; Sun, X. Unsupervised Scale-Conditioned Hyperspectral and Multispectral Image Fusion via a Frequency–Spatial Dual-Domain Network. Remote Sens. 2026, 18, 3187. https://doi.org/10.3390/rs18183187

AMA Style

Tang P, Zheng K, Li J, Yu H, Sun X. Unsupervised Scale-Conditioned Hyperspectral and Multispectral Image Fusion via a Frequency–Spatial Dual-Domain Network. Remote Sensing. 2026; 18(18):3187. https://doi.org/10.3390/rs18183187

Chicago/Turabian Style

Tang, Peng, Ke Zheng, Jiaxin Li, Haoyang Yu, and Xu Sun. 2026. "Unsupervised Scale-Conditioned Hyperspectral and Multispectral Image Fusion via a Frequency–Spatial Dual-Domain Network" Remote Sensing 18, no. 18: 3187. https://doi.org/10.3390/rs18183187

APA Style

Tang, P., Zheng, K., Li, J., Yu, H., & Sun, X. (2026). Unsupervised Scale-Conditioned Hyperspectral and Multispectral Image Fusion via a Frequency–Spatial Dual-Domain Network. Remote Sensing, 18(18), 3187. https://doi.org/10.3390/rs18183187

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop