The experiments cover enhancement quality on iSAID-dark, a fixed-detector object-detection evaluation on high-resolution scenes, supplementary results on five general low-light benchmarks, and controlled analyses of the DDIP components and design parameters.
4.1. Datasets and Experimental Setup
Datasets. The primary evaluation was conducted on the publicly available iSAID-dark dataset released by Yao et al. [
4]. This dataset was constructed from the iSAID aerial image benchmark by applying a physically motivated illumination degradation model to high-quality daytime images, producing varying exposure levels and spatially uneven lighting. The prepared split used by our enhancement experiments contained 3755 paired training images and 66 paired validation images at a spatial resolution of
pixels. The scenes contain complex terrain textures and dense multi-scale objects, making illumination correction and detail preservation challenging.
General Benchmarks. Supplementary results are reported on five low-light datasets: LOL-v1 [
33], LOL-v2-Real [
34], LOL-v2-Syn [
34], LSRW-Nikon [
35], and LSRW-Huawei [
35] (training/test splits: 485/15, 689/100, 900/100, 3150/20, and 2450/30, respectively).
Comparison Baselines. The controlled results were the adjacent backbone/DDIP pairs, for which the baseline and DDIP-equipped variant were trained and evaluated under matched settings within each pair. FECNet, SNR-Net, FourLLIE, LLFormer, and ZERO-IG are literature-based contextual references reported with citations to their source methods, whereas DFFN was evaluated using the authors’ publicly released iSAID checkpoint. These contextual rows were not treated as results reproduced under the matched host-backbone protocols. Accordingly, only the adjacent backbone/DDIP pairs and their explicitly reported “+” gains were interpreted as controlled comparisons in this study; no gain was computed against a contextual method row.
Training Configuration. For each controlled backbone/DDIP pair, the baseline and DDIP-equipped variant used matched data splits, preprocessing, optimization schedules, patch and batch settings, iteration budgets, and evaluation scripts. No official pre-trained weights were used for either member of a controlled pair. DDIP was integrated as an illumination-prior branch before the main restoration blocks. This description excludes the separately reported DFFN reference, for which the released checkpoint was used as stated below. Unless otherwise stated, the principal reported setup used PyTorch 1.13 [
36] and an NVIDIA RTX 4090 GPU. The learning rate was initialized at
and decreased to
using cosine annealing [
37]. During training, input images were randomly cropped to
patches. We employed the Adam optimizer [
38] with
and
. The principal iteration budget was
with a batch size of four. The Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) [
39] were used as standard evaluation metrics. In all tables, the upward arrow (↑) indicates that higher values are preferable, and the downward arrow (↓) indicates that lower values are preferable.
For the complementary perceptual and illumination-oriented analysis, we additionally report the Learned Perceptual Image Patch Similarity (LPIPS) [
40], the mean absolute error of the CIELAB lightness channel (Lab
MAE), and the signed mean
bias on the 66 paired images in the iSAID-dark validation set under the Standard setting. LPIPS was computed using the official AlexNet linear model. LPIPS and Lab
MAE compare the enhanced result with the paired normal-light reference, while the signed mean
bias indicates whether the output is globally darker or brighter than that reference; a bias magnitude closer to zero indicates closer global lightness. In addition to the controlled backbone/DDIP pairs, DFFN was evaluated using the authors’ publicly released iSAID checkpoint as an external architecture-level reference. These metrics were computed after inference and were not used as training losses.
4.2. Low-Light Remote Sensing Image Enhancement
Experiment Settings. DDIP was evaluated with CNN-, Transformer-, and Mamba-based backbones on iSAID-dark [
4] under two settings. In the Standard setting, models were trained and tested on
images. In the High-resolution setting, the models trained at
were applied directly to full-resolution
aerial images without fine-tuning. This cross-resolution setting tested whether illumination correction and fine-grained geospatial detail could be retained at a substantially higher input resolution.
Analysis of Results.
Table 1 reports the quantitative results on iSAID-dark. Each DDIP-equipped backbone improves upon its corresponding baseline in the reported PSNR and SSIM comparisons. In the
Standard setting, Retinexformer+DDIP achieves the highest PSNR of 26.04 dB, while SPJFNet [
23] + DDIP improves by 2.0 dB over its base network.
Complementary Illumination and Perceptual Analysis.
Table 2 provides complementary evidence beyond the primary PSNR/SSIM evaluation. With RetinexMamba, DDIP reduces Lab
MAE from 4.6589 to 4.2976 and the absolute mean
bias from 2.1474 to 1.3362, while LPIPS increases from 0.2222 to 0.2261. With MIRNet-v2, DDIP reduces LPIPS from 0.2223 to 0.2143, Lab
MAE from 4.6245 to 4.2304, and the absolute mean
bias from 1.2311 to 1.1555. With Restormer, DDIP reduces Lab
MAE by 12.47% and the absolute mean
bias by 22.22%, while LPIPS remains nearly unchanged. With Retinexformer, DDIP reduces LPIPS by 4.98%, but Lab
MAE increases by 3.92% and the absolute mean
bias changes from 0.9762 to 1.2401. Thus, changes in the complementary metrics vary across host architectures, and no single perceptual or illumination metric is used to claim uniform superiority. DFFN is included as an external architecture-level reference rather than a paired DDIP ablation. All values were computed after inference using saved outputs and did not enter the training objectives.
DFFN is a representative dual-domain architecture for low-light remote sensing image enhancement. In the Standard setting of
Table 1, the PSNR values of Retinexformer+DDIP (26.04 dB) and MIRNet-v2+DDIP (25.94 dB) exceed the contextual DFFN result (25.30 dB) by 0.74 dB and 0.60 dB, respectively. Because DFFN was evaluated as a separate end-to-end architecture, these differences are not treated as paired DDIP gains.
The qualitative comparisons in
Figure 8 and
Figure 9 further illustrate the reported results on the iSAID-dark dataset under the Standard and High-resolution settings. The baseline SPJFNet [
23] shows a performance decline when applied to the high-resolution inputs (PSNR decreasing from 21.50 dB to 18.55 dB). In this setting, SPJFNet+DDIP’s PSNR is 1.69 dB higher than that of SPJFNet. These results suggest that dual-domain illumination calibration merits further study for remote-sensing imagery at varying resolutions. The gains reported for the tested CNN-, Transformer-, and Mamba-based backbones are specific to iSAID-dark and the stated experimental configuration. DDIP is jointly optimized with each host backbone; the frequency-domain branch is intended to calibrate global brightness, while the spatial-domain branch is intended to address localized illumination variation.
4.3. Downstream Object Detection Evaluation
Experiment Settings. To assess whether the tested enhancement outputs benefited a downstream remote-sensing task, we trained a YOLOv8n detector once on normal-illumination iSAID images derived from DOTA v1.0, using the corresponding official iSAID instance annotations. The detector was then fixed; no detector retraining was performed for any low-light or enhanced input. The test set comprised 24 high-resolution iSAID-dark validation scenes, partitioned into 144 detector tiles. For each scene, we evaluated the raw synthetic low-light images at three illumination levels (Low-15, Low-20, and Low-30), together with the outputs of RetinexMamba, RetinexMamba+DDIP, MIRNet-v2, MIRNet-v2+DDIP, Restormer, Restormer+DDIP, Retinexformer, Retinexformer+DDIP, and DFFN. All variants used the same image tiles and annotations. DFFN used the authors’ publicly released iSAID checkpoint and was included as an external architecture-level reference; it was not treated as a controlled DDIP pair. We report COCO-style
(denoted as mAP),
, and small-object AP. These checkpoints are used here as application-oriented comparisons of enhancement pipelines; this experiment does not replace the controlled backbone-level ablations in
Section 4.5.
Analysis of Results. The fixed-detector evaluation in
Table 3 shows that the tested enhancement pipelines substantially improve detection performance relative to the raw synthetic low-light inputs under the evaluated conditions. Relative to the corresponding host backbones, DDIP increases mean mAP by 0.48, 0.46, and approximately 0.08 percentage points for MIRNet-v2, Restormer, and RetinexMamba, respectively; the last value was calculated from the unrounded results. For Retinexformer, the mean mAP changes only marginally, whereas its mean
and mean
decrease. For RetinexMamba, the mean
increases from 14.83% to 15.41%, and the largest gain occurs under the Low-20 condition; its mean
decreases slightly from 45.27% to 45.12%. The improvement is therefore more evident for the MIRNet-v2 and Restormer hosts and is not uniform across all backbones, degradation levels, or metrics. In particular, the present experiment does not support the claim that DDIP consistently improves small-object AP. The official DFFN checkpoint provides an external architecture-level reference of 29.75% mean mAP and 45.60% mean
; this row is not used to infer a DDIP gain.
4.4. Supplementary Evaluation on Standard Low-Light Benchmarks
We additionally evaluated DDIP on five standard low-light benchmarks. While the primary contribution targets remote sensing imagery (
Section 4.2), these experiments assess DDIP in the reported non-remote-sensing settings and do not constitute additional remote-sensing validation.
Experiment Settings. MIRNet-v2 [
11], Restormer [
13], Retinexformer [
12], and EvLight [
41] served as base networks, with DDIP inserted before each for controlled backbone-versus-backbone+DDIP comparisons (see
Figure 7). Additional baseline methods are listed in
Section 4.1. A cross-dataset evaluation was performed between LOL-v2-Real [
34] and LSRW-Huawei [
35]: models trained on one dataset were tested on the other under the specified train–test pairing.
Analysis of Results.
Table 4 and
Table 5 report the quantitative results on the LOL-v1, LOL-v2-Real, LOL-v2-Syn, LSRW-Nikon, and LSRW-Huawei datasets. Across these benchmarks, equipping the reported base networks with DDIP yields positive PSNR and SSIM differences. The largest PSNR increases on the five datasets are 0.94 dB, 0.61 dB, 1.34 dB, 0.41 dB, and 0.38 dB, respectively. The qualitative comparison in
Figure 10 suggests that DDIP contributes to correcting the illumination distribution of the base networks while preserving spatial structures. These results are limited to the evaluated datasets and training settings.
Table 6 reports the cross-dataset results. FECNet, FourLLIE, and LLFormer are included as reference baselines; because their architectures are not used as modular host backbones here, DDIP is integrated into MIRNet-v2, Restormer, and Retinexformer for the paired comparisons. In this setting, the DDIP-equipped backbones retain positive PSNR and SSIM gains over their corresponding baselines. This finding is restricted to the two reported train–test dataset pairs.
4.5. Ablation Study
To examine the contribution of each component, we conducted controlled backbone-level ablations on the iSAID-dark [
4] dataset with MIRNet-v2 under the Standard setting.
Table 7 evaluates FIDP, SIDP, and SCFF within the full enhancement pipeline.
Table 8 provides a separate standalone fusion analysis on LOL-v1 [
33]; because it uses a different dataset and evaluation configuration, it is not combined with
Table 7 to estimate the total DDIP overhead. The ablation focused on controlled component and fusion comparisons. We did not extract modules from other complete architectures for module-only comparisons because doing so would not have reproduced their original end-to-end systems.
Component Ablation. The “Only FIDP” and “Only SIDP” configurations isolate the frequency-domain and spatial-domain priors, respectively, while “No SCFF” retains both prior branches without the selective fusion module.
Table 7 shows that the full DDIP configuration performs better than either single-domain configuration and the version without SCFF. It achieves a PSNR of 25.94 dB and SSIM of 0.8570, compared with 25.50 dB and 0.7913 for the baseline. Within this controlled setting, the results support the joint use of the two illumination priors and selective cross-domain fusion.
SIDP Scale and Order Sensitivity. The default sequence
combines medium-scale (
), fine-scale (
), and global calibration so that the spatial branch can model heterogeneous local illumination together with scene-wide brightness. In the retraining-controlled rows of
Table 9, removing the fine-scale level changes PSNR by
dB and SSIM by
. Thus, after retraining, the tested scale-count change produces only a small difference. The fixed-checkpoint rows examine processing order while preserving the same learned parameters. The default, local-order-reversed, and global-first sequences obtain 25.9471/0.8522, 25.9481/0.8522, and 25.9436/0.8522, respectively. The 0.0045 dB PSNR span and unchanged four-decimal SSIM show limited order sensitivity for the tested images. This fixed-checkpoint result is interpreted as an inference-time diagnostic rather than an independently trained architectural ablation, and values from the two protocols are not compared directly. The three-level formulation remains the hierarchical default rather than being presented as uniquely optimal; the preferred spatial partition may depend on image resolution and illumination degradation.
Amplitude-Loss-Weight Sensitivity.
Table 10 evaluates
. The maximum PSNR variation is 0.0986 dB. The setting
obtains the highest observed PSNR, whereas
obtains the highest observed SSIM. These small differences indicate that DDIP is not highly sensitive to the tested range of amplitude-loss weights; no statistical-significance claim is made from these single runs. The reported main results and the scale comparison retain
to preserve the original training protocol.
SCFF Ablation. We conducted a component-level comparison on LOL-v1 [
33] between SCFF and three alternative aggregation strategies: Sum, Concat, and Selective Kernel Feature Fusion (SKFF) [
10]. Only the fusion strategy was replaced, while the remaining standalone dual-domain analysis configuration was kept unchanged. Therefore, the FLOPs and parameter counts in
Table 8 should be interpreted as the total complexity of the standalone fusion-analysis configuration rather than the parameter size of each fusion block alone. The differences among rows reflect the relative cost introduced by each fusion strategy. Compared with SKFF, SCFF obtains a 0.12 dB higher PSNR and 0.003 higher SSIM under the reported configuration, with the same rounded FLOPs.
SCFF Parameter Sensitivity.
Table 11 reports a fixed-checkpoint inference diagnostic. Varying
from 1.0 to 4.0 leaves the measured PSNR and SSIM unchanged at the reported precision. For the fidelity coefficient, the default
gives the highest PSNR among the tested values. Setting
disables the explicit input-fidelity path and reduces PSNR by 0.0899 dB, whereas increasing
to 0.6 reduces both PSNR and SSIM. The small SSIM increase at
reflects a PSNR–SSIM trade-off rather than a uniform gain. This fixed-checkpoint analysis characterizes local sensitivity around the selected configuration and is not interpreted as a retraining-controlled architectural ablation.
Inference Overhead of DDIP.
Table 12 reports end-to-end inference latency and throughput for MIRNet-v2 with and without DDIP at two representative resolutions. At
, the reported DDIP configuration adds 14.6 ms of latency (a 5.6% relative increase), with FPS changing from 3.82 to 3.62. At
, the reported absolute latency difference is 126.0 ms, with FPS changing from 0.48 to 0.45. Static inspection of the tested implementation further shows that DDIP adds 502 parameters (from 5,858,560 to 5,859,062 parameters; 0.009% of the host parameter count). The numbers of incremental parameters are 81 in FIDP, 36 in SIDP, 65 in SCFF, and 320 in the prior-projection layer. These measurements describe the tested implementation, hardware, and host backbone; they do not provide module-level profiling or establish latency behavior for untested backbones or resolutions.