ACBDT: SAR-Optical Cross-Modal Distillation for Sentinel-1/2 Building-Footprint Mapping in Heterogeneous Yangtze River Delta Cities
Abstract
1. Introduction
2. Study Area and Dataset
2.1. Study Area Description

2.2. Satellite Data and Preprocessing
2.2.1. Sentinel-1 SAR Data
| City | Sentinel-1A IW SLC Products | Relative Orbit(s) | Sentinel-1 Date Range (2024) | Coherence Pairs | Sentinel-2 L2A Products | Sentinel-2 Date Range (2024) |
|---|---|---|---|---|---|---|
| Hangzhou | 10 | 69 | 7 June–23 September | 9 | 50 | 1 June–29 September |
| Hefei | 20 | 142 | 12 June–28 September | 9 | 203 | 2 June–30 September |
| Nanjing | 40 | 69, 142 | 7 June–28 September | 18 | 52 | 1 June–29 September |
| Ningbo | 42 | 69, 171 | 2 June–30 September | 19 | 98 | 1 June–29 September |
| Shanghai | 20 | 171 | 2 June–30 September | 10 | 49 | 1 June–29 September |
| Suzhou | 40 | 69, 171 | 2 June–30 September | 19 | 100 | 1 June–29 September |
| Wenzhou | 31 | 69, 171 | 2 June–30 September | 19 | 197 | 1 June–29 September |
| Wuxi | 28 | 69, 171 | 2 June–30 September | 15 | 100 | 1 June–29 September |
| Total | 231 | 69, 142, 171 | 2 June–30 September | 118 | 849 | 1 June–30 September |
| # | Channel | Definition | Physical Role |
|---|---|---|---|
| 1 | VV_dB | 10 log10(sigma0_VV), terrain-corrected, temporally averaged | Co-polarized backscatter; strong double-bounce from building-ground dihedrals |
| 2 | VH_dB | 10 log10(sigma0_VH), terrain-corrected, temporally averaged | Cross-polarized backscatter; sensitive to volume scattering and depolarizing targets |
| 3 | VV_Coh | Interferometric coherence magnitude, VV, 12-day pair, 5 × 5 window | Temporal stability of co-polarized scattering; high over rigid built structures |
| 4 | VH_Coh | Interferometric coherence magnitude, VH, 12-day pair, 5 × 5 window | Temporal stability of cross-polarized scattering; low over vegetation |
| 5 | H | -sum_i p_i log2(p_i), p_i = lambda_i/(lambda_1 + lambda_2 + eps), C2 over 5 × 5, eps = 1 × 10−15 | Entropy-like scattering randomness (dual-pol approximation, bounded by 1) |
| 6 | Alpha | sum_i p_i alpha_i, alpha_i from the C2 eigenvectors | Alpha-like mean scattering angle; separates surface, dipole, and dihedral regimes (dual-pol approximation) |
| 7 | Pd | (lambda_1 − lambda_2)/(lambda_1 + lambda_2 + eps) | Normalized eigenvalue contrast; dominance of a single scattering mechanism (anisotropy analog) |
2.2.2. Sentinel-2 Optical Data
2.2.3. Multi-Modal Data Co-Registration
2.3. Reference Data Preparation
| Category | N Footprints | Min (m2) | 25th Pct | Median | Mean | 75th Pct | Max (m2) |
|---|---|---|---|---|---|---|---|
| Residential | 132,640 | 100 | 128 | 176 | 264 | 292 | 6350 |
| Commercial/industrial | 18,920 | 102 | 410 | 920 | 1720 | 2140 | 48,600 |
| All retained buildings | 151,560 | 100 | 132 | 202 | 446 | 338 | 48,600 |
| Excluded below threshold | 24,310 | – | – | – | – | – | 99.9 |
| Record | Spatial Unit | N | Role | Reference Basis | Grid | Aggregation | Notes |
|---|---|---|---|---|---|---|---|
| Eight-city primary inventory | 256 × 256 tiles | 2680 | Training/validation/test | Corrected building polygons | 10 m | 2144/268/268 | Spatially disjoint block partition |
| Held-out primary test | 256 × 256 tiles | 268 | Reference benchmark + repeated-seed robustness | Same reference masks | 10 m | Tile-aggregated; dataset-level BF1 | Used by benchmark and repeated-seed analyses; not used for model selection |
| Recovered-checkpoint density audit | 256 × 256 tiles | 416 | Density diagnostic | Frozen audit masks | 10 m | Tile-aggregated | Seed-42 retained density audit only |
| Pixel-pooled diagnostic | Pixels | 136,511,488 | Commission/omission diagnosis | Same test-mask geometry | 10 m | Micro-averaged | Not directly comparable with tile-aggregated IoU |
3. Methods
3.1. Overview

3.2. Adaptive Teacher Generator
3.2.1. Asymmetric Fusion Network
3.2.2. Hybrid Attention Multi-Scale Fusion
3.2.3. Time-Step-Conditioned Teacher Network
3.3. Cross-Modal Distillation Bridge
3.4. Student Refinement Decoder
3.5. Loss Functions and Training Strategy

3.5.1. Teacher Supervision
3.5.2. Knowledge Distillation [35,36]
3.5.3. Student Supervision and Overall Objective
3.6. Training and Inference Protocol
3.7. Benchmark and Validation Protocol


4. Results and Analysis
| Record | Inventory | Partition | Aggregation | Reported Record | Headline Use | Evidential Role |
|---|---|---|---|---|---|---|
| Checkpoint-selection log | 2680-patch eight-city set | Validation | Tile-aggregated | OA 97.60; IoU 81.08; F1 89.55; Kappa 0.8819 | No | Checkpoint selection only |
| Pixel-pooled diagnostic | 2083-tile single-city archive | Full archive (136,511,488 px) | Pixel-pooled (micro) | OA 99.22; Precision 96.09; Recall 97.36 | No | Error diagnosis only; not comparable with tile metrics |
| Same-protocol benchmark | 2680-patch eight-city set | Held-out test | Tile-aggregated; dataset BF1 | OA 95.47; IoU 85.63; F1 92.24; Kappa 0.9017; BF1 83.71 | Reference | Broad architecture comparison; single-run values |
| Repeated-seed robustness and key ablations | 2680-patch eight-city set | Held-out test | Tile-aggregated; dataset BF1 | 3 runs each: ACBDT, FTransUNet, w/o CMDB, w/o time-step | Yes | Mean ± SD robustness evidence |
| ACBDT implementation (Section 4.5.3) | Archived ACBDT checkpoint | Inference record | Direct ACBDT profiling | 32.402 M; 33.80 ± 4.31 ms/tile; 11,633.6 km2/min | No | ACBDT only; no cross-model efficiency claim |
| Density diagnostic | Same 416-tile audit | Five density strata | Tile-aggregated by stratum | Full model > optical-only in all strata | No | Observed density dependence of SAR contribution |
4.1. Training Convergence and Evaluation Protocol
4.2. Pixel-Level and Tile-Level Error Structure
| Reference/Prediction | Predicted Background | Predicted Building |
|---|---|---|
| Reference background | 119,707,039 | 639,784 |
| Reference building | 427,314 | 15,737,351 |
4.3. Qualitative Assessment of Representative Validation Samples

4.4. City-Scale Building Extraction Across the Yangtze River Delta

4.5. Quantitative Evaluation: Benchmark Comparison and Observed Diagnostics
4.5.1. Overall Benchmark Comparison
| Method | Modality | OA (%) | IoU (%) | F1 (%) | Kappa | BF1 (%) | Params (M) |
|---|---|---|---|---|---|---|---|
| U-Net (OPT) | OPT | 92.14 | 74.83 | 85.62 | 0.8134 | 72.31 | 31.0 |
| U-Net (SAR) | SAR | 89.67 | 68.24 | 81.13 | 0.7589 | 65.47 | 31.0 |
| DeepLab v3+ (CAT) | OPT + SAR | 93.28 | 77.56 | 87.34 | 0.8367 | 74.82 | 59.3 |
| HRNet (FUSE) | OPT + SAR | 93.71 | 78.93 | 88.19 | 0.8478 | 76.14 | 65.8 |
| SegFormer (B3) | OPT + SAR | 94.12 | 80.47 | 89.21 | 0.8612 | 77.53 | 47.2 |
| CMX | OPT + SAR | 94.56 | 82.13 | 90.17 | 0.8743 | 79.68 | 73.4 |
| FTransUNet | OPT + SAR | 94.83 | 83.27 | 90.84 | 0.8821 | 81.23 | 88.6 |
| DiffusionSeg | OPT + SAR | 94.31 | 81.79 | 89.94 | 0.8698 | 78.94 | 94.2 |
| ACBDT (Ours) | OPT + SAR | 95.47 | 85.63 | 92.24 | 0.9017 | 83.71 | 32.402 |
4.5.2. Controlled Ablation and Repeated-Seed Analysis
| Configuration | Protocol | Runs | IoU (%) | ΔIoU (pp) | F1 (%) | Kappa | BF1 (%) |
|---|---|---|---|---|---|---|---|
| Full ACBDT | Held-out test | 3 | 85.61 ± 0.32 | — | 92.24 ± 0.19 | 0.9015 ± 0.0028 | 83.74 ± 0.34 |
| FTransUNet | Held-out test | 3 | 83.21 ± 0.24 | −2.40 | 90.82 ± 0.16 | 0.8818 ± 0.0025 | 81.19 ± 0.29 |
| w/o CMDB | Held-out test | 3 | 79.01 ± 0.42 | −6.60 | 88.27 ± 0.31 | 0.8546 ± 0.0037 | 77.63 ± 0.46 |
| w/o time-step conditioning | Held-out test | 3 | 84.53 ± 0.20 | −1.08 | 91.62 ± 0.14 | 0.8924 ± 0.0021 | 82.47 ± 0.26 |
| Full ACBDT † | 416-tile audit | 1 | 82.7 | — | 90.46 | 0.8738 | 81.35 |
| Optical only † | 416-tile audit | 1 | 80.61 | −2.09 | 89.18 | 0.8587 | 79.42 |
| SAR only † | 416-tile audit | 1 | 67.2 | −15.50 | 80.38 | 0.7316 | 64.73 |
| w/o HAMSF † | 416-tile audit | 1 | 81.88 | −0.82 | 89.94 | 0.8681 | 80.57 |
4.5.3. Computational Audit
4.6. Density-Dependent Performance and Boundary Behavior
| Building-Density Stratum | N Tiles | Full Model IoU (%) | Optical-Only IoU (%) | SAR-Channel Gain (pp) |
|---|---|---|---|---|
| Very sparse (<1%) | 4 | 74.95 | 67.49 | +7.47 |
| Sparse (1–5%) | 104 | 79.54 | 77.29 | +2.26 |
| Moderate (5–15%) | 180 | 82.90 | 81.16 | +1.73 |
| Dense (15–30%) | 120 | 83.76 | 81.59 | +2.18 |
| Very dense (>30%) | 8 | 80.81 | 77.85 | +2.96 |
5. Discussion
5.1. Modality Complementarity and the Role of Cross-Modal Alignment
5.2. Error Modes and Their Physical Interpretation
5.3. Limitations and Future Work
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Seto, K.C.; Güneralp, B.; Hutyra, L.R. Global forecasts of urban expansion to 2030 and direct impacts on biodiversity and carbon pools. Proc. Natl. Acad. Sci. USA 2012, 109, 16083–16088. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pesaresi, M.; Huadong, G.; Blaes, X.; Ehrlich, D.; Ferri, S.; Gueguen, L.; Halkia, M.; Kauffmann, M.; Kemper, T.; Lu, L.; et al. A Global Human Settlement Layer from Optical HR/VHR RS Data: Concept and First Results. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2013, 6, 2102–2131. [Google Scholar] [CrossRef] [Scilit]
- Gong, P.; Li, X.; Wang, J.; Bai, Y.; Chen, B.; Hu, T.; Liu, X.; Xu, B.; Yang, J.; Zhang, W.; et al. Annual maps of global artificial impervious area (GAIA) between 1985 and 2018. Remote Sens. Environ. 2020, 236, 111510. [Google Scholar] [CrossRef] [Scilit]
- Torres, R.; Snoeij, P.; Geudtner, D.; Bibby, D.; Davidson, M.; Attema, E.; Potin, P.; Rommen, B.; Floury, N.; Brown, M.; et al. GMES Sentinel-1 mission. Remote Sens. Environ. 2012, 120, 9–24. [Google Scholar] [CrossRef] [Scilit]
- Drusch, M.; Del Bello, U.; Carlier, S.; Colin, O.; Fernandez, V.; Gascon, F.; Hoersch, B.; Isola, C.; Laberinti, P.; Martimort, P.; et al. Sentinel-2: ESA’s Optical High-Resolution Mission for GMES Operational Services. Remote Sens. Environ. 2012, 120, 25–36. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.-S. Digital Image Enhancement and Noise Filtering by Use of Local Statistics. IEEE Trans. Pattern Anal. Mach. Intell. 1980, PAMI-2, 165–168. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zebker, H.A.; Villasenor, J. Decorrelation in interferometric radar echoes. IEEE Trans. Geosci. Remote Sens. 1992, 30, 950–959. [Google Scholar] [CrossRef] [Scilit]
- Cloude, S.R.; Pottier, E. An entropy based classification scheme for land applications of polarimetric SAR. IEEE Trans. Geosci. Remote Sens. 1997, 35, 68–78. [Google Scholar] [CrossRef] [Scilit]
- Khaleghi, B.; Khamis, A.; Karray, F.O.; Razavi, S.N. Multisensor data fusion: A review of the state-of-the-art. Inf. Fusion 2013, 14, 28–44. [Google Scholar] [CrossRef] [Scilit]
- Schmitt, M.; Zhu, X.X. Data Fusion and Remote Sensing: An ever-growing relationship. IEEE Geosci. Remote Sens. Mag. 2016, 4, 6–23. [Google Scholar] [CrossRef] [Scilit]
- Kussul, N.; Lavreniuk, M.; Skakun, S.; Shelestov, A. Deep Learning Classification of Land Cover and Crop Types Using Remote Sensing Data. IEEE Geosci. Remote Sens. Lett. 2017, 14, 778–782. [Google Scholar] [CrossRef] [Scilit]
- Inglada, J.; Vincent, A.; Arias, M.; Tardy, B.; Morin, D.; Rodes, I. Operational High Resolution Land Cover Map Production at the Country Scale Using Satellite Image Time Series. Remote Sens. 2017, 9, 95. [Google Scholar] [CrossRef] [Scilit]
- Ienco, D.; Interdonato, R.; Gaetano, R.; Ho Tong Minh, D. Combining Sentinel-1 and Sentinel-2 Satellite Image Time Series for land cover mapping via a multi-source deep learning architecture. ISPRS J. Photogramm. Remote Sens. 2019, 158, 11–22. [Google Scholar] [CrossRef] [Scilit]
- Hafner, S.; Ban, Y.; Nascetti, A. Unsupervised domain adaptation for global urban extraction using Sentinel-1 SAR and Sentinel-2 MSI data. Remote Sens. Environ. 2022, 280, 113192. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Matgen, P.; Chini, M. Extraction of built-up areas using Sentinel-1 and Sentinel-2 data with automated training data sampling and label noise robust cross-fusion neural networks. Int. J. Appl. Earth Obs. Geoinf. 2025, 139, 104524. [Google Scholar] [CrossRef] [Scilit]
- Quan, Y.; Zhang, R.; Li, J.; Ji, S.; Guo, H.; Yu, A. Learning SAR-Optical Cross Modal Features for Land Cover Classification. Remote Sens. 2024, 16, 431. [Google Scholar] [CrossRef] [Scilit]
- Zhou, W.; Jin, J.; Lei, J.; Yu, L. CIMFNet: Cross-Layer Interaction and Multiscale Fusion Network for Semantic Segmentation of High-Resolution Remote Sensing Images. IEEE J. Sel. Top. Signal Process. 2022, 16, 666–676. [Google Scholar] [CrossRef] [Scilit]
- Ayala, C.; Aranda, C.; Galar, M. Towards fine-grained road maps extraction using Sentinel-2 imagery. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2021, V-3-2021, 9–14. [Google Scholar] [CrossRef] [Scilit]
- Luo, Z.; Pan, J.; Hu, Y.; Deng, L.; Li, Y.; Qi, C.; Wang, X. RS-Dseg: Semantic segmentation of high-resolution remote sensing images based on a diffusion model component with unsupervised pretraining. Sci. Rep. 2024, 14, 18609. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 3431–3440. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2015; Volume 9351, pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 2481–2495. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Computer Vision—ECCV 2018; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2018; Volume 11211, pp. 833–851. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.X.; Tuia, D.; Mou, L.; Xia, G.-S.; Zhang, L.; Xu, F.; Fraundorfer, F. Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources. IEEE Geosci. Remote Sens. Mag. 2017, 5, 8–36. [Google Scholar] [CrossRef] [Scilit]
- Ma, L.; Liu, Y.; Zhang, X.; Ye, Y.; Yin, G.; Johnson, B.A. Deep learning in remote sensing applications: A meta-analysis and review. ISPRS J. Photogramm. Remote Sens. 2019, 152, 166–177. [Google Scholar] [CrossRef] [Scilit]
- Audebert, N.; Le Saux, B.; Lefèvre, S. Beyond RGB: Very high resolution urban remote sensing with multimodal deep networks. ISPRS J. Photogramm. Remote Sens. 2018, 140, 20–32. [Google Scholar] [CrossRef] [Scilit]
- Marmanis, D.; Schindler, K.; Wegner, J.D.; Galliani, S.; Datcu, M.; Stilla, U. Classification with an edge: Improving semantic image segmentation with boundary detection. ISPRS J. Photogramm. Remote Sens. 2018, 135, 158–172. [Google Scholar] [CrossRef] [Scilit]
- Maggiori, E.; Tarabalka, Y.; Charpiat, G.; Alliez, P. High-Resolution Aerial Image Labeling with Convolutional Neural Networks. IEEE Trans. Geosci. Remote Sens. 2017, 55, 7092–7103. [Google Scholar] [CrossRef] [Scilit]
- Volpi, M.; Tuia, D. Dense Semantic Labeling of Subdecimeter Resolution Images with Convolutional Neural Networks. IEEE Trans. Geosci. Remote Sens. 2017, 55, 881–893. [Google Scholar] [CrossRef] [Scilit]
- Kampffmeyer, M.; Salberg, A.-B.; Jenssen, R. Semantic Segmentation of Small Objects and Modeling of Uncertainty in Urban Remote Sensing Images Using Deep Convolutional Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Las Vegas, NV, USA, 26 June–1 July 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 680–688. [Google Scholar] [CrossRef] [Scilit]
- Cui, J.; Liu, J.; Ni, Y.; Sun, Y.; Guo, M. MCKTNet: Multiscale Cross-Modal Knowledge Transfer Network for Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4406015. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Qu, Y.; Zhang, L. Multispectral Scene Classification via Cross-Modal Knowledge Distillation. IEEE Trans. Geosci. Remote Sens. 2022, 60, 4406015. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Zhang, X.; Pun, M.-O.; Liu, M. A Multilevel Multimodal Fusion Transformer for Remote Sensing Semantic Segmentation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5403215. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Yue, J.; Xia, S.; Ghamisi, P.; Xie, W.; Fang, L. Diffusion Models Meet Remote Sensing: Principles, Methods, and Perspectives. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4708322. [Google Scholar] [CrossRef] [Scilit]
- Hinton, G.; Vinyals, O.; Dean, J. Distilling the Knowledge in a Neural Network. arXiv 2015, arXiv:1503.02531. [Google Scholar] [CrossRef] [Scilit]
- Romero, A.; Ballas, N.; Kahou, S.E.; Chassang, A.; Gatta, C.; Bengio, Y. FitNets: Hints for Thin Deep Nets. In Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015; Available online: https://arxiv.org/abs/1412.6550 (accessed on 20 June 2026).
- Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. In Proceedings of the Advances in Neural Information Processing Systems 33 (NeurIPS 2020), Virtual, 6–12 December 2020; pp. 6840–6851. Available online: https://proceedings.neurips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html (accessed on 20 June 2026).
- Song, Y.; Sohl-Dickstein, J.; Kingma, D.P.; Kumar, A.; Ermon, S.; Poole, B. Score-Based Generative Modeling through Stochastic Differential Equations. In Proceedings of the 9th International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021; Available online: https://openreview.net/forum?id=PxTIG12RRHS (accessed on 20 June 2026).
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 10674–10685. [Google Scholar] [CrossRef] [Scilit]
- Tucker, C.J. Red and photographic infrared linear combinations for monitoring vegetation. Remote Sens. Environ. 1979, 8, 127–150. [Google Scholar] [CrossRef] [Scilit]
- McFeeters, S.K. The use of the Normalized Difference Water Index (NDWI) in the delineation of open water features. Int. J. Remote Sens. 1996, 17, 1425–1432. [Google Scholar] [CrossRef] [Scilit]
- Zha, Y.; Gao, J.; Ni, S. Use of normalized difference built-up index in automatically mapping urban areas from TM imagery. Int. J. Remote Sens. 2003, 24, 583–594. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 7132–7141. [Google Scholar] [CrossRef] [Scilit]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Computer Vision—ECCV 2018; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2018; Volume 11211, pp. 3–19. [Google Scholar] [CrossRef] [Scilit]
- Neupane, B.; Horanont, T.; Aryal, J. Deep Learning-Based Semantic Segmentation of Urban Features in Satellite Images: A Review and Meta-Analysis. Remote Sens. 2021, 13, 808. [Google Scholar] [CrossRef] [Scilit]
- Chini, M.; Pelich, R.; Hostache, R.; Matgen, P.; Lopez-Martinez, C. Towards a 20 m Global Building Map from Sentinel-1 SAR Data. Remote Sens. 2018, 10, 1833. [Google Scholar] [CrossRef] [Scilit]
- Jacob, A.W.; Vicente-Guijalba, F.; Lopez-Martinez, C.; Lopez-Sanchez, J.M.; Litzinger, M.; Kristen, H.; Mestre-Quereda, A.; Ziolkowski, D.; Lavalle, M.; Notarnicola, C.; et al. Sentinel-1 InSAR Coherence for Land Cover Mapping: A Comparison of Multiple Feature-Based Classifiers. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 13, 535–552. [Google Scholar] [CrossRef] [Scilit]
- Ji, K.; Wu, Y. Scattering Mechanism Extraction by a Modified Cloude-Pottier Decomposition for Dual Polarization SAR. Remote Sens. 2015, 7, 7447–7470. [Google Scholar] [CrossRef] [Scilit]
- Ainsworth, T.L.; Kelly, J.P.; Lee, J.-S. Classification comparisons between dual-pol, compact polarimetric and quad-pol SAR imagery. ISPRS J. Photogramm. Remote Sens. 2009, 64, 464–471. [Google Scholar] [CrossRef] [Scilit]
- Main-Knorn, M.; Pflug, B.; Louis, J.; Debaecker, V.; Müller-Wilm, U.; Gascon, F. Sen2Cor for Sentinel-2. In Proceedings of the Image and Signal Processing for Remote Sensing XXIII, Warsaw, Poland, 11–13 September 2017; SPIE: Bellingham, WA, USA, 2017; Volume 10427, p. 1042704. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Z.; Woodcock, C.E. Object-based cloud and cloud shadow detection in Landsat imagery. Remote Sens. Environ. 2012, 118, 83–94. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Z.; Wang, S.; Woodcock, C.E. Improvement and expansion of the Fmask algorithm: Cloud, cloud shadow, and snow detection for Landsats 4–7, 8, and Sentinel 2 images. Remote Sens. Environ. 2015, 159, 269–277. [Google Scholar] [CrossRef] [Scilit]
- Gorelick, N.; Hancher, M.; Dixon, M.; Ilyushchenko, S.; Thau, D.; Moore, R. Google Earth Engine: Planetary-scale geospatial analysis for everyone. Remote Sens. Environ. 2017, 202, 18–27. [Google Scholar] [CrossRef] [Scilit]
- Haklay, M.; Weber, P. OpenStreetMap: User-Generated Street Maps. IEEE Pervasive Comput. 2008, 7, 12–18. [Google Scholar] [CrossRef] [Scilit]
- Verma, A.; Bhattacharya, A.; Dey, S.; López-Martínez, C.; Gamba, P. Built-up area mapping using Sentinel-1 SAR data. ISPRS J. Photogramm. Remote Sens. 2023, 203, 55–70. [Google Scholar] [CrossRef] [Scilit]
- Ploton, P.; Mortier, F.; Réjou-Méchain, M.; Barbier, N.; Picard, N.; Rossi, V.; Dormann, C.; Cornu, G.; Viennois, G.; Bayol, N.; et al. Spatial validation reveals poor predictive performance of large-scale ecological mapping models. Nat. Commun. 2020, 11, 4540. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.-C.; Papandreou, G.; Schroff, F.; Adam, H. Rethinking Atrous Convolution for Semantic Image Segmentation. arXiv 2017, arXiv:1706.05587. [Google Scholar] [CrossRef] [Scilit]
- Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Proceedings of the Advances in Neural Information Processing Systems 27 (NIPS 2014), Montreal, QC, Canada, 8–13 December 2014; pp. 2672–2680. Available online: https://papers.nips.cc/paper_files/paper/2014/hash/f033ed80deb0234979a61f95710dbe25-Abstract.html (accessed on 20 June 2026).
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 2999–3007. [Google Scholar] [CrossRef] [Scilit]
- Milletari, F.; Navab, N.; Ahmadi, S.-A. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation. In Proceedings of the Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA, 25–28 October 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 565–571. [Google Scholar] [CrossRef] [Scilit]
- Cohen, J. A Coefficient of Agreement for Nominal Scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
- Congalton, R.G. A review of assessing the accuracy of classifications of remotely sensed data. Remote Sens. Environ. 1991, 37, 35–46. [Google Scholar] [CrossRef] [Scilit]
- Olofsson, P.; Foody, G.M.; Herold, M.; Stehman, S.V.; Woodcock, C.E.; Wulder, M.A. Good practices for estimating area and assessing accuracy of land change. Remote Sens. Environ. 2014, 148, 42–57. [Google Scholar] [CrossRef] [Scilit]
- Csurka, G.; Larlus, D.; Perronnin, F. What is a good evaluation measure for semantic segmentation? In Proceedings of the British Machine Vision Conference (BMVC), Bristol, UK, 9–13 September 2013; pp. 32.1–32.11. [Google Scholar] [CrossRef] [Scilit]
- Padilla, R.; Netto, S.L.; da Silva, E.A.B. A Survey on Performance Metrics for Object-Detection Algorithms. In Proceedings of the 27th International Conference on Systems, Signals and Image Processing (IWSSIP), Niterói, Brazil, 1–3 July 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 237–242. [Google Scholar] [CrossRef] [Scilit]
- Reichstein, M.; Camps-Valls, G.; Stevens, B.; Jung, M.; Denzler, J.; Carvalhais, N.; Prabhat. Deep learning and process understanding for data-driven Earth system science. Nature 2019, 566, 195–204. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Marconcini, M.; Metz-Marconcini, A.; Üreyen, S.; Palacios-Lopez, D.; Hanke, W.; Bachofer, F.; Zeidler, J.; Esch, T.; Gorelick, N.; Kakarla, A.; et al. Outlining where humans live, the World Settlement Footprint 2015. Sci. Data 2020, 7, 242. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, J.; Liu, H.; Yang, K.; Hu, X.; Liu, R.; Stiefelhagen, R. CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers. IEEE Trans. Intell. Transp. Syst. 2023, 24, 14679–14694. [Google Scholar] [CrossRef] [Scilit]
- Karasiak, N.; Dejoux, J.-F.; Monteil, C.; Sheeren, D. Spatial dependence between training and test sets: Another pitfall of classification accuracy assessment in remote sensing. Mach. Learn. 2022, 111, 2715–2740. [Google Scholar] [CrossRef] [Scilit]
- Esch, T.; Brzoska, E.; Dech, S.; Leutner, B.; Palacios-Lopez, D.; Metz-Marconcini, A.; Marconcini, M.; Roth, A.; Zeidler, J. World Settlement Footprint 3D—A first three-dimensional survey of the global building stock. Remote Sens. Environ. 2022, 270, 112877. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, X.; Pan, B.; Li, J. ACBDT: SAR-Optical Cross-Modal Distillation for Sentinel-1/2 Building-Footprint Mapping in Heterogeneous Yangtze River Delta Cities. Remote Sens. 2026, 18, 2868. https://doi.org/10.3390/rs18172868
Zhang X, Pan B, Li J. ACBDT: SAR-Optical Cross-Modal Distillation for Sentinel-1/2 Building-Footprint Mapping in Heterogeneous Yangtze River Delta Cities. Remote Sensing. 2026; 18(17):2868. https://doi.org/10.3390/rs18172868
Chicago/Turabian StyleZhang, Xianlong, Bin Pan, and Jianhua Li. 2026. "ACBDT: SAR-Optical Cross-Modal Distillation for Sentinel-1/2 Building-Footprint Mapping in Heterogeneous Yangtze River Delta Cities" Remote Sensing 18, no. 17: 2868. https://doi.org/10.3390/rs18172868
APA StyleZhang, X., Pan, B., & Li, J. (2026). ACBDT: SAR-Optical Cross-Modal Distillation for Sentinel-1/2 Building-Footprint Mapping in Heterogeneous Yangtze River Delta Cities. Remote Sensing, 18(17), 2868. https://doi.org/10.3390/rs18172868

