DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation
Highlights
- The proposed DBCS-T-based fusion method can effectively extract complementary polarization, spectral and intensity features. The fused multimodal image has higher information entropy and lower distribution divergence compared with original inputs, while maintaining high consistency with source data.
- The generated fusion image achieves accurate semantic segmentation of real vegetation, artificial foliage and same-color metallic objects under natural illumination, verifying the effectiveness of the designed network and fusion strategy.
- This method fully exploits the complementary advantages of spectro-polarimetric data, providing a feasible solution to the problem of insufficient utilization of multimodal information in spectro-polarimetric imaging.
- It shows great application potential in remote sensing image interpretation, and can serve as a powerful foundation for various high-level computer vision and remote sensing analysis tasks.
Abstract
1. Introduction
- We propose a dual-branch Transformer-CNN hybrid network (DBCS-T), which integrates a Cross-Channel Transposed Attention (CCTA) module for cross-modal interaction between Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) features and a Multi-Scale Polarization Feature Adaptive Modulation (MPAM) module for local detail enhancement, thereby generating a high-quality CPI.
- We construct a spectral prior-guided differential feature extraction strategy that effectively leverages the physical knowledge of target spectral radiance curves to directly derive the CSI, significantly improving the discrimination capability for metameric targets.
- We establish a PCA-based multimodal feature fusion framework that achieves synergistic utilization and redundancy elimination among polarization, spectral, and intensity information. The PCA-generated pseudo-color Multimodal Fusion Image retains essential information while effectively removing linear redundancy, and its practicality and effectiveness are validated in the corresponding experimental scenarios.
2. Methodology
| Algorithm 1: Multimodal Fusion Method Process |
|
2.1. CPI Extraction Based on the DBCS-T Network
2.2. CSI Extraction via Spectral Prior-Guided Difference Strategy
2.3. CII Extraction Derived from Stokes Vector
3. Experiment
3.1. Dataset Preparation
3.2. Loss Function Design
3.3. Network Training
3.4. Analysis of Metrics for Comparison and Ablation Experiments
4. Results and Analysis
4.1. Multimodal Image Acquisition
4.2. Multimodal Image Fusion
4.3. Fusion Result Analysis
4.3.1. Subjective Visual Evaluation
4.3.2. Objective Quantitative Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Fan, Y.; Huang, W.; Zhu, F.; Liu, X.; Jin, C.; Guo, C.; An, Y.; Kivshar, Y.; Qiu, C.W.; Li, W. Dispersion-assisted high-dimensional photodetector. Nature 2024, 630, 77–83. [Google Scholar] [CrossRef] [PubMed]
- Wang, F.; Fang, S.; Zhang, Y.; Wang, Q.J. 2D computational photodetectors enabling multidimensional optical information perception. Nat. Commun. 2025, 16, 6791. [Google Scholar] [CrossRef] [PubMed]
- Zhu, R.z.; Feng, H.g.; Xu, F. Deep learning-based multimode fiber imaging in multispectral and multipolarimetric channels. Opt. Lasers Eng. 2023, 161, 107386. [Google Scholar] [CrossRef]
- Li, S.; Tang, H. Multimodal Alignment and Fusion: A Survey. Int. J. Comput. Vis. 2026, 134, 103. [Google Scholar] [CrossRef]
- Altaqui, A.; Sen, P.; Schrickx, H.; Rech, J.; Lee, J.W.; Escuti, M.; You, W.; Kim, B.J.; Kolbas, R.; O’Connor, B.T.; et al. Mantis shrimp–inspired organic photodetector for simultaneous hyperspectral and polarimetric imaging. Sci. Adv. 2021, 7, eabe3196. [Google Scholar] [CrossRef] [PubMed]
- Li, S.; Jiao, J.; Wang, C. Research on Polarized Multi-Spectral System and Fusion Algorithm for Remote Sensing of Vegetation Status at Night. Remote Sens. 2021, 13, 3510. [Google Scholar] [CrossRef]
- Song, J.; Xue, Q.; Lu, F.; Li, K. Research on high-spectral polarization detection and classification of submerged oil based on multi-dimensional information. Opt. Laser Technol. 2025, 192, 113750. [Google Scholar] [CrossRef]
- Li, S.; Jiao, J.; Wang, C. Research on the Detection Algorithm of Camouflage Scattered Landmines in Vegetation Environment Based on Polarization Spectral Fusion. IEEE Geosci. Remote Sens. Lett. 2024, 21, 6011305. [Google Scholar] [CrossRef]
- Liu, Y.; Gao, K.; Wang, H.; Yang, Z.; Wang, P.; Ji, S.; Huang, Y.; Zhu, Z.; Zhao, X. A Transformer-based multi-modal fusion network for semantic segmentation of high-resolution remote sensing imagery. Int. J. Appl. Earth Obs. Geoinf. 2024, 133, 104083. [Google Scholar] [CrossRef]
- Saputra, M.R.U.; Bhaswara, I.D.; Nasution, B.I.; Ern, M.A.L.; Husna, N.L.R.; Witra, T.; Feliren, V.; Owen, J.R.; Kemp, D.; Lechner, A.M. Multi-modal deep learning approaches to semantic segmentation of mining footprints with multispectral satellite imagery. Remote Sens. Environ. 2025, 318, 114584. [Google Scholar] [CrossRef]
- Chen, Q.; Pang, M.; Liu, X.; Zhang, Z. A polarization-spectrum fusion framework based on multiscale transform and generative adversarial network for improving water and different vegetation distinguishability. Int. J. Appl. Earth Obs. Geoinf. 2023, 123, 103468. [Google Scholar] [CrossRef]
- Xu, H.; Yuan, J.; Ma, J. MURF: Mutually Reinforcing Multi-Modal Image Registration and Fusion. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12148–12166. [Google Scholar] [CrossRef] [PubMed]
- Wang, X.; Xu, J.; Ding, J. Polarization-based Camouflaged Object Detection with high-resolution adaptive fusion Network. Eng. Appl. Artif. Intell. 2025, 146, 110245. [Google Scholar] [CrossRef]
- Wang, X.; Zhang, Z.; Gao, J. Polarization-based Camouflaged Object Detection. Pattern Recognit. Lett. 2023, 174, 106–111. [Google Scholar] [CrossRef]
- Karim, S.; Tong, G.; Li, J.; Yu, Y.; Ibrar, M.; Mehmood, F. Dense Network-Based Spectral-Polarization Image Fusion: Multispectral Data Enhancement via Encoder-Decoder Approach. In Proceedings of the 6th International Conference on Information Technologies and Electrical Engineering; ACM: New York, NY, USA, 2023; pp. 441–446. [Google Scholar]
- Karim, S.; Tong, G.; Li, J.; Qadir, A.; Farooq, U.; Yu, Y. Current advances and future perspectives of image fusion: A comprehensive review. Inf. Fusion 2023, 90, 185–217. [Google Scholar] [CrossRef]
- Wang, R.; Zhou, Z.; Li, S.; Zhang, Z. Advances and challenges in infrared-visible image fusion: A comprehensive review of techniques and applications. Artif. Intell. Rev. 2025, 59, 18. [Google Scholar] [CrossRef]
- Mo, Y.; Kang, X.; Duan, P.; Sun, B.; Li, S. Attribute filter based infrared and visible image fusion. Inf. Fusion 2021, 75, 41–54. [Google Scholar] [CrossRef]
- Chen, J.; Ding, J.; Ma, J. HitFusion: Infrared and Visible Image Fusion for High-Level Vision Tasks Using Transformer. IEEE Trans. Multimed. 2024, 26, 10145–10159. [Google Scholar] [CrossRef]
- Kang, H.; Li, H.; Wu, X.; Xu, T.; Wang, R.; Cheng, C.; Kittler, J. Grformer: A Novel Transformer on Grassmann Manifold for Infrared and Visible Image Fusion. Inf. Fusion 2026, 125, 103402. [Google Scholar] [CrossRef]
- Ma, Q.; Li, X.; Li, B.; Zhu, Z.; Wu, J.; Huang, F.; Hu, H. STAMF: Synergistic transformer and mamba fusion network for RGB-Polarization based underwater salient object detection. Inf. Fusion 2025, 122, 103182. [Google Scholar] [CrossRef]
- Liu, X.; Wang, L. Infrared linear polarization small target enhancement algorithm in the cloudy background. J. Opt. Soc. Am. A 2023, 40, 859. [Google Scholar] [CrossRef] [PubMed]
- Shen, Y.; Liu, X.; Zhang, S.; Xu, Y.; Zeng, D.; Wang, S.; Huang, F. Real-Time Segmentation of Artificial Targets Using a Dual-Modal Efficient Attention Fusion Network. Remote Sens. 2023, 15, 4398. [Google Scholar] [CrossRef]
- Xiao, K.; Kang, X.; Liu, H.; Duan, P. MOFA: A novel dataset for Multi-modal Image Fusion Applications. Inf. Fusion 2023, 96, 144–155. [Google Scholar] [CrossRef]
- Liu, J.; Li, S.; Dian, R.; Song, Z. DT-F Transformer: Dual transpose fusion transformer for polarization image fusion. Inf. Fusion 2024, 106, 102274. [Google Scholar] [CrossRef]
- Tong, G.; Yao, X.; Li, B.; Fu, J.; Wang, Y.; Hao, J.; Karim, S.; Yu, Y. MSPFusion: A feature transformer for multidimensional spectral-polarization image fusion. Expert Syst. Appl. 2025, 275, 127079. [Google Scholar] [CrossRef]
- Shahdoosti, H.R.; Ghassemian, H. Combining the spectral PCA and spatial PCA fusion methods by an optimal filter. Inf. Fusion 2016, 27, 150–160. [Google Scholar] [CrossRef]
- Wu, J.; Li, B.; Ni, W.; Yan, W.; Zhang, H. Optimal Segmentation Scale Selection for Object-Based Change Detection in Remote Sensing Images Using Kullback–Leibler Divergence. IEEE Geosci. Remote Sens. Lett. 2020, 17, 1124–1128. [Google Scholar] [CrossRef]
- Zhao, Q.; Sbert, M.; Feixas, M.; Xu, Q. Multi-Exposure Image Fusion Based on Information-Theoretic Channel. In Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP); IEEE: New York, NY, USA, 2018; pp. 1872–1876. [Google Scholar]
- Wang, Z.; Bovik, A.; Sheikh, H.; Simoncelli, E. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [PubMed]
- Qiao, S.; Chen, R.; Xue, Z.; Wang, D. DPSR: Dual-branch network for robust polarization image super-resolution via color–polarization fusion. Opt. Laser Technol. 2026, 195, 114428. [Google Scholar] [CrossRef]
- Li, X.; Li, H.; Lin, Y.; Guo, J.; Yang, J.; Yue, H.; Li, K.; Li, C.; Cheng, Z.; Hu, H.; et al. Learning-based denoising for polarimetric images. Opt. Express 2020, 28, 16309. [Google Scholar] [CrossRef] [PubMed]
- Meng, J.; Ren, W.; Yu, R.; Ma, X.; Arce, G.R.; Wu, D.; Zhang, R.; Xie, Y. Learning based polarization image fusion under an alternative paradigm. Opt. Laser Technol. 2024, 168, 109969. [Google Scholar] [CrossRef]
- Yang, J.; Qiu, S.; Jin, W.; Wang, X.; Xue, F. Polarization imaging model considering the non-ideality of polarizers. Appl. Opt. 2020, 59, 306. [Google Scholar] [CrossRef] [PubMed]
- Dabov, K.; Foi, A.; Katkovnik, V.; Egiazarian, K. Image Denoising by Sparse 3-D Transform-Domain Collaborative Filtering. IEEE Trans. Image Process. 2007, 16, 2080–2095. [Google Scholar] [CrossRef] [PubMed]







| Category | Details | Configuration |
|---|---|---|
| Hardware | CPU | 11th Gen Intel Core i7-11700K |
| GPU | NVIDIA GeForce RTX 3060 12 GB | |
| Software Environment | Programming Language | Python 3.8 |
| Network Architecture | Transformer | |
| Training Parameters | Epochs | 30 |
| Learning Rate | 0.00005 | |
| Batch size | 1 |
| Method 1 | Method 2 | Method 3 | DT-F | DBCS-T | ||
|---|---|---|---|---|---|---|
| Entropy (bit) | Source 1 | 5.9004 | 5.9004 | 5.9004 | 5.9004 | 5.9004 |
| Source 2 | 5.7813 | 5.7813 | 5.7813 | 5.7813 | 5.7813 | |
| Average source | 6.8293 | 6.7692 | 6.8339 | 6.8369 | 6.8526 | |
| PSNR (dB) | Source 1 | 24.8928 | 21.9852 | 22.8205 | 20.7187 | 19.0155 |
| Source 2 | 5.3350 | 5.6551 | 4.8954 | 6.0593 | 7.0368 | |
| Average source | 15.1139 | 13.8201 | 13.8580 | 13.3890 | 13.0262 | |
| KL Divergence | Source 1 | 1.2561 | 1.4176 | 1.5213 | 1.0351 | 1.9482 |
| Source 2 | 2.1707 | 3.9125 | 2.8492 | 2.0901 | 2.3000 | |
| Average source | 1.2067 | 2.1584 | 1.6786 | 1.0559 | 1.6174 | |
| JS Divergence | Source 1 | 0.2083 | 0.2054 | 0.2711 | 0.1736 | 0.3255 |
| Source 2 | 0.4509 | 0.5175 | 0.4787 | 0.4678 | 0.4510 | |
| Average source | 0.2204 | 0.2585 | 0.2772 | 0.2105 | 0.2984 | |
| Conditional Entropy (bit) | Source 1 | 4.3672 | 4.4634 | 4.4314 | 4.8442 | 4.5359 |
| Source 2 | 5.5861 | 5.5701 | 5.5508 | 5.5586 | 5.6028 | |
| MS-SSIM | Source 1 | 0.9531 | 0.9360 | 0.9510 | 0.8878 | 0.9258 |
| Source 2 | 0.3644 | 0.4013 | 0.3800 | 0.4012 | 0.4302 | |
| Average source | 0.5272 | 0.5618 | 0.5430 | 0.5526 | 0.5874 |
| Entropy (bit) | PSNR (dB) | KL Divergence | JS Divergence | Conditional Entropy (bit) | MS-SSIM | ||
|---|---|---|---|---|---|---|---|
| DBCS-T Fusion | AoLP | 6.8550 | 10.9263 | 2.8476 | 0.3646 | 6.4445 (5.99%) | 0.5123 |
| DoLP | 4.7243 | 10.7838 | 5.4926 | 0.5715 | 4.3795 (7.30%) | 0.6106 | |
| CPI | 6.7389 | \ | 3.6110 | 0.4136 | \ | 0.5571 | |
| Multimodal Fusion | CPI | 6.7389 | 13.4823 | 0.4069 | 0.1109 | 6.6194 (1.77%) | 0.3367 |
| CSI | 6.7686 | 13.8214 | 1.1559 | 0.2667 | 6.2682 (7.39%) | 0.5006 | |
| MFI | 6.7816 | \ | 0.3714 | 0.0936 | \ | 0.4026 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, Y.; Gao, B.; Yang, X.; Li, H.; Yu, W.; Xu, H. DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation. Remote Sens. 2026, 18, 2700. https://doi.org/10.3390/rs18162700
Liu Y, Gao B, Yang X, Li H, Yu W, Xu H. DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation. Remote Sensing. 2026; 18(16):2700. https://doi.org/10.3390/rs18162700
Chicago/Turabian StyleLiu, Yiming, Bo Gao, Xiao Yang, Hang Li, Weixing Yu, and Huangrong Xu. 2026. "DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation" Remote Sensing 18, no. 16: 2700. https://doi.org/10.3390/rs18162700
APA StyleLiu, Y., Gao, B., Yang, X., Li, H., Yu, W., & Xu, H. (2026). DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation. Remote Sensing, 18(16), 2700. https://doi.org/10.3390/rs18162700

