TCM-Net: Mixed Global–Local Learning for Salient Object Detection in Optical Remote Sensing Images
Abstract
1. Introduction
- (1)
- A novel approach is proposed for ORSI-SOD that combines the strengths of the CNN and transformer. Rather than a simple combination of two networks, the U-shaped codec network architecture can learn better representation by fusing and refining the features from local details and global contexts across different layers.
- (2)
- To comprehensively aggregate the global contexts and local details from encoder layers while reducing the redundancy of the multi-scale features, we propose the LGFF module and the AG module. In addition, we tailored a hybrid loss function that incorporates three supervision strategies, global, local and output, for enhanced representation learning.
- (3)
2. Related Works
2.1. SOD Models in NSIs
- (1)
- Traditional Approaches: Over the past two decades, the theoretical system of SOD has undergone diversified development [27]. Itti et al. [28] proposed one of the earliest saliency models, which featured the well-known center-surround difference mechanism to locate salient objects. Since then, other models have emerged, such as the Kullback–Leibler divergence [29] and saliency tree [30]. In addition, hand-crafted features or visual priors, such as background prior [31], color histograms [32], color and brightness frequency-tuned detection [33] and color compactness [34], have been utilized to represent the saliency attribute of an object and generate bottom-up SOD models. However, these models have lower metrics or performance than deep-learning-based approaches.
- (2)
- Deep-Learning-Based Approaches: Recently, deep learning has achieved significant progress in SOD. Earlier deep-learning SOD models [35] utilized MLP classifiers to predict the saliency scores of deep features extracted from each image processing unit. A more effective and efficient approach is based on the full convolutional network (FCN) model [5,12,13], which uses the U-shaped network architecture with top-down feature encoding and skip connections to achieve semantic feature reuse at different levels. The attention mechanism has been widely applied to various computer vision tasks, including SOD. Liu [14] proposed a pixelwise contextual attention network for SOD by selectively focusing on contextual information of pixels. Meanwhile, the transformer model [24] has also been successful in global feature representation of SOD. Liu [10] proposed a visual saliency transformer for RGB and RGB-D SOD, and Xie [11] designed a pyramidal grafting network that combines a CNN and transformer to extract multi-scale features. These methods had achieved good results, but direct migration of NSI-SOD solutions to ORSI-SOD often resulted in unsatisfactory performance, and they still suffered from issues like blurred edges or inaccurate localization.
2.2. SOD Models in ORSIs
- (1)
- Traditional Approaches: Unlike a large number of traditional methods based on NSI-SOD, there are few works focusing on the ORSI-SOD. Zhao et al. [36] proposed a sparsity-guided saliency model that integrates saliency maps by incorporating global and background cues. Ma et al. [37] introduced a superpixel-to-pixel saliency model for detecting regions of interest based on texture and color features. Zhang et al. [38] aimed to determine the location of airports by integrating saliency results obtained from both vision-oriented and knowledge-oriented approaches. Zhang et al. [39] employed adaptive feature fusion of color, intensity, texture and global contrast using low-rank matrix recovery to generate the saliency map.
- (2)
- Deep-Learning-Based Approaches: Li et al. [18] first built the publicly available dataset called ORSSD for ORSI-SOD. Based on this work, Zhang et al. [19] extended the ORSSD dataset named EORSSD, containing some more challenging images, Tu et al. [20] also constructed the ORSI-4199 dataset with more complex scenarios and goals.Owing to the three public datasets, ORSI-SOD has received increasing attention [18,19,20,22,40,41,42,43,44,45,46]. Li et al. [18] introduced an end-to-end LV-Net for ORSI-SOD, which comprises a two-stream pyramid module and an encoder–decoder module. Zhou et al. [22] combined three strategies, namely image pyramid, feature pyramid and edge learning, to improve the performance of SOD. Cong et al. [41] is the first to explore the use of graph convolution networks for ORSI-SOD.
3. Methodology
3.1. Architecture Overview
3.2. Local and Global Feature Fusion Module (LGFF)
3.3. Attention Gate (AG) Module
3.4. Loss Function
4. Experiments
4.1. Experimental Settings
- (1)
- Datasets: To fully validate our model, we conducted an extensive comparison of three public benchmark ORSI datasets.ORSSD [18] is a dataset that includes 800 ORSIs, depicting significant objects, each accompanied by its corresponding ground truth. Of these, 600 images are used for training, and the remaining 200 images are reserved for testing.EORSSD [19] is an extension of ORSSD, which includes 2000 more comprehensive and diverse scenes ORSIs with the corresponding GT. It consists of 1400 images as the training subset and the other 600 images for testing.ORSI-4199 [20] is the latest and most challenging ORSI-SOD dataset. The dataset contains a total of 4199 images and the corresponding GTs, of which 2000 images are used for training and the remaining 2199 images are used as a test subset. In addition, it defines nine different scene attributes, which helps us to objectively evaluate the SOD models with various attributes.
- (2)
- Experimental details: Our model was implemented using PyTorch on a machine comprising an Intel(R) Core(TM) i9-10900X 3.70 GHz × 20 CPU, 128 GB of RAM (Kunshan, Jiangsu, China) and an NVIDIA GTX 2080Ti GPU (Suzhou, Jiangsu, China). The encoder was optimized using ResNet-34 [47] pre-training weights from PyTorch and Swin transformer [48] pre-training weights from Swin-B_224. The proposed TCM-Net can be trained end-to-end, and the network is optimized using Adam’s algorithm [55]. The maximum learning rate was set to 0.03 for the Swin backbone and 0.03 for the remaining components. During training, the learning initially increased and then decayed, with momentum set to 0.9 and weight decay to 0.001. The batch size was set to 8. Furthermore, during the training phase, each training image was resized to 1024 × 1024 and 224 × 224. The high-resolution images were input into the ResNet34 branch, while the low-resolution images were input into the Swin transformer branch. Due to the small size of the ORSSD dataset, we applied a strategy based on random flipping and rotation described in [18,19], which produced seven additional variants of the original training data to obtain a training set of 4800 training images. Then, three datasets were enhanced by cropping and multi-scale input images.
4.2. Evaluation Metrics
4.3. Comparison with SOTA Methods
- (1)
- Quantitative comparison:
- (a)
- Quantitative comparison of ORSSD: We present the performance of our method TCM-Net for the ORSSD [18] dataset by quantitatively comparing it with other methods. The evaluation was conducted based on three metrics: MAE, max F-measure () and S-measure (), and the results are shown in Table 1. Our method demonstrates superior performance compared to all other methods on all three metrics; while MCCNet [40] is the best performer among the remaining 17 methods, our method exhibits even better performance. Specifically, our method outperforms MCCNet with a 2.14% improvement in , a 13.8% lower MAE and a 0.07% better . Despite RRNet [41] and DAFNet [19] achieving higher scores of 0.9210 and 0.9166, our method surpasses them by 1.62% and 2.11% in the metric, respectively. Furthermore, Figure 7 illustrates the PR curve and the F-measure curve. In the three datasets with 17 methods, our model’s PR curve is positioned closer to the upper right corner, while our F-measure curve covers a larger area compared to other methods.
- (b)
- Quantitative comparison of EORSSD: We present a quantitative comparison of our approach with other methods on the EORSSD [19] dataset, as shown in the middle column of Table 1. Additionally, we plotted the PR curve and the F-measurement curve in Figure 7, which highlight our method’s superior performance. Among the compared methods, our approach achieves the best performance in two metrics, MAE and , and also has the best PR curve and F-measurement curve. However, our metric is slightly behind ACCoNet [42] by 0.03%. Our value is 0.9332, whereas ACCoNet’s is 0.9335. Despite this, our outperforms ACCoNet by 3.70%, and our MAE is 12% lower. Notably, RRNet [41] has the highest value of 0.9031 among all the compared methods, except our approach. However, our is still superior to RRNet by 1.25%.
- (c)
- Quantitative comparison of ORSI-4199: Based on the ORSI-4199 dataset [20], we performed a quantitative comparison of the three metrics and present the results in Table 1. The corresponding PR curve and F-measure curve are shown in Figure 7 on the right. Overall, our method outperforms other methods, achieving values of 0.0275, 0.8717 and 0.8760, respectively. In addition, our method showed the best performance in the PR curve F-measure curve. Among the other methods, ACCoNet [42] achieved the best results with MAE, and values of 0.0328, 0.8584 and 0.8800, respectively. Compared to ACCoNet, our was 0.46% lower, but our was 1.55% higher and MAE was 19.3% lower. Moreover, our comparison with VSTNet [10], which showed the highest performance among the NSI-SOD methods, demonstrated the high applicability of the transformer model in the SOD task. VSTNet achieved values of 0.8543, 0.0306 and 0.8752 for MAE, and , respectively. Our method achieved a better (improved by 2.04%), lower MAE (reduced by 11.3%) and slightly improved (by 0.09%) compared to VSTNet.
- (2)
- Visual comparison: To qualitatively compare all methods, we selected experimental results from the ORSI-4199 dataset, which includes representative and challenging scenes, as shown in Figure 8. These scenes involve multiple objects, complex scenes, low contrast, shadow occlusion and more. Our model’s predictions demonstrate greater completeness and accuracy in detecting salient objects compared to other methods shown in Figure 8. The specific advantages are reflected in the following aspects.
- (a)
- Superiority in scenes with multiple or multiscale objects: In the first and third examples of Figure 8, both traditional models (e.g., LC [32] and MBD [58] shown in Figure 8q,s) and deep-learning-based models (e.g., EMFINet [22], DAFNet [19] and F3Net [5] presented in Figure 8h,j,m) fail to highlight the foreground regions accurately, resulting in prominent errors. In contrast, our model provides complete and accurate inference for salient objects. Similarly, in the eighth and thirteenth examples of Figure 8, some state-of-the-art deep-learning models (e.g., ACCoNet [42], MCCNet [40] and MSCNet [43] depicted in Figure 8d,e,g) either incorrectly detect multiple objects or fail to fully pop out the salient objects. In stark contrast, the model shown in Figure 8 can still successfully highlight the salient objects, and our results exhibit clear boundaries, particularly in preserving the overall integrity of multiple small airplanes and a single large building. This is clearly attributed to the superiority of our U-shaped encoder–decoder architecture and the control of redundant information by the AG module.
- (b)
- Superiority in cluttered background and low contrast scenes: In the fourth, fifth and sixth examples of Figure 8, traditional models (e.g., LC [32] and FT [33] shown in Figure 8r,s) completely fail to detect salient objects, while deep-learning-based models (e.g., ACCoNet [42], MSCNet [43], PoolNet [7] presented in Figure 8d,g,n) either provide incomplete detection or incorrectly highlight background regions. In contrast, our model successfully detects salient bridges, four ships and two airplanes from the aforementioned three examples. Similarly, in Figure 8, for the fourteenth, fifteenth and seventeenth examples, state-of-the-art deep-learning models (e.g., ACCoNet [42], MCCNet [40] and RRNet [41]) either erroneously highlight background regions or fail to clearly distinguish salient objects, as shown in Figure 8d–f). In contrast, the model depicted in Figure 8c can fully and clearly pop out all salient objects. This is clearly attributed to our effective control of global contextual and local detailed information, as well as the LGFF module’s ability to fuse information from both sources effectively.
- (c)
- Superiority in salient regions with complicated edges or irregular topology: In the tenth, eleventh and twelfth examples of Figure 8, both traditional models (e.g., LC [32] and MBD [58] shown in Figure 8q,s) and deep-learning-based models (e.g., ACCoNet [42], MCCNet [40], EMFINet [22], MJRBMNet [20] and GateNet [9]) presented in Figure 8d,e,h,i,l fail to accurately delineate the lake regions comprehensively and also fall short in detecting irregular rivers and buildings. In contrast, our model, as shown in Figure 8c, outperforms the other models. The forest regions are effectively suppressed, irregular lakes are fully highlighted and, for irregular topological structures of rivers and buildings, our model is capable of generating a more complete saliency map with more accurate boundaries. It is evident that these superior results are attributed to the addition of edge supervision in our hybrid loss and our effective control of global-to-local information for image localization.
- (3)
- Attribute-based study: In the latest ORSI-4199 dataset [20], each image is meticulously categorized according to the distinctive attributes commonly found in ORSIs. These attributes include the presence of big salient objects (BSOs), small salient objects (SSOs), off center (OC) objects, complex salient objects (CSOs), complex scenes (CSs), narrow salient objects (NSOs), multiple salient objects (MSOs), low contrast scenes (LCSs) and incomplete salient objects (ISOs). These annotations allow us to compare the strengths and weaknesses of our proposed model and other models under different conditions. Table 2 shows the values of our model and other state-of-the-art models. Our model ranks first in seven of the nine attributes and second in scores for the remaining two attributes. Additionally, we used radar plots for the first time to depict the values of the top five ranked methods for different attributes, and the results confirm that our model outperforms existing models in the majority of challenging scenarios, as shown in Figure 9.
4.4. Extension Experiment on NSI Datasets
4.5. Ablation Study
- (1)
- Module ablation experiments: We conducted module ablation experiments on the EORSSD dataset to evaluate the effectiveness of our proposed ResNet34+Swin transformer model compared to baseline models using either ResNet34 or the Swin transformer alone. Additionally, we evaluated the U-shaped encoder–decoder network architecture to determine its effectiveness, with the modular ablation study focusing on two specific modules, namely LGFF and AG. To visually demonstrate the effectiveness of our proposed architecture and modules, we present saliency maps in Figure 11, highlighting the impact of different ablation modules and baselines.
- (2)
- Loss ablation experiments: We performed ablation experiments on the EORSSD dataset [19] to demonstrate the effectiveness of our hybrid loss function strategy, which is specifically designed for the network structure. To demonstrate the complementarity of BCE, IoU, SSIM and Fm in the loss function, we evaluated our loss function using four common loss function strategies: (1) training our model only with BCE loss, e.g., [6]; (2) training our model with BCE-IoU loss, e.g., [42]; (3) training our model jointly with the three loss functions of BCE, IoU and SSIM, e.g., [22]; (4) training our model jointly with the three loss functions of BCE, IoU and Fm, e.g., [40]; and (5) using our selected hybrid function and allocation strategy that includes BCE, IoU, SSIM and Fm. The results are shown in Table 5.
4.6. Cross-Validation Analysis of the Model
4.7. Evaluation of Loss Weighting
4.8. Complexity Analysis
4.9. Failure Case
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Zeng, Y.; Zhuge, Y.; Lu, H.; Zhang, L. Joint learning of saliency detection and weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7223–7233. [Google Scholar]
- Fang, Y.; Lin, W.; Chen, Z.; Tsai, C.M.; Lin, C.W. A video saliency detection model in compressed domain. IEEE Trans. Circuits Syst. Video Technol. 2013, 24, 27–38. [Google Scholar] [CrossRef] [Scilit]
- Yuan, Y.; Lu, Y.; Wang, Q. Tracking as a whole: Multi-target tracking by modeling group behavior with sequential detection. IEEE Trans. Intell. Transp. Syst. 2017, 18, 3339–3349. [Google Scholar] [CrossRef] [Scilit]
- Yang, S.; Jiang, Q.; Lin, W.; Wang, Y. SGDNet: An End-to-End Saliency-Guided Deep Neural Network for No-Reference Image Quality Assessment. In Proceedings of the 27th ACM International Conference on Multimedia, MM 2019, Nice, France, 21–25 October 2019; Amsaleg, L., Huet, B., Larson, M.A., Gravier, G., Hung, H., Ngo, C., Ooi, W.T., Eds.; ACM: New York, NY, USA, 2019; pp. 1383–1391. [Google Scholar] [CrossRef] [Scilit]
- Wei, J.; Wang, S.; Huang, Q. F3Net: Fusion, feedback and focus for salient object detection. Proc. AAAI Conf. Artif. Intell. 2020, 34, 12321–12328. [Google Scholar] [CrossRef] [Scilit]
- Wu, Z.; Su, L.; Huang, Q. Cascaded partial decoder for fast and accurate salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 3907–3916. [Google Scholar]
- Liu, J.J.; Hou, Q.; Cheng, M.M.; Feng, J.; Jiang, J. A simple pooling-based design for real-time salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 3917–3926. [Google Scholar]
- Wu, Z.; Su, L.; Huang, Q. Stacked cross refinement network for edge-aware salient object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7264–7273. [Google Scholar]
- Zhao, X.; Pang, Y.; Zhang, L.; Lu, H.; Zhang, L. Suppress and balance: A simple gated network for salient object detection. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part II 16. Springer: Berlin/Heidelberg, Germany, 2020; pp. 35–51. [Google Scholar]
- Liu, N.; Zhang, N.; Wan, K.; Shao, L.; Han, J. Visual saliency transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtual, 11–17 October 2021; pp. 4722–4732. [Google Scholar]
- Xie, C.; Xia, C.; Ma, M.; Zhao, Z.; Chen, X.; Li, J. Pyramid grafting network for one-stage high resolution saliency detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 11717–11726. [Google Scholar]
- Qin, X.; Zhang, Z.; Huang, C.; Gao, C.; Dehghan, M.; Jagersand, M. Basnet: Boundary-aware salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 7479–7489. [Google Scholar]
- Qin, X.; Zhang, Z.; Huang, C.; Dehghan, M.; Zaiane, O.R.; Jagersand, M. U2-Net: Going deeper with nested U-structure for salient object detection. Pattern Recognit. 2020, 106, 107404. [Google Scholar] [CrossRef] [Scilit]
- Liu, N.; Han, J.; Yang, M.H. Picanet: Learning pixel-wise contextual attention for saliency detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake CIty, UT, USA, 18–22 June 2018; pp. 3089–3098. [Google Scholar]
- Hou, Q.; Cheng, M.M.; Hu, X.; Borji, A.; Tu, Z.; Torr, P.H. Deeply supervised salient object detection with short connections. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 3203–3212. [Google Scholar]
- Zhao, C. Advances of research and application in remote sensing for agriculture. Nongye Jixie Xuebao Trans. Chin. Soc. Agric. Mach. 2014, 45, 277–293. [Google Scholar]
- Bello, O.M.; Aina, Y.A. Satellite remote sensing as a tool in disaster management and sustainable development: Towards a synergistic approach. Procedia Soc. Behav. Sci. 2014, 120, 365–373. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Cong, R.; Hou, J.; Zhang, S.; Qian, Y.; Kwong, S. Nested network with two-stream pyramid for salient object detection in optical remote sensing images. IEEE Trans. Geosci. Remote Sens. 2019, 57, 9156–9166. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Cong, R.; Li, C.; Cheng, M.M.; Fang, Y.; Cao, X.; Zhao, Y.; Kwong, S. Dense attention fluid network for salient object detection in optical remote sensing images. IEEE Trans. Image Process. 2020, 30, 1305–1317. [Google Scholar] [CrossRef] [Scilit]
- Tu, Z.; Wang, C.; Li, C.; Fan, M.; Zhao, H.; Luo, B. ORSI salient object detection via multiscale joint region and boundary model. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5607913. [Google Scholar] [CrossRef] [Scilit]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2881–2890. [Google Scholar]
- Zhou, X.; Shen, K.; Liu, Z.; Gong, C.; Zhang, J.; Yan, C. Edge-Aware Multiscale Feature Integration Network for Salient Object Detection in Optical Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5605315. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 834–848. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Chen, K.; Zou, Z.; Shi, Z. Building Extraction from Remote Sensing Images with Sparse Token Transformers. Remote. Sens. 2021, 13, 4441. [Google Scholar] [CrossRef] [Scilit]
- Fang, J.; Lin, H.; Chen, X.; Zeng, K. A Hybrid Network of CNN and Transformer for Lightweight Image Super-Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2022, New Orleans, LA, USA, 19–20 June 2022; pp. 1102–1111. [Google Scholar] [CrossRef] [Scilit]
- Borji, A.; Cheng, M.M.; Jiang, H.; Li, J. Salient object detection: A benchmark. IEEE Trans. Image Process. 2015, 24, 5706–5722. [Google Scholar] [CrossRef] [Scilit]
- Itti, L.; Koch, C.; Niebur, E. A model of saliency-based visual attention for rapid scene analysis. IEEE Trans. Pattern Anal. Mach. Intell. 1998, 20, 1254–1259. [Google Scholar] [CrossRef] [Scilit]
- Klein, D.A.; Frintrop, S. Center-surround divergence of feature statistics for salient object detection. In Proceedings of the 2011 International Conference on Computer Vision, Virtual, 11–17 October 2021; pp. 2214–2219. [Google Scholar]
- Liu, Z.; Zou, W.; Le Meur, O. Saliency tree: A novel saliency detection framework. IEEE Trans. Image Process. 2014, 23, 1937–1952. [Google Scholar]
- Zhu, W.; Liang, S.; Wei, Y.; Sun, J. Saliency optimization from robust background detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 2814–2821. [Google Scholar]
- Zhai, Y.; Shah, M. Visual attention detection in video sequences using spatiotemporal cues. In Proceedings of the 14th ACM International Conference on MULTIMEDIA, Santa Barbara, CA, USA, 23–27 October 2006; pp. 815–824. [Google Scholar]
- Achanta, R.; Hemami, S.; Estrada, F.; Susstrunk, S. Frequency-tuned salient region detection. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 1597–1604. [Google Scholar]
- Zhou, L.; Yang, Z.; Yuan, Q.; Zhou, Z.; Hu, D. Salient region detection via integrating diffusion-based compactness and local contrast. IEEE Trans. Image Process. 2015, 24, 3308–3320. [Google Scholar] [CrossRef] [Scilit]
- Liu, T.; Sun, J.; Zheng, N.N.; Tang, X.; Shum, H.Y. Learning to Detect A Salient Object. In Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MI, USA, 18–23 June 2007; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
- Zhao, D.; Wang, J.; Shi, J.; Jiang, Z. Sparsity-guided saliency detection for remote sensing images. J. Appl. Remote Sens. 2015, 9, 95055. [Google Scholar] [CrossRef] [Scilit]
- Ma, L.; Du, B.; Chen, H.; Soomro, N.Q. Region-of-interest detection via superpixel-to-pixel saliency analysis for remote sensing image. IEEE Geosci. Remote Sens. Lett. 2016, 13, 1752–1756. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Zhang, L.; Shi, W.; Liu, Y. Airport Extraction via Complementary Saliency Analysis and Saliency-Oriented Active Contour Model. IEEE Geosci. Remote. Sens. Lett. 2018, 15, 1085–1089. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Liu, Y.; Zhang, J. Saliency detection based on self-adaptive multiple feature fusion for remote sensing images. Int. J. Remote Sens. 2019, 40, 8270–8297. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Liu, Z.; Lin, W.; Ling, H. Multi-content complementation network for salient object detection in optical remote sensing images. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5614513. [Google Scholar] [CrossRef] [Scilit]
- Cong, R.; Zhang, Y.; Fang, L.; Li, J.; Zhao, Y.; Kwong, S. RRNet: Relational reasoning network with parallel multiscale attention for salient object detection in optical remote sensing images. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5613311. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Liu, Z.; Zeng, D.; Lin, W.; Ling, H. Adjacent context coordination network for salient object detection in optical remote sensing images. IEEE Trans. Cybern. 2022, 53, 526–538. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, Y.; Sun, H.; Liu, N.; Bian, Y.; Cen, J.; Zhou, H. A lightweight multi-scale context network for salient object detection in optical remote sensing images. In Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR), Montreal, QC, Canada, 21–25 August 2022; pp. 238–244. [Google Scholar]
- Bai, Z.; Li, G.; Liu, Z. Global–local–global context-aware network for salient object detection in optical remote sensing images. ISPRS J. Photogramm. Remote Sens. 2023, 198, 184–196. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Liu, Y.; Xiong, Z.; Yuan, Y. Hybrid Feature Aligned Network for Salient Object Detection in Optical Remote Sensing Imagery. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5624915. [Google Scholar] [CrossRef] [Scilit]
- Huang, Z.; Chen, H.; Liu, B.; Wang, Z. Semantic-Guided Attention Refinement Network for Salient Object Detection in Optical Remote Sensing Images. Remote Sens. 2021, 13, 2163. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtual, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
- Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N.Y.; Kainz, B.; et al. Attention u-net: Learning where to look for the pancreas. arXiv 2018, arXiv:1804.03999. [Google Scholar]
- De Boer, P.T.; Kroese, D.P.; Mannor, S.; Rubinstein, R.Y. A tutorial on the cross-entropy method. Ann. Oper. Res. 2005, 134, 19–67. [Google Scholar] [CrossRef] [Scilit]
- Yu, J.; Jiang, Y.; Wang, Z.; Cao, Z.; Huang, T. Unitbox: An advanced object detection network. In Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands, 15–19 October 2016; pp. 516–520. [Google Scholar]
- Wang, Z.; Simoncelli, E.P.; Bovik, A.C. Multiscale structural similarity for image quality assessment. In Proceedings of the Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, Pacific Grove, CA, USA, 9–12 November 2003; Volume 2, pp. 1398–1402. [Google Scholar]
- Zhao, K.; Gao, S.; Wang, W.; Cheng, M.M. Optimizing the F-measure for threshold-free salient object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 8849–8857. [Google Scholar]
- Canny, J. A computational approach to edge detection. IEEE Trans. Pattern Anal. Mach. Intell. 1986, 8, 679–698. [Google Scholar] [CrossRef] [Scilit]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
- Fan, D.P.; Cheng, M.M.; Liu, Y.; Li, T.; Borji, A. Structure-measure: A new way to evaluate foreground maps. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 4548–4557. [Google Scholar]
- Perazzi, F.; Krähenbühl, P.; Pritch, Y.; Hornung, A. Saliency filters: Contrast based filtering for salient region detection. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012; pp. 733–740. [Google Scholar]
- Zhang, J.; Sclaroff, S.; Lin, Z.; Shen, X.; Price, B.; Mech, R. Minimum barrier salient object detection at 80 fps. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 1404–1412. [Google Scholar]
- Yang, C.; Zhang, L.; Lu, H.; Ruan, X.; Yang, M.H. Saliency detection via graph-based manifold ranking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA, 23–28 June 2013; pp. 3166–3173. [Google Scholar]
- Wang, L.; Lu, H.; Wang, Y.; Feng, M.; Wang, D.; Yin, B.; Ruan, X. Learning to detect salient objects with image-level supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 136–145. [Google Scholar]
- Li, G.; Yu, Y. Visual saliency based on multiscale deep features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 5455–5463. [Google Scholar]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.; et al. Segment Anything. arXiv 2023, arXiv:2304.02643. [Google Scholar] [CrossRef] [Scilit]












| Methods | ORSSD [18] | EORSSD [19] | ORSI-4199 [20] | ||||||
|---|---|---|---|---|---|---|---|---|---|
| ↑ | MAE↓ | ↑ | ↑ | MAE↓ | ↑ | ↑ | MAE↓ | ↑ | |
| LC (2006) [32] | 0.4162 | 0.1243 | 0.5904 | 0.4450 | 0.0871 | 0.5926 | 0.3532 | 0.1904 | 0.5248 |
| FT (2009) [33] | 0.4023 | 0.3889 | 0.4269 | 0.3955 | 0.4221 | 0.4061 | 0.4303 | 0.4109 | 0.4600 |
| MBD (2015) [58] | 0.6412 | 0.0766 | 0.7018 | 0.5575 | 0.0526 | 0.6758 | 0.5508 | 0.1353 | 0.6247 |
| CPDNet (2019) [6] | 0.8259 | 0.0219 | 0.8651 | 0.7466 | 0.0201 | 0.8237 | 0.7996 | 0.0538 | 0.8245 |
| SCRNet (2019) [8] | 0.8245 | 0.0232 | 0.8558 | 0.8016 | 0.0127 | 0.8403 | 0.8176 | 0.0443 | 0.8452 |
| PoolNet (2019) [7] | 0.8293 | 0.0328 | 0.8660 | 0.7421 | 0.0191 | 0.8258 | 0.7213 | 0.0701 | 0.7664 |
| F3Net (2020) [5] | 0.8820 | 0.0161 | 0.9011 | 0.8764 | 0.0094 | 0.9085 | 0.8242 | 0.0414 | 0.8419 |
| GateNet (2020) [9] | 0.8885 | 0.0139 | 0.9133 | 0.8483 | 0.0102 | 0.8953 | 0.8429 | 0.0392 | 0.8595 |
| VST (2021) [10] | 0.8765 | 0.0115 | 0.9152 | 0.8749 | 0.0070 | 0.9143 | 0.8543 | 0.0306 | 0.8752 |
| MJRBMNet (2021) [20] | 0.8874 | 0.0131 | 0.9270 | 0.8709 | 0.0097 | 0.9247 | 0.8386 | 0.0384 | 0.8637 |
| EMFINet (2022) [22] | 0.9066 | 0.0125 | 0.9364 | 0.8717 | 0.0073 | 0.9279 | 0.8327 | 0.0375 | 0.8566 |
| RRNet (2021) [41] | 0.9210 | 0.0117 | 0.9268 | 0.9031 | 0.0082 | 0.9203 | 0.8070 | 0.0505 | 0.8317 |
| DAFNet (2020) [19] | 0.9166 | 0.0119 | 0.9154 | 0.8752 | 0.0088 | 0.8868 | 0.8252 | 0.0449 | 0.8471 |
| LVNet (2019) [18] | 0.8263 | 0.0207 | 0.8813 | 0.7824 | 0.0145 | 0.8650 | - | - | - |
| ACCoNet (2022) [42] | 0.9112 | 0.0103 | 0.9336 | 0.8818 | 0.0075 | 0.9335 | 0.8584 | 0.0328 | 0.8800 |
| MSCNet (2022) [43] | 0.8962 | 0.0120 | 0.9293 | 0.8597 | 0.0084 | 0.9180 | 0.8487 | 0.0377 | 0.8584 |
| MCCNet (2021) [40] | 0.9163 | 0.0094 | 0.9421 | 0.8875 | 0.0068 | 0.9329 | 0.8301 | 0.0379 | 0.8517 |
| Ours | 0.9359 | 0.0081 | 0.9428 | 0.9144 | 0.0066 | 0.9332 | 0.8717 | 0.0275 | 0.8760 |
| Attr | BSO | CS | CSO | ISO | LSO | MSO | NSO | OC | SSO | Avg |
|---|---|---|---|---|---|---|---|---|---|---|
| ACCoNet [42] | 0.9207 | 0.8976 | 0.8864 | 0.9012 | 0.7713 | 0.8486 | 0.8751 | 0.8329 | 0.7843 | 0.8576 |
| DAFNet [19] | 0.8875 | 0.8682 | 0.8631 | 0.8598 | 0.7444 | 0.8048 | 0.8458 | 0.7808 | 0.7370 | 0.8213 |
| EMFINet [22] | 0.9185 | 0.8790 | 0.8870 | 0.9018 | 0.7418 | 0.8188 | 0.8432 | 0.7562 | 0.7286 | 0.8305 |
| RRNet [41] | 0.8533 | 0.8413 | 0.8359 | 0.8314 | 0.7354 | 0.8017 | 0.8393 | 0.7697 | 0.7237 | 0.8035 |
| MJRBNet [20] | 0.8951 | 0.8800 | 0.8601 | 0.8835 | 0.7468 | 0.8415 | 0.8471 | 0.8282 | 0.7724 | 0.8394 |
| MCCNet [40] | 0.9056 | 0.8749 | 0.8685 | 0.8858 | 0.7383 | 0.8075 | 0.8477 | 0.8060 | 0.7414 | 0.8306 |
| MSCNet [43] | 0.9014 | 0.8886 | 0.8676 | 0.8961 | 0.7659 | 0.8375 | 0.8811 | 0.8123 | 0.7699 | 0.8467 |
| VSTNet [10] | 0.9107 | 0.8978 | 0.9071 | 0.9269 | 0.7743 | 0.8299 | 0.9022 | 0.8103 | 0.7435 | 0.8559 |
| F3Net [5] | 0.8852 | 0.8625 | 0.8578 | 0.8796 | 0.7495 | 0.8249 | 0.8668 | 0.7861 | 0.7426 | 0.8283 |
| PoolNet [7] | 0.7473 | 0.7436 | 0.7365 | 0.7096 | 0.6442 | 0.7464 | 0.7669 | 0.6962 | 0.6605 | 0.7168 |
| GateNet [9] | 0.9140 | 0.8841 | 0.8807 | 0.8954 | 0.7598 | 0.8372 | 0.8535 | 0.8059 | 0.7624 | 0.8437 |
| SCRNNet [8] | 0.8945 | 0.8647 | 0.8658 | 0.8752 | 0.7283 | 0.8181 | 0.8296 | 0.7875 | 0.7354 | 0.8221 |
| CPDNet [6] | 0.8850 | 0.8479 | 0.8482 | 0.8631 | 0.7176 | 0.7839 | 0.8009 | 0.7740 | 0.7070 | 0.8031 |
| Ours | 0.9324 | 0.9142 | 0.9037 | 0.9275 | 0.7845 | 0.8515 | 0.8987 | 0.8386 | 0.7920 | 0.8715 |
| Methods | DUT-OMRON [59] | DUTS-TE [60] | HKU-IS [61] | ||||||
|---|---|---|---|---|---|---|---|---|---|
| ↑ | MAE↓ | ↑ | ↑ | MAE↓ | ↑ | ↑ | MAE↓ | ↑ | |
| CPDNet [6] | 0.7513 | 0.0550 | 0.8237 | 0.8374 | 0.0434 | 0.8675 | 0.9115 | 0.0336 | 0.9078 |
| SCRNet [8] | 0.7628 | 0.0626 | 0.8300 | 0.8556 | 0.0440 | 0.8791 | 0.9220 | 0.0343 | 0.9166 |
| PoolNet [7] | 0.7280 | 0.0597 | 0.8049 | 0.8187 | 0.0457 | 0.8529 | 0.8911 | 0.0386 | 0.8912 |
| F3Net [5] | 0.7777 | 0.0578 | 0.8346 | 0.8714 | 0.0373 | 0.8874 | 0.9287 | 0.0279 | 0.9205 |
| GateNet [9] | 0.7748 | 0.0535 | 0.8370 | 0.8723 | 0.0370 | 0.8895 | 0.9239 | 0.0323 | 0.9188 |
| VST [10] | 0.8066 | 0.0549 | 0.8562 | 0.8780 | 0.0370 | 0.8966 | 0.9371 | 0.0296 | 0.9288 |
| MJRBMNet [20] | 0.7493 | 0.0663 | 0.8192 | 0.8267 | 0.0521 | 0.8607 | 0.9071 | 0.0376 | 0.9066 |
| EMFINet [22] | 0.7697 | 0.0598 | 0.8281 | 0.8278 | 0.0500 | 0.8585 | 0.9138 | 0.0336 | 0.9066 |
| RRNet [41] | 0.7728 | 0.0622 | 0.8326 | 0.7830 | 0.0621 | 0.8279 | 0.8671 | 0.0523 | 0.8717 |
| DAFNet [19] | 0.7268 | 0.0778 | 0.8004 | 0.7523 | 0.0770 | 0.8084 | 0.8616 | 0.0622 | 0.8674 |
| ACCoNet [42] | 0.7635 | 0.0631 | 0.8253 | 0.8282 | 0.0500 | 0.8588 | 0.9128 | 0.0340 | 0.9062 |
| MSCNet [43] | 0.7711 | 0.0632 | 0.8254 | 0.8188 | 0.0537 | 0.8492 | 0.9014 | 0.0415 | 0.8940 |
| MCCNet [40] | 0.7750 | 0.0576 | 0.8296 | 0.8356 | 0.0465 | 0.8616 | 0.9167 | 0.0326 | 0.9054 |
| Ours | 0.7973 | 0.0535 | 0.8426 | 0.8839 | 0.0336 | 0.8953 | 0.9378 | 0.0243 | 0.9283 |
| Backbone | LGFF | AG | ↑ | MAE↓ | ↑ |
|---|---|---|---|---|---|
| Swin transformer | 0.5409 | 0.0214 | 0.7433 | ||
| Swin + U | 0.8654 | 0.083 | 0.8998 | ||
| ResNet34 | 0.7981 | 0.0151 | 0.8669 | ||
| ResNet34 + U | 0.8563 | 0.0124 | 0.8990 | ||
| Baseline | 0.8980 | 0.0077 | 0.9235 | ||
| Baseline | ✓ | 0.9029 | 0.0075 | 0.9246 | |
| Baseline | ✓ | 0.9126 | 0.0067 | 0.9318 | |
| Baseline | ✓ | ✓ | 0.9144 | 0.0066 | 0.9332 |
| No. | BCE | IoU | SSIM | Fm | ↑ | MAE↓ | ↑ |
|---|---|---|---|---|---|---|---|
| 1 | ✓ | 0.8459 | 0.0095 | 0.8864 | |||
| 2 | ✓ | ✓ | 0.9119 | 0.0070 | 0.9309 | ||
| 3 | ✓ | ✓ | ✓ | 0.9080 | 0.0069 | 0.9295 | |
| 4 | ✓ | ✓ | ✓ | 0.9103 | 0.0071 | 0.9314 | |
| 5 | ✓ | ✓ | ✓ | ✓ | 0.9038 | 0.0071 | 0.9272 |
| 6 (Ours) | ✓ | ✓ | ✓ | ✓ | 0.9144 | 0.0066 | 0.9332 |
| NO. | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.9317 | 0.9230 | 0.9259 | 0.9295 | 0.9310 | 0.9299 | 0.9201 | 0.8811 | 0.9033 | 0.9149 | 0.9190 | |
| MAE | 0.0079 | 0.0070 | 0.0064 | 0.0086 | 0.0059 | 0.0066 | 0.0068 | 0.0104 | 0.0078 | 0.0139 | 0.0081 |
| 0.9372 | 0.9344 | 0.9403 | 0.9353 | 0.9468 | 0.9371 | 0.9363 | 0.9125 | 0.9312 | 0.9341 | 0.9345 |
| Weight | 0.1 | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | 0.7 | 0.8 | 0.9 | 1 |
|---|---|---|---|---|---|---|---|---|---|---|
| 0.9051 | 0.9144 | 0.9089 | 0.9007 | 0.8819 | 0.9018 | 0.9029 | 0.9011 | 0.8951 | 0.9061 | |
| MAE | 0.0075 | 0.0066 | 0.0072 | 0.0074 | 0.0108 | 0.0078 | 0.0076 | 0.0080 | 0.0076 | 0.0075 |
| 0.9248 | 0.9332 | 0.9297 | 0.9212 | 0.9121 | 0.9237 | 0.9241 | 0.9231 | 0.9200 | 0.9258 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
He, J.; Zhao, L.; Hu, W.; Zhang, G.; Wu, J.; Li, X. TCM-Net: Mixed Global–Local Learning for Salient Object Detection in Optical Remote Sensing Images. Remote Sens. 2023, 15, 4977. https://doi.org/10.3390/rs15204977
He J, Zhao L, Hu W, Zhang G, Wu J, Li X. TCM-Net: Mixed Global–Local Learning for Salient Object Detection in Optical Remote Sensing Images. Remote Sensing. 2023; 15(20):4977. https://doi.org/10.3390/rs15204977
Chicago/Turabian StyleHe, Junkang, Lin Zhao, Wenjing Hu, Guoyun Zhang, Jianhui Wu, and Xinping Li. 2023. "TCM-Net: Mixed Global–Local Learning for Salient Object Detection in Optical Remote Sensing Images" Remote Sensing 15, no. 20: 4977. https://doi.org/10.3390/rs15204977
APA StyleHe, J., Zhao, L., Hu, W., Zhang, G., Wu, J., & Li, X. (2023). TCM-Net: Mixed Global–Local Learning for Salient Object Detection in Optical Remote Sensing Images. Remote Sensing, 15(20), 4977. https://doi.org/10.3390/rs15204977

