HR2SIOD-CL: A Compressed Learning Framework for Object Detection in High-Resolution Remote Sensing Images
Highlights
- We introduce an end-to-end compressed learning framework named HR2SIOD-CL, which enables direct object detection from compressed measurements of high-resolution remote sensing images without explicit image reconstruction.
- Compared with conventional “reconstruct-then-infer” pipelines, the proposed measurement-domain object detection paradigm drastically reduces computational and memory overheads while achieving comparable or higher detection accuracy.
- This work challenges the traditional “reconstruct-then-infer” paradigm, revealing that explicit pixel-level reconstruction is not a prerequisite for achieving competitive performance in high-resolution remote sensing object detection.
- The proposed reconstruction-free detection method offers a promising alternative for high-resolution remote sensing object detection under resource-constrained scenarios.
Abstract
1. Introduction
- From an algorithmic rather than physical implementation perspective, a systematic study is conducted on direct object detection from CS measurements of high-resolution RSIs, focusing on whether explicit image reconstruction is indispensable when detection is the primary goal in the image post-acquisition compression scenario.
- We propose an end-to-end CL framework dubbed HR2SIOD-CL for high-resolution remote sensing object detection, which enables direct object detection from CS measurements via joint optimization of the adaptive sampling module and measurement-domain detection backbone.
- An entropy-driven and content-aware adaptive sampling strategy is introduced to allocate sensing resources by local information distribution, enhancing the robustness of detection under a limited measurement budget. A measurement-domain detection backbone is designed by incorporating a Measurement-to-Embedding (M2E) stem, four hierarchical Stages, and downsampling to capture multi-scale discriminative semantic features from CS measurements while maintaining scalability to high-resolution RSIs.
- We provide comprehensive evaluations for various methods like the image-domain, CS reconstruction-based and measurement-domain detection schemes on public remote sensing benchmarks in terms of detection accuracy as well as computational and memory costs, demonstrating the excellent trade-off between efficiency and accuracy of direct measurement-domain detection.
2. Related Work
2.1. Object Detection in Remote Sensing Images
2.2. Compressed Sensing
2.2.1. Non-Adaptive Sampling
2.2.2. Adaptive Sampling
2.3. Compressed Learning
3. Methodology
3.1. Problem Formulation and Design Requirements
- Scalability under high spatial resolution. The detection framework should sustain high computational efficiency across varying image spatial sizes. Although various DL-based reconstruction methods [16,17,18] achieve promising performance on small and medium-sized images, they encounter severe resource bottlenecks when processing high-resolution inputs. Traditional reconstruction-based pipelines exhibit weak scalability on high-resolution images due to heavy overhead in both computation and memory consumption. Accordingly, a practical detection framework should avoid intermediate representations that incur heavy computational overhead proportional to spatial dimensions.
- Task-oriented representations over pixel fidelity. For scenarios that prioritize final detection outputs without concern with high-quality reconstruction, extracting semantic features for object recognition and localization is of vital importance. Meanwhile mandatory image reconstruction introduces a redundant intermediate process that incurs extra computational costs and may even trigger privacy leakage.
- Constrained and structured encoding. In real-world remote sensing systems, the measurement budget is restricted by transmission bandwidth and terminal processing power. To ensure scalability and parallelism, the block-wise sampling mechanism is generally adopted for high-resolution images. As a result, the detection framework should perform efficiently under a limited measurement budget while effectively managing non-uniform information distribution across different image regions.
- Decoupled sampling and detection. In practical applications, the sampling module and the detection module may run on distinct platforms or terminals. Therefore, although these two modules are jointly optimized during training, they should be decoupled in the inference stage to enable flexible deployment.
3.2. Overview of the HR2SIOD-CL Framework
3.3. Entropy-Driven Content-Aware Adaptive Sampling
3.4. Measurement-Domain Detection Backbone
3.5. Reconstruction-Based Detection Baseline
4. Experimental Results
4.1. Datasets
4.2. Experimental Settings
4.3. Effects of Different Block Size
4.4. Reconstruction Performance Validation for the Proposed Baseline
4.5. Object Detection Based on Original Images, Reconstructed Images, and CS Measurements
4.6. Cross-Architecture Generalization
4.7. Comparison with State-of-the-Art Detection Methods
4.8. Ablation Studies
4.8.1. Effect of the Saliency Metrics
4.8.2. Sensitivity Analysis of Minimum Guaranteed Ratio
4.8.3. Ablation Study on the Sampling Strategy and Detection Backbone
4.9. Analysis of Failure Cases and Model Limitations
5. Discussion
5.1. Necessity of Explicit Image Reconstruction
5.2. Mechanism Analysis of Core Modules
5.3. Trade-Off Between Detection Accuracy and Computational Efficiency
5.4. Model Limitations
5.5. Comparison with Existing CL-Based Methods
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Lin, A.; Sun, X.; Wu, H.; Luo, W.; Wang, D.; Zhong, D.; Wang, Z.; Zhao, L.; Zhu, J. Identifying urban building function by integrating remote sensing imagery and POI data. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2021, 14, 8864–8875. [Google Scholar] [CrossRef] [Scilit]
- Wang, B.; Xu, Y.; Wu, Z.; Zhan, T.; Wei, Z. Spatial-spectral local domain adaption for cross domain few shot hyperspectral images classification. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5539515. [Google Scholar] [CrossRef] [Scilit]
- Hatić, D.; Polushko, V.; Rauhut, M.; Hagen, H. Post-disaster building damage assessment: Multi-class object detection vs. Object localization and classification. Remote Sens. 2025, 17, 3957. [Google Scholar] [CrossRef] [Scilit]
- Shimoni, M.; Haelterman, R.; Perneel, C.J.I.G.; Magazine, R.S. Hypersectral imaging for military and security applications: Combining myriad processing and sensing techniques. IEEE Geosci. Remote Sens. Mag. 2019, 7, 101–117. [Google Scholar] [CrossRef] [Scilit]
- Ding, J.; Xue, N.; Xia, G.S.; Bai, X.; Yang, W.; Yang, M.Y.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; et al. Object detection in aerial images: A large-scale benchmark and challenges. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 7778–7796. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, G.; Bai, Z.; Liu, Z. Texture-semantic collaboration network for ORSI salient object detection. IEEE Trans. Circuits Syst. II Express Briefs 2024, 71, 2464–2468. [Google Scholar] [CrossRef] [Scilit]
- Zhang, G.; Zhang, W.; Li, W.; Wang, L.; Cui, H. A dynamic attention mechanism for object detection in road or strip environments. Vis. Comput. 2025, 41, 4171–4181. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Zhou, P.; Han, J. Learning rotation-invariant convolutional neural networks for object detection in VHR optical remote sensing images. IEEE Trans. Geosci. Remote Sens. 2016, 54, 7405–7415. [Google Scholar] [CrossRef] [Scilit]
- Li, K.; Wan, G.; Cheng, G.; Meng, L.; Han, J. Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS J. Photogramm. Remote Sens. 2020, 159, 296–307. [Google Scholar] [CrossRef] [Scilit]
- Lunga, D.; Gerrand, J.; Yang, L.; Layton, C.; Stewart, R. Apache spark accelerated deep learning inference for large scale satellite image analytics. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 13, 271–283. [Google Scholar] [CrossRef] [Scilit]
- Donoho, D.L. Compressed sensing. IEEE Trans. Inf. Theory 2006, 52, 1289–1306. [Google Scholar] [CrossRef] [Scilit]
- Li, S.L.; Li, K.; Zhang, F.; Zhang, L.; Xiao, L.L.; Huang, D.P. Innovative remote sensing imaging method based on compressed sensing. Opt. Laser Technol. 2014, 63, 83–89. [Google Scholar] [CrossRef] [Scilit]
- Li, F.; Xin, L.; Liu, Y.; Fu, J.; Liu, Y.; Guo, Y. High efficient optical remote sensing images acquisition for nano-satellite framework. In Proceedings of the Sensors, Systems, and Next-Generation Satellites XXI, Warsaw, Poland, 11–14 September 2017; SPIE: Bellingham, WA, USA, 2017; Volume 10423, pp. 341–347. [Google Scholar]
- Ghahremani, M.; Liu, Y.; Yuen, P.; Behera, A. Remote sensing image fusion via compressive sensing. ISPRS J. Photogramm. Remote Sens. 2019, 152, 34–48. [Google Scholar] [CrossRef] [Scilit]
- Xiao, S.; Zhang, Y.; Chang, X. Ship detection based on compressive sensing measurements of optical remote sensing scenes. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 8632–8649. [Google Scholar] [CrossRef] [Scilit]
- Shi, W.; Jiang, F.; Liu, S.; Zhao, D. Image compressed sensing using convolutional neural network. IEEE Trans. Image Process. 2020, 29, 375–388. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Song, J.; Mou, C.; Wang, S.; Ma, S.; Zhang, J. Optimization-inspired cross-attention Transformer for compressive sensing. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 6174–6184. [Google Scholar]
- Shen, M.; Gan, H.; Ma, C.; Ning, C.; Li, H.; Liu, F. MTC-CSNet: Marrying Transformer and convolution for image compressed sensing. IEEE Trans. Cybern. 2024, 54, 4949–4961. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sankaranarayanan, A.C.; Turaga, P.K.; Baraniuk, R.G.; Chellappa, R. Compressive acquisition of dynamic scenes. In Proceedings of the 11th European Conference on Computer Vision (ECCV), Heraklion, Greece, 5–11 September 2010; pp. 129–142. [Google Scholar]
- Hahn, J.; Rosenkranz, S.; Zoubir, A.M. Adaptive compressed classification for hyperspectral imagery. In Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy, 4–9 May 2014; pp. 1020–1024. [Google Scholar]
- Tang, S.; Cheang, C.F.; Yu, X.; Liang, Y.; Feng, Q.; Chen, Z. TransCS-Net: A hybrid transformer-based privacy-protecting network using compressed sensing for medical image segmentation. Biomed. Signal Process. Control 2023, 86, 105131. [Google Scholar] [CrossRef] [Scilit]
- Calderbank, R.; Jafarpour, S.; Schapire, R. Compressed Learning: Universal Sparse Dimensionality Reduction and Learning in the Measurement Domain. Preprint, 2009. Available online: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.481.8129&rep=rep1&type=pdf (accessed on 4 May 2026).
- Davenport, M.A.; Duarte, M.F.; Wakin, M.B.; Laska, J.N.; Takhar, D.; Kelly, K.F.; Baraniuk, R.G. The smashed filter for compressive classification and target recognition. In Proceedings of the SPIE Computational Imaging, San Jose, CA, USA, 29–31 January 2007; pp. 142–153. [Google Scholar]
- Cui, Y.; Xu, W.; Wang, Y.; Lin, J.; Lu, L. Performance bounds of compressive classification under perturbation. Signal Process. 2021, 180, 107855. [Google Scholar] [CrossRef] [Scilit]
- Wimalajeewa, T.; Chen, H.; Varshney, P.K. Performance limits of compressive sensing-based signal classification. IEEE Trans. Signal Process. 2012, 60, 2758–2770. [Google Scholar] [CrossRef] [Scilit]
- Lohit, S.; Kulkarni, K.; Turaga, P. Direct inference on compressive measurements using convolutional neural networks. In Proceedings of the IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA, 25–28 September 2016; pp. 1913–1917. [Google Scholar]
- Zisselman, E.; Adler, A.; Elad, M. Compressed learning for image classification: A deep neural network approach. In Handbook of Numerical Analysis; Elsevier: Amsterdam, The Netherlands, 2018; Volume 19, pp. 3–17. [Google Scholar]
- Duan, Z.; Ma, Z.; Zhu, F. Unified architecture adaptation for compressed domain semantic inference. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 4108–4121. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xiao, D.; Zhang, M.; Zhang, M.; Chen, L. CFMVOR: Federated multi-view 3D object recognition based on compressed learning. In Proceedings of the 7th Chinese Conference on Pattern Recognition and Computer Vision (PRCV), Urumqi, China, 18–20 October 2024; Springer: Singapore, 2025; pp. 280–293. [Google Scholar]
- Zhang, X.; Yu, X.; Feng, J.; Chen, W.; Li, Y. TriNeXt: An efficient three-path fusion module for multi-scale feature enhancement in object detection. Sci. Prog. 2025, 108, 00368504251395116. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, J.; Yu, Y.; Cheng, J.; Li, J.; Tang, J. PillarBAPI: Enhancing pillar-based 3D object detection through attentive pseudo-image feature extraction. Multimed. Syst. 2025, 31, 263. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Shi, S.; Wu, Y.; Lin, W.; Bai, Z. Lightweight ORSI salient object detection via frequency and mutual assistance attention. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5617112. [Google Scholar] [CrossRef] [Scilit]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar]
- Huang, Y.; Jiao, D.; Huang, X.; Tang, T.; Gui, G. A hybrid CNN-Transformer network for object detection in optical remote sensing images: Integrating local and global feature fusion. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 241–254. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Tian, P.; Song, R.; Xu, H.; Li, Y.; Du, Q. PCViT: A pyramid convolutional vision Transformer detector for object detection in remote-sensing imagery. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5608115. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, real-Time object detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-time end-to-end object detection. In Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024), Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
- Tian, Y.; Ye, Q.; Doermann, D. Yolov12: Attention-centric real-time object detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with Transformers. In Proceedings of the European Conference on Computer Vision (ECCV), Cham, Switzerland, 23–28 August 2020; pp. 213–229. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar]
- Tropp, J.A.; Gilbert, A.C. Signal recovery from random measurements via orthogonal matching pursuit. IEEE Trans. Inform. Theory 2007, 53, 4655–4666. [Google Scholar] [CrossRef] [Scilit]
- Blumensath, T. Accelerated iterative hard thresholding. Signal Process. 2012, 92, 752–756. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Rao, B.D. Extension of SBL algorithms for the recovery of block sparse signals with intra-block correlation. IEEE Trans. Signal Process. 2013, 61, 2009–2015. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Zhao, C.; Gao, W. Optimization-Inspired Compact Deep Compressive Sensing. IEEE J. Sel. Top. Signal Process. 2020, 14, 765–774. [Google Scholar] [CrossRef] [Scilit]
- Shen, M.; Gan, H.; Ning, C.; Hua, Y.; Zhang, T. TransCS: A Transformer-based hybrid architecture for image compressed sensing. IEEE Trans. Image Process. 2022, 31, 6991–7005. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Song, J.; Chen, B.; Zhang, J. Deep memory-augmented proximal unrolling network for compressive sensing. Int. J. Comput. Vis. 2023, 131, 1477–1496. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Hua, Z.; Li, Y.; Chen, Y.; Zhou, Y. AMS-Net: Adaptive multi-scale network for image compressive sensing. IEEE Trans. Multimed. 2023, 25, 5676–5689. [Google Scholar] [CrossRef] [Scilit]
- Chen, B.; Zhang, J. Content-aware scalable deep compressed sensing. IEEE Trans. Image Process. 2022, 31, 5412–5426. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, K.; Hua, Z.; Li, Y.; Zhang, Y.; Zhou, Y. Uformer-ICS: A U-shaped Transformer for image compressive sensing service. IEEE Trans. Serv. Comput. 2024, 17, 2974–2988. [Google Scholar] [CrossRef] [Scilit]
- Qiu, C.; Hu, X. AdaCS: Adaptive compressive sensing with restricted isometry property-based error-clamping. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 4702–4719. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mdrafi, R.; Gurbuz, A.C. Compressed classification from learned measurements. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada, 11–17 October 2021; pp. 4038–4047. [Google Scholar]
- Tran, D.T.; Yamaç, M.; Degerli, A.; Gabbouj, M.; Iosifidis, A. Multilinear compressive learning. IEEE Trans. Neural Netw. Learn. Syst. 2021, 32, 1512–1524. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mou, C.; Zhang, J. TransCL: Transformer Makes Strong and Flexible Compressive Learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 5236–5251. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, X.; Gong, Y.; Wu, W.; Zhu, S.; Zhao, Y. CSDet: A compressed sensing object detection architecture with lightweight networks. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 2355–2368. [Google Scholar] [CrossRef] [Scilit]
- van den Oord, A.; Vinyals, O.; Kavukcuoglu, K. Neural discrete representation learning. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Bengio, Y.; Léonard, N.; Courville, A. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv 2013, arXiv:1308.3432. [Google Scholar]
- Shi, W.; Caballero, J.; Huszár, F.; Totz, J.; Aitken, A.P.; Bishop, R.; Rueckert, D.; Wang, Z. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 1874–1883. [Google Scholar]
- Ding, X.; Zhang, X.; Ma, N.; Han, J.; Ding, G.; Sun, J. RepVGG: Making VGG-style convnets great again. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 13733–13742. [Google Scholar]
- Chu, X.; Li, L.; Zhang, B. Make RepVGG greater again: A quantization-aware approach. arXiv 2022, arXiv:2212.01593. [Google Scholar]
- Yu, W.; Luo, M.; Zhou, P.; Si, C.; Zhou, Y.; Wang, X.; Feng, J.; Yan, S. Metaformer is actually what you need for vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 10819–10829. [Google Scholar]
- Li, Y.; Hu, J.; Wen, Y.; Evangelidis, G.; Salahi, K.; Wang, Y.; Tulyakov, S.; Ren, J. Rethinking vision Transformers for MobileNet size and speed. arXiv 2022, arXiv:2212.08059. [Google Scholar]
- Mehta, S.; Rastegari, M. Separable self-attention for mobile vision transformers. arXiv 2022, arXiv:2206.02680. [Google Scholar]
- Han, K.; Wang, Y.; Tian, Q.; Guo, J.; Xu, C.; Xu, C. GhostNet: More features from cheap operations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 1580–1589. [Google Scholar]
- Tang, Y.; Han, K.; Guo, J.; Xu, C.; Xu, C.; Wang, Y. GhostNetv2: Enhance cheap operation with long-range attention. In Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022), New Orleans, LA, USA, 28 November–9 December 2022. [Google Scholar]
- Finder, S.E.; Amoyal, R.; Treister, E.; Freifeld, O. Wavelet convolutions for large receptive fields. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 363–380. [Google Scholar]
- Hu, J.; Shen, L.; Albanie, S.; Sun, G.; Wu, E. Squeeze-and-excitation networks. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 2011–2023. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vasu, P.K.A.; Gabriel, J.; Zhu, J.; Tuzel, O.; Ranjan, A. FastViT: A fast hybrid vision Transformer using structural reparameterization. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 5785–5795. [Google Scholar]
- Shaker, A.; Maaz, M.; Rasheed, H.; Khan, S.; Yang, M.H.; Khan, F.S. SwiftFormer: Efficient additive attention for Transformer-based real-time mobile vision applications. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 17425–17436. [Google Scholar]












| Block Size | Params. | MACs. | SR | |||
|---|---|---|---|---|---|---|
| 0.25 | 0.1 | 0.05 | 0.018 | |||
| 8 × 8 | 12.29 K | 402.65 M | 88.2/53.2 | 86.3/51.7 | 84.3/50.4 | 83.3/49.1 |
| 16 × 16 | 196.61 K | 1.61 G | 88.9/53.8 | 86.5/51.0 | 79.8/46.7 | 79.9/45.6 |
| 32 × 32 | 3.15 M | 6.44 G | 86.3/51.6 | 84.2/49.7 | 79.0/46.8 | 81.0/45.7 |
| 64 × 64 | 50.33 M | 25.77 G | 77.4/44.0 | 78.9/44.2 | 76.4/43.3 | 76.5/42.6 |
| SR | CSNet+ | Ours | ||
|---|---|---|---|---|
| PSNR | SSIM | PSNR | SSIM | |
| 0.018 | 23.78 | 0.5290 | 23.94 | 0.5466 |
| 0.05 | 25.82 | 0.6286 | 26.25 | 0.6758 |
| 0.1 | 25.42 | 0.7267 | 27.47 | 0.7266 |
| 0.25 | 29.59 | 0.8339 | 33.09 | 0.9110 |
| Average | 26.15 | 0.6796 | 27.69 | 0.7150 |
| Group | SR | Params. | MACs. | Memory | NWPU VHR-10 | DIOR | ||
|---|---|---|---|---|---|---|---|---|
| Image domain | / | 14.87 M | 32.34 G | 4.7 G | 90.8 | 57.1 | 74.3 | 51.3 |
| Reconstruction-based | 0.25 | 15.31 M | 470.02 G | 13.29 G | 86.5 | 53.0 | 72.2 | 49.6 |
| 0.1 | 85.1 | 51.5 | 72.2 | 49.4 | ||||
| 0.05 | 83.4 | 47.5 | 71.1 | 48.6 | ||||
| 0.018 | 82.9 | 45.5 | 67.2 | 44.6 | ||||
| Measurement domain | 0.25 | 14.89 M (−0.42) | 32.74 G (−437.28) | 4.75 G (−8.54) | 88.2 | 53.2 | 73.5 | 51.2 |
| 0.1 | 86.3 | 51.7 | 73.0 | 50.2 | ||||
| 0.05 | 84.3 | 50.4 | 72.1 | 49.5 | ||||
| 0.018 | 83.3 | 49.1 | 70.3 | 47.2 | ||||
| Group | Detector | SR | Params. | MACs. | Memory | NWPU VHR-10 | DIOR | ||
|---|---|---|---|---|---|---|---|---|---|
| Image domain | YOLO12s | / | 14.87 M | 32.34 G | 4.7 G | 93.5 | 62.2 | 83.0 | 64.4 |
| RT-DETR-L | 95.7 | 64.3 | 82.6 | 64.7 | |||||
| Reconstruction-based | YOLO12s | 0.25 | 15.31 M | 470.02 G | 13.29 G | 86.9 | 53.9 | 81.6 | 62.4 |
| 0.1 | 84.8 | 53.5 | 80.7 | 61.7 | |||||
| 0.05 | 85.2 | 52.9 | 79.5 | 60.2 | |||||
| 0.018 | 81.1 | 49.5 | 78.3 | 59.0 | |||||
| RT-DETR-L | 0.25 | 15.31 M | 470.02 G | 13.29 G | 92.0 | 60.0 | 80.2 | 61.9 | |
| 0.1 | 91.8 | 60.4 | 81.0 | 62.9 | |||||
| 0.05 | 89.4 | 57.3 | 79.9 | 61.6 | |||||
| 0.018 | 87.2 | 55.0 | 77.3 | 58.8 | |||||
| Measurement domain | YOLO12s | 0.25 | 14.89 M | 32.74 G | 4.75 G | 86.6 | 54.0 | 81.7 | 63.4 |
| 0.1 | 86.0 | 54.0 | 81.0 | 62.3 | |||||
| 0.05 | 84.7 | 53.1 | 79.4 | 60.9 | |||||
| 0.018 | 81.5 | 48.8 | 78.1 | 59.1 | |||||
| RT-DETR-L | 0.25 | 14.89 M | 32.74 G | 4.75 G | 91.7 | 61.0 | 81.7 | 63.7 | |
| 0.1 | 90.7 | 59.8 | 81.1 | 62.9 | |||||
| 0.05 | 89.8 | 57.5 | 79.8 | 61.3 | |||||
| 0.018 | 87.1 | 56.5 | 77.7 | 58.9 | |||||
| Method | NWPU VHR-10 | DIOR | ||
|---|---|---|---|---|
| GhostNetV2 1.0× (NeurIPS 2022) [65] | 84.7 | 46.3 | 72.7 | 48.4 |
| FastViT-S12 (ICCV 2023) [68] | 87.9 | 49.4 | 73.8 | 48.1 |
| SwiftFormer-S (ICCV2023) [69] | 70.9 | 42.4 | 74.5 | 52.3 |
| EfficientFormer V2 S1 (ICCV 2023) [62] | 91.3 | 58.0 | 75.2 | 52.9 |
| RT-DETR-L (CVPR 2024) [41] | 92.2 | 63.4 | 81.3 | 63.3 |
| YOLO12s (NeurIPS 2025) [39] | 94.7 | 66.6 | 82.5 | 64.4 |
| Ours-25 | 91.7 | 61.0 | 81.7 | 63.7 |
| Ours-10 | 90.7 | 59.8 | 81.1 | 62.9 |
| Ours-5 | 89.8 | 57.5 | 79.8 | 61.3 |
| Ours-1.8 | 87.1 | 56.5 | 77.7 | 58.9 |
| Metric | SR = 0.1 | SR = 0.05 | ||
|---|---|---|---|---|
| Mean | 86.1 | 50.0 | 83.9 | 48.3 |
| Max | 86.6 | 50.5 | 84.5 | 49.5 |
| Variance | 85.0 | 51.6 | 84.0 | 49.6 |
| Entropy (Ours) | 86.3 | 51.7 | 84.3 | 50.4 |
| SR = 0.1 | SR = 0.05 | SR = 0.018 | ||||
|---|---|---|---|---|---|---|
| 0.3 | 86.4 | 50.7 | 85.0 | 49.5 | 82.3 | 47.4 |
| 0.7 | 85.8 | 50.4 | 84.7 | 46.5 | 82.8 | 47.8 |
| 0.5 (Ours) | 86.3 | 51.7 | 84.3 | 50.4 | 83.3 | 49.1 |
| SR | Adaptive Sampling | Non-Adaptive Sampling | Our Backbone | GhostNetV2 | NWPU VHR-10 | DIOR | ||
|---|---|---|---|---|---|---|---|---|
| 0.25 | ✓ | ✓ | 88.2 | 53.2 | 73.5 | 51.2 | ||
| ✓ | ✓ | 87.7 | 52.1 | 71.9 | 50.3 | |||
| ✓ | ✓ | 80.2 | 42.7 | 68.6 | 44.7 | |||
| ✓ | ✓ | 79.3 | 41.4 | 67.7 | 43.9 | |||
| 0.1 | ✓ | ✓ | 86.3 | 51.7 | 73.0 | 50.2 | ||
| ✓ | ✓ | 83.9 | 48.1 | 70.6 | 48.4 | |||
| ✓ | ✓ | 78.5 | 40.3 | 66.7 | 42.1 | |||
| ✓ | ✓ | 77.1 | 40.0 | 65.9 | 41.6 | |||
| Method | Domain | Params (M) | MACs (G) | Inference Time (ms) |
|---|---|---|---|---|
| YOLO12s | Image | 9.28 | 34.01 | 11.82 |
| RT-DETR-L | 32.97 | 137.16 | 29.93 | |
| Ours + YOLO12s | Reconstruction-based | 19.78 | 480.80 | 107.13 |
| Measurement | 19.36 | 42.18 | 72.02 | |
| Ours + RT-DETR-L | Reconstruction-based | 35.52 | 551.22 | 119.78 |
| Measurement | 35.11 | 112.60 | 82.40 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jing, Y.; Wu, X.; Wang, H.; Wang, K.; You, D.; Kan, H. HR2SIOD-CL: A Compressed Learning Framework for Object Detection in High-Resolution Remote Sensing Images. Remote Sens. 2026, 18, 2851. https://doi.org/10.3390/rs18172851
Jing Y, Wu X, Wang H, Wang K, You D, Kan H. HR2SIOD-CL: A Compressed Learning Framework for Object Detection in High-Resolution Remote Sensing Images. Remote Sensing. 2026; 18(17):2851. https://doi.org/10.3390/rs18172851
Chicago/Turabian StyleJing, Yanhao, Xiangjun Wu, Hui Wang, Kunshu Wang, Datao You, and Haibin Kan. 2026. "HR2SIOD-CL: A Compressed Learning Framework for Object Detection in High-Resolution Remote Sensing Images" Remote Sensing 18, no. 17: 2851. https://doi.org/10.3390/rs18172851
APA StyleJing, Y., Wu, X., Wang, H., Wang, K., You, D., & Kan, H. (2026). HR2SIOD-CL: A Compressed Learning Framework for Object Detection in High-Resolution Remote Sensing Images. Remote Sensing, 18(17), 2851. https://doi.org/10.3390/rs18172851

