FBR-DETR: An Efficient End-to-End Network for Real-Time Small-Object Detection in UAV Imagery
Highlights
- FBR-DETR integrates frequency-domain inverse convolution, binary attention, and re-parameterized cross-scale fusion to enhance real-time small-object detection in UAV imagery.
- Experiments on VisDrone2019-DET and HIT-UAV show that FBR-DETR outperforms the RT-DETR baseline in detection accuracy while reducing the parameter count by 30.8% and boosting inference speed.
- The three-stage collaborative optimization framework effectively strengthens the feature representation of dense, weak-texture small objects in complex aerial backgrounds.
- FBR-DETR achieves a favorable accuracy–efficiency trade-off with a lightweight architecture, which is suitable for real-time deployment on resource-constrained UAV-embedded platforms.
Abstract
1. Introduction
- We propose a frequency-domain inverse-convolution-enhanced feature extraction network (FICE-Net), which introduces Converse2D frequency-domain inverse convolution into the backbone through the FICE Block, thereby enhancing the model’s perception of global spectral structures and fine-grained spatial features and improving small-object feature representation at the source.
- We develop a binary attention-based intra-scale feature interaction module (Binary-AIFI), which binarizes the query and key matrices in the attention mechanism to preserve global contextual modelling capability while controlling additional computational cost.
- We introduce a re-parameterized cross-scale feature fusion module (RCFF), which incorporates GCConv into the multi-scale fusion path, enhances cross-scale feature alignment and representation during training, and maintains efficient inference through structural re-parameterization, thereby balancing detection performance and real-time capability.
2. Related Work
2.1. Feature Extraction Methods
2.2. Attention Mechanism
2.3. Multi-Scale Feature Fusion
3. Methodology
3.1. Overall Framework of FBR-DETR
3.2. Frequency-Domain Inverse-Convolution-Enhanced Feature Extraction Network (FICE-Net)
3.3. Binary Attention-Based Intra-Scale Feature Interaction Module (Binary-AIFI)
3.4. Re-Parameterized Cross-Scale Feature Fusion Module (RCFF)
4. Experiments
4.1. Datasets
4.2. Experimental Details and Evaluation Metrics
5. Results and Analysis
5.1. Ablation Experiments
5.1.1. Ablation Study on the Effectiveness of the Proposed Modules
5.1.2. Ablation Study on the Influence of the Proposed Modules on Different Object Categories
5.2. Comparison Experiments
5.2.1. Comparison with Different Network Models
5.2.2. Comparison of Different Lightweight Backbone Convolutions
5.2.3. Comparison of Different Attention Mechanisms
5.2.4. Supplementary Comparative Verification of Core Modules
5.3. Visualization Analysis
5.3.1. Visual Analysis on the VisDrone2019-DET Dataset
5.3.2. Visual Analysis on the HIT-UAV Dataset
5.3.3. Heatmap Visual Analysis
5.4. Cross-Scene Visual Case Analysis
5.5. Edge Deployment and Real-Time Performance Analysis
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Wang, B.; Sui, H.; Ma, G.; Zhou, Y.; Zhou, M. GMODet: A Real-Time Detector for Ground-Moving Objects in Optical Remote Sensing Images with Regional Awareness and Semantic–Spatial Progressive Interaction. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5605623. [Google Scholar] [CrossRef] [Scilit]
- Xiao, Y.; Wang, J.; Zhao, Z.; Jiang, B.; Li, C.; Tang, J. UAV Video Vehicle Detection: Benchmark and Baseline. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5609814. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Zhao, D.; Yuan, B.; Jiang, Z. RescueADI: Adaptive disaster interpretation in remote sensing images with autonomous agents. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5611814. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Liu, C.; Li, W.; Xu, Q.; Deng, H. DTSSNet: Dynamic training sample selection network for UAV object detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5902516. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Li, W.; Hong, D.; Tao, R.; Du, Q. Deep learning for unmanned aerial vehicle-based object detection and tracking: A survey. IEEE Geosci. Remote Sens. Mag. 2021, 10, 91–124. [Google Scholar] [CrossRef] [Scilit]
- Ma, W.; Wang, X.; Zhu, H.; Yang, X.; Yi, X.; Jiao, L. Significant feature elimination and sample assessment for remote sensing small objects’ detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5615115. [Google Scholar] [CrossRef] [Scilit]
- Zhang, T.; Zhang, X.; Zhu, X.; Wang, G.; Han, X.; Tang, X.; Jiao, L. Multistage enhancement network for tiny object detection in remote sensing images. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5611512. [Google Scholar] [CrossRef] [Scilit]
- Shen, F.; Jiang, X.; He, X.; Ye, H.; Wang, C.; Du, X.; Li, Z.; Tang, J. Imagdressing-v1: Customizable virtual dressing. In Proceedings of the AAAI Conference on Artificial Intelligence; The Association for the Advancement of Artificial Intelligence: Washington, DC, USA, 2025; pp. 6795–6804. [Google Scholar]
- Xin, C.; Hartel, A.; Kasneci, E. Dart: An automated end-to-end object detection pipeline with data diversification, open-vocabulary bounding box annotation, pseudo-label review, and model training. Expert Syst. Appl. 2024, 258, 125124. [Google Scholar] [CrossRef] [Scilit]
- Hong, D.; Gao, L.; Hang, R.; Zhang, B.; Chanussot, J. Deep encoder–decoder networks for classification of hyperspectral and LiDAR data. IEEE Geosci. Remote Sens. Lett. 2020, 19, 5500205. [Google Scholar] [CrossRef] [Scilit]
- Cai, Z.; Vasconcelos, N. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2018; pp. 6154–6162. [Google Scholar]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2017; pp. 2961–2969. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. Efficientdet: Scalable and efficient object detection. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2020; pp. 10781–10790. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2016; pp. 779–788. [Google Scholar]
- Wang, H.; Shi, J.; Karimian, H.; Liu, F.; Wang, F. YOLOSAR-Lite: A lightweight framework for real-time ship detection in SAR imagery. Int. J. Digit. Earth 2024, 17, 2405525. [Google Scholar] [CrossRef] [Scilit]
- Liang, S.; Wang, X.; Sun, J.; Liu, H.; Yang, H. EDM-Net: A Multi-Scale Network for Object Detection in Remote Sensing Images. Sensors 2026, 26, 3927. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2020; pp. 213–229. [Google Scholar]
- Wang, H.; Shi, J.; Karimian, H.; Wang, F.; Javed, F.; Liu, B.; Shi, S.; Li, Z.; Tao, Y. CitrusNet: A vision transformer-CNN approach for citrus detection from multi-source imagery with multi-scale feature integration. Comput. Electron. Agric. 2026, 241, 111260. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. Detrs beat yolos on real-time object detection. In Proceedings of the 2024 IEEE/CVF Conference On Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 16965–16974. [Google Scholar]
- Zhang, H.; Zhang, H.; Liu, K.; Gan, Z.; Zhu, G.-N. UAV-DETR: Efficient end-to-end object detection for unmanned aerial vehicle imagery. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2025; pp. 15143–15149. [Google Scholar]
- Hu, L.; Yuan, J.; Cheng, B.; Xu, Q. CSFPR-RTDETR: Real-Time Small Object Detection Network for UAV Images Based on Cross Spatial Frequency Domain and Position Relation. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5638219. [Google Scholar] [CrossRef] [Scilit]
- Xia, Y.; Liu, C.; Xiang, T.; Tu, Z. EFSI-DETR: Efficient Frequency-Semantic Integration for Real-Time Small Object Detection in UAV Imagery. arXiv 2026, arXiv:2601.18597. [Google Scholar]
- Ma, S.; Zhang, Y.; Peng, L.; Sun, C.; Ding, B.; Zhu, Y. OWRT-DETR: A novel real-time transformer network for small-object detection in open-water search and rescue from UAV aerial imagery. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4205313. [Google Scholar] [CrossRef] [Scilit]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Lin, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y. VisDrone-DET2019: The vision meets drone object detection in image challenge results. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); IEEE: New York, NY, USA, 2019; pp. 213–226. [Google Scholar]
- Suo, J.; Wang, T.; Zhang, X.; Chen, H.; Zhou, W.; Shi, W. HIT-UAV: A high-altitude infrared thermal dataset for Unmanned Aerial Vehicle-based object detection. Sci. Data 2023, 10, 227. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, Z.; Zhang, Y.; Zhou, J.; Li, D.; Cai, H.; Wang, N.; Zhang, J.; Xue, Y.; Jiang, H.; Lv, X. UIT-GAN: A cross-modal generative model for ultrasonic infrared thermography detection of impact damage in composites. Compos. Part B Eng. 2026, 313, 113389. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Chen, H. HPS-DETR: Enhancing small object detection with lightweight feature extraction and transformer integration. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5937420. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Lang, C.; Wu, M.; Xie, X.; Yao, X.; Han, J. Feature enhancement network for object detection in optical remote sensing images. J. Remote Sens. 2021, 2021, 9805389. [Google Scholar] [CrossRef] [Scilit]
- Gao, T.; Xia, S.; Liu, M.; Zhang, J.; Chen, T.; Li, Z. Msnet: Multi-scale network for object detection in remote sensing images. Pattern Recognit. 2025, 158, 110983. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Shi, Y.; Hong, Q.; Jia, Y. A Scale-Aware Multi-Domain DETR for Small Object Detection in UAV Remote Sensing Imagery. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4421520. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Sun, H.; Gao, H.; Gao, L.; Zhang, B. Hyperspectral Remote Sensing Object Detection via Cross Domain Learning from Visible Images. IEEE Trans. Geosci. Remote Sens. 2026, 63, 5510518. [Google Scholar] [CrossRef] [Scilit]
- Wang, D.; Zhuang, L.; Gao, L.; Sun, X.; Zhao, X.; Plaza, A. Sliding dual-window-inspired reconstruction network for hyperspectral anomaly detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5504115. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Ma, J.; Zhang, J. Elnet: An efficient and lightweight network for small object detection in UAV imagery. Remote Sens. 2025, 17, 2096. [Google Scholar] [CrossRef] [Scilit]
- Huang, X.; Liu, S.; Zhang, K.; Tai, Y.; Yang, J.; Zeng, H.; Zhang, L. Reverse Convolution and Its Applications to Image Restoration. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2025; pp. 10507–10516. [Google Scholar]
- Li, F.; Wang, X.; Wang, H.; Karimian, H.; Shi, J.; Zha, G. LMVMamba: A Hybrid U-Shape Mamba for Remote Sensing Segmentation with Adaptation Fine-Tuning. Remote Sens. 2025, 17, 3367. [Google Scholar] [CrossRef] [Scilit]
- Cai, Z.; Quan, S.; Wang, J.; Xing, S.; Su, X.; Li, Y.; Liu, L. MaOutCNN: A physical mechanism coupled MambaOut-CNN detection network with maiden-released interference polarized ship detection dataset. ISPRS J. Photogramm. Remote Sens. 2026, 236, 500–527. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2018; pp. 7132–7141. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. Cbam: Convolutional block attention module. In Proceedings of the 2018 European Conference on Computer Vision (ECCV); IEEE: New York, NY, USA, 2018; pp. 3–19. [Google Scholar]
- Li, M.; Gao, Y.; Guo, X.; Chen, Z.; Deng, L.; Dong, M.; Zhu, L. Edge-Semantic Synergy Network with Edge-Aware Attention for Infrared Small Target Detection. IEEE Trans. Geosci. Remote Sens. 2025, 64, 5000417. [Google Scholar] [CrossRef] [Scilit]
- Xiao, C.; Zhang, Z.; Zhang, L. BinaryAttention: One-bit QK-attention for vision and diffusion transformers. In Proceedings of the 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2026; pp. 12106–12117. [Google Scholar]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2014; pp. 740–755. [Google Scholar]
- Wang, J.; Yang, W.; Guo, H.; Zhang, R.; Xia, G.-S. Tiny object detection in aerial images. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR); IEEE: New York, NY, USA, 2021; pp. 3791–3798. [Google Scholar]
- Huang, J.; Zhang, F.; Zhu, H.; Yan, T. FSDETR: Frequency-Spatial Feature Enhancement for Small Object Detection. arXiv 2026, arXiv:2604.14884. [Google Scholar]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2017; pp. 2117–2125. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path aggregation network for instance segmentation. In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2018; pp. 8759–8768. [Google Scholar]
- Liu, H.-I.; Tseng, Y.-W.; Chang, K.-C.; Wang, P.-J.; Shuai, H.-H.; Cheng, W.-H. A denoising fpn with transformer r-cnn for tiny object detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4704415. [Google Scholar] [CrossRef] [Scilit]
- Bai, Y.; Song, C.; Wang, Y.; Li, P. Efficient object detection in remote sensing images based on feature weaving and redundancy suppression. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5607016. [Google Scholar] [CrossRef] [Scilit]
- Yang, G.; Wang, Y.; Shi, D.; Wang, Y. Golden cudgel network for real-time semantic segmentation. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2025; pp. 25367–25376. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Stoken, A.; Borovec, J.; Kwon, Y.; Michael, K.; Fang, J.; Yifu, Z.; Wong, C.; Montes, D. Ultralytics/Yolov5: V7. 0-Yolov5 Sota Realtime Instance Segmentation. Zenodo 2022. Available online: https://zenodo.org/records/7347926 (accessed on 24 August 2026).
- Varghese, R.; Sambath, M. Yolov8: A novel object detection algorithm with enhanced performance and robustness. In Proceedings of the 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS); IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. Yolov10: Real-time end-to-end object detection. Adv. Neural Inf. Process. Syst. 2024, 37, 107984–108011. [Google Scholar] [CrossRef] [Scilit]
- Khanam, R.; Hussain, M. Yolov11: An overview of the key architectural enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar]
- Sapkota, R.; Cheppally, R.H.; Sharda, A.; Karkee, M. YOLO26: Key architectural enhancements and performance benchmarking for real-time object detection. arXiv 2025, arXiv:2509.25164. [Google Scholar]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.-Y. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv 2022, arXiv:2203.03605. [Google Scholar]
- Yan, X.; Sun, S.; Zhu, H.; Hu, Q.; Ying, W.; Li, Y. DMF-YOLO: Dynamic multi-scale feature fusion network-driven small target detection in UAV aerial images. Remote Sens. 2025, 17, 2385. [Google Scholar] [CrossRef] [Scilit]
- Feijoo, D.; Benito, J.C.; Garcia, A.; Conde, M.V. Darkir: Robust low-light image restoration. In Proceedings of the 2025 Computer Vision and Pattern Recognition Conference (CVPR); IEEE: New York, NY, USA, 2025; pp. 10879–10889. [Google Scholar]
- Chen, L.; Gu, L.; Li, L.; Yan, C.; Fu, Y. Frequency dynamic convolution for dense image prediction. In Proceedings of the 2025 Computer Vision and Pattern Recognition Conference (CVPR); IEEE: New York, NY, USA, 2025; pp. 30178–30188. [Google Scholar]
- Wang, A.; Chen, H.; Lin, Z.; Han, J.; Ding, G. Lsnet: See large, focus small. In Proceedings of the 2025 Computer Vision and Pattern Recognition Conference (CVPR); IEEE: New York, NY, USA, 2025; pp. 9718–9729. [Google Scholar]
- Huang, H.; Xia, T.; Ren, P. Partial channel network: Compute fewer, perform better. arXiv 2025, arXiv:2502.01303. [Google Scholar]
- Yuan, X.; Zheng, Z.; Li, Y.; Liu, X.; Liu, L.; Li, X.; Hou, Q.; Cheng, M.-M. Strip R-CNN: Large strip convolution for remote sensing object detection. In Proceedings of the AAAI Conference on Artificial Intelligence; The Association for the Advancement of Artificial Intelligence: Washington, DC, USA, 2026; pp. 12259–12267. [Google Scholar]
- Pan, Z.; Cai, J.; Zhuang, B. Fast vision transformers with hilo attention. Adv. Neural Inf. Process. Syst. 2022, 35, 14541–14554. [Google Scholar] [CrossRef] [Scilit]
- Xu, L.; Zhang, D.; Song, Z. Pushing Trade-Off Boundaries: Compact yet Effective Remote Sensing Change Detection. In Proceedings of the 33rd ACM International Conference on Multimedia; Association for Computing Machinery: New York, NY, USA; IEEE: New York, NY, USA, 2025; pp. 641–649. [Google Scholar]
- Yin, D.; Hu, L.; Li, B.; Zhang, Y.; Yang, X. 5% > 100%: Breaking performance shackles of full fine-tuning on visual recognition tasks. In Proceedings of the 2025 Computer Vision and Pattern Recognition Conference (CVPR); IEEE: New York, NY, USA, 2025; pp. 20071–20081. [Google Scholar]
- Wang, W.; Chen, W.; Qiu, Q.; Chen, L.; Wu, B.; Lin, B.; He, X.; Liu, W. Crossformer++: A versatile vision transformer hinging on cross-scale attention. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 3123–3136. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, Z.; Ding, T.; Lu, Y.; Pai, D.; Zhang, J.; Wang, W.; Yu, Y.; Ma, Y.; Haeffele, B. Token statistics transformer: Linear-time attention via variational rate reduction. In Proceedings of the International Conference on Learning Representations; Curran Associates, Inc.: Red Hook, NY, USA, 2025; pp. 38536–38560. [Google Scholar]
















| Method | Params/M | VisDrone2019-DET | ||||||
|---|---|---|---|---|---|---|---|---|
| P/% | R/% | mAP0.5/% | mAP0.5:0.95/% | APs/% | FPS | GFLOPs | ||
| Baseline | 19.8 | 59.7 | 44.7 | 45.9 | 28.0 | 21.1 | 50.5 | 57.0 |
| +FICE-Net | 13.7 | 61.5 | 47.5 | 48.3 | 28.8 | 22.0 | 60.1 | 46.3 |
| +FICE-Net + Binary-AIFI | 13.7 | 64.9 | 47.7 | 49.1 | 29.8 | 22.6 | 59.8 | 46.5 |
| +FICE-Net + Binary-AIFI + RCFF | 13.7 | 66.0 | 50.1 | 51.2 | 32.6 | 23.5 | 66.9 | 46.5 |
| Method | Params/M | HIT-UAV | ||||||
| P/% | R/% | mAP0.5/% | mAP0.5:0.95/% | APs/% | FPS | GFLOPs | ||
| Baseline | 19.8 | 85.1 | 75.6 | 77.5 | 49.3 | 37.5 | 50.5 | 57.0 |
| +FICE-Net | 13.7 | 84.3 | 72.9 | 77.8 | 49.9 | 38.1 | 60.1 | 46.3 |
| +FICE-Net + Binary-AIFI | 13.7 | 78.3 | 79.2 | 81.5 | 50.6 | 41.5 | 59.8 | 46.5 |
| +FICE-Net + Binary-AIFI + RCFF | 13.7 | 85.3 | 80.0 | 83.4 | 56.3 | 42.4 | 66.9 | 46.5 |
| FICE-Net | Binary-AIFI | RCFF | mAP /% | AP/% | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PT | PE | BC | CA | VA | TK | TC | AT | BU | MO | ||||
| × | × | × | 45.9 | 54.4 | 46.8 | 21.2 | 84.8 | 48.2 | 35.6 | 32.1 | 17.2 | 59.6 | 58.6 |
| √ | × | × | 48.3 | 56.7 | 48.7 | 21.8 | 85.5 | 48.4 | 41.0 | 34.8 | 19.7 | 66.2 | 60.6 |
| × | √ | × | 47.5 | 56.0 | 49.4 | 21.2 | 86.0 | 51.1 | 37.1 | 34.7 | 18.7 | 59.8 | 60.4 |
| × | × | √ | 48.0 | 56.5 | 48.5 | 22.0 | 85.6 | 51.0 | 39.3 | 33.6 | 17.4 | 66.1 | 59.6 |
| √ | √ | × | 49.1 | 56.3 | 49.6 | 25.0 | 77.9 | 48.3 | 41.1 | 40.7 | 27.3 | 63.8 | 60.9 |
| √ | × | √ | 49.3 | 57.2 | 51.7 | 24.1 | 78.8 | 49.1 | 40.8 | 40.2 | 25.4 | 63.3 | 62.2 |
| × | √ | √ | 48.2 | 54.9 | 49.6 | 23.6 | 78.2 | 49.7 | 39.2 | 39.0 | 25.8 | 61.8 | 60.1 |
| √ | √ | √ | 51.2 | 60.0 | 50.1 | 22.8 | 85.7 | 53.3 | 43.5 | 40.4 | 24.9 | 70.4 | 60.9 |
| FICE-Net | Binary-AIFI | RCFF | mAP /% | AP/% | ||||
|---|---|---|---|---|---|---|---|---|
| PE | CA | OT | BI | DO | ||||
| × | × | × | 77.5 | 93.0 | 98.5 | 79.3 | 91.2 | 25.6 |
| √ | × | × | 77.8 | 92.8 | 98.6 | 80.8 | 91.9 | 25.1 |
| × | √ | × | 79.5 | 93.6 | 98.7 | 66.8 | 93.2 | 45.3 |
| × | × | √ | 78.1 | 94.0 | 98.7 | 58.8 | 92.0 | 47.3 |
| √ | √ | × | 81.5 | 92.6 | 98.7 | 87.9 | 92.0 | 36.0 |
| √ | × | √ | 81.6 | 94.3 | 98.9 | 86.0 | 92.3 | 36.3 |
| × | √ | √ | 80.9 | 93.7 | 98.8 | 66.1 | 91.1 | 54.6 |
| √ | √ | √ | 83.4 | 94.6 | 99.1 | 96.9 | 93.4 | 33.1 |
| Method | Params/M | VisDrone2019-DET | HIT-UAV | ||
|---|---|---|---|---|---|
| mAP0.5/% | mAP0.5:0.95/% | mAP0.5/% | mAP0.5:0.95/% | ||
| YOLOv5m [53] | 25.1 | 42.6 | 25.5 | 77.2 | 47.5 |
| YOLOv8m [54] | 25.8 | 43.6 | 26.2 | 78.2 | 51.4 |
| YOLOv10m [55] | 15.3 | 43.6 | 26.2 | 73.6 | 45.9 |
| YOLO11m [56] | 20.0 | 45.3 | 27.8 | 76.1 | 48.9 |
| YOLO26m [57] | 20.4 | 45.0 | 27.4 | 76.9 | 48.0 |
| Faster R-CNN [13] | 41.3 | 41.4 | 21.9 | 56.2 | 27.6 |
| DETR [21] | 41.3 | 41.3 | 23.1 | 71.4 | 41.4 |
| DINO [58] | 47.5 | 47.5 | 28.4 | 78.9 | 48.4 |
| RT-DETR-L [23] | 32.0 | 41.1 | 24.2 | 75.5 | 48.4 |
| RT-DETR-R50 [23] | 41.9 | 42.8 | 25.5 | 77.5 | 48.9 |
| RT-DETR-R18 [23] | 19.8 | 45.9 | 28.0 | 77.5 | 49.3 |
| DMF-YOLO * [59] | 17.6 | 50.1 | 29.7 | 81.4 | 52.8 |
| UAV-DETR * [24] | 42.0 | 51.1 | 31.5 | 74.9 | — |
| CSFPR-RTDETR * [25] | 14.09 | 42.3 | 24.9 | 83.1 | — |
| Ours (FBR-DETR) | 13.7 | 51.2 | 32.6 | 83.4 | 56.3 |
| Method | mAP/% | AP/% | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| PT | PE | BC | CA | VA | TK | TC | AT | BU | MO | ||
| YOLOv5m [53] | 42.6 | 47.8 | 36.5 | 15.1 | 82.2 | 47.7 | 39.7 | 29.6 | 18.3 | 59.6 | 49.5 |
| YOLOv8m [54] | 43.6 | 49.8 | 38.2 | 15.8 | 82.8 | 48.2 | 39.8 | 32.4 | 17.2 | 60.7 | 50.7 |
| YOLOv10m [55] | 43.6 | 49.9 | 38.5 | 18.0 | 82.5 | 45.9 | 40.0 | 32.0 | 15.7 | 62.5 | 51.3 |
| YOLO11m [56] | 45.3 | 50.8 | 37.5 | 18.2 | 82.6 | 50.7 | 44.5 | 35.1 | 19.9 | 62.1 | 51.4 |
| YOLO26m [57] | 45.0 | 53.4 | 40.4 | 19.3 | 83.4 | 47.5 | 39.3 | 33.0 | 18.4 | 62.3 | 53.1 |
| DETR [21] | 41.3 | 38.5 | 32.4 | 17.4 | 79.5 | 42.5 | 34.2 | 29.4 | 17.2 | 71.5 | 50.4 |
| DINO [58] | 47.5 | 54.4 | 37.8 | 21.2 | 84.6 | 49.4 | 40.5 | 34.6 | 19.8 | 76.4 | 56.3 |
| RT-DETR-L [23] | 41.1 | 40.3 | 40.8 | 12.0 | 79.5 | 42.6 | 39.2 | 31.3 | 24.1 | 51.4 | 49.4 |
| RT-DETR-R50 [23] | 42.8 | 42.4 | 40.9 | 13.5 | 79.7 | 45.1 | 40.3 | 37.0 | 24.1 | 54.1 | 51.2 |
| RT-DETR-R18 [23] | 45.9 | 54.4 | 46.8 | 21.2 | 84.8 | 48.2 | 35.6 | 32.1 | 17.2 | 59.6 | 58.6 |
| DMF-YOLO * [59] | 50.1 | 56.6 | 47.4 | 24.6 | 87.5 | 53.1 | 41.7 | 39.8 | 25.4 | 65.9 | 59.7 |
| Ours (FBR-DETR) | 51.2 | 60.0 | 50.1 | 22.8 | 85.7 | 53.3 | 43.5 | 40.4 | 24.9 | 70.4 | 60.9 |
| Method | mAP /% | AP/% | ||||
|---|---|---|---|---|---|---|
| PE | CA | OT | BI | DO | ||
| YOLOv5m [53] | 77.2 | 87.3 | 98.2 | 67.3 | 86.9 | 46.1 |
| YOLOv8m [54] | 78.2 | 89.5 | 96.9 | 65.4 | 90.2 | 49.0 |
| YOLOv10m [55] | 73.6 | 86.4 | 97.6 | 68.0 | 85.2 | 30.6 |
| YOLO11m [56] | 76.1 | 87.0 | 98.0 | 70.0 | 88.0 | 37.8 |
| YOLO26m [57] | 76.9 | 87.3 | 97.2 | 84.3 | 87.0 | 28.6 |
| DETR [21] | 71.4 | 91.6 | 96.2 | 86.1 | 43.1 | 40.0 |
| DINO [58] | 78.9 | 92.4 | 96.4 | 88.5 | 57.3 | 59.9 |
| RT-DETR-L [23] | 75.5 | 93.9 | 98.3 | 73.9 | 91.7 | 19.5 |
| RT-DETR-R50 [23] | 77.5 | 93.5 | 98.8 | 74.7 | 91.0 | 29.3 |
| RT-DETR-R18 [23] | 77.5 | 93.0 | 98.5 | 79.3 | 91.2 | 25.6 |
| CSFPR-RTDETR * [25] | 83.1 | 94.4 | 96.2 | 63.4 | 91.0 | 70.6 |
| Ours (FBR-DETR) | 83.4 | 94.6 | 99.1 | 96.9 | 93.4 | 33.1 |
| Backbone Conv | Params/M | GFLOPs | P/% | R/% | mAP0.5/% | mAP0.75/% | mAP0.5:0.95/% | FPS |
|---|---|---|---|---|---|---|---|---|
| EBlock [60] | 13.7 | 44.6 | 55.1 | 39.7 | 40.4 | 21.1 | 22.1 | 48.1 |
| FDConv [61] | 14.1 | 41.6 | 56.4 | 40.7 | 40.9 | 23.6 | 24.0 | 25.8 |
| LSBlock [62] | 13.3 | 44.9 | 58.6 | 42.5 | 43.7 | 26.4 | 26.3 | 38.2 |
| PartialNetBlock [63] | 12.9 | 42.5 | 59.0 | 41.5 | 43.2 | 26.1 | 25.8 | 43.2 |
| Strip [64] | 14.3 | 47.7 | 58.7 | 40.8 | 42.1 | 25.1 | 25.0 | 54.1 |
| Converse2D | 13.7 | 46.3 | 61.5 | 47.5 | 48.3 | 29.2 | 28.8 | 60.1 |
| Attention | Params/M | GFLOPs | P/% | R/% | mAP0.5/% | mAP0.75/% | mAP0.5:0.95/% | FPS |
|---|---|---|---|---|---|---|---|---|
| HiLo [65] | 19.8 | 57.1 | 59.7 | 45.0 | 45.7 | 28.1 | 27.7 | 85.1 |
| EGSA [66] | 19.6 | 57.0 | 53.2 | 34.7 | 35.5 | 20.0 | 20.5 | 90.3 |
| Mona [67] | 19.9 | 57.0 | 58.4 | 41.5 | 42.8 | 25.3 | 25.3 | 49.9 |
| DPB [68] | 19.8 | 57.2 | 60.9 | 44.4 | 46.0 | 28.2 | 27.7 | 81.4 |
| TSSA [69] | 19.7 | 57.1 | 59.8 | 43.9 | 45.5 | 27.2 | 27.0 | 89.3 |
| Binary-AIFI | 19.8 | 57.2 | 60.7 | 44.9 | 47.5 | 29.3 | 28.6 | 54.4 |
| Module Name | Peak Memory/MB | PyTorch Latency/ms | TensorRT FP16 Latency/ms | Operating Power/W |
|---|---|---|---|---|
| AIFI | 86.50 | 1.15 | 0.58 | 14.8 |
| Binary-AIFI | 40.14 | 0.28 | 0.32 | 11.4 |
| Kernel Size k | Params/M | GFLOPs | mAP0.5/% | mAP0.5:0.95/% |
|---|---|---|---|---|
| k = 1 × 1 | 13.74 | 46.3 | 47.1 | 29.6 |
| k = 3 × 3 | 13.76 | 46.3 | 51.2 | 32.6 |
| k = 5 × 5 | 13.79 | 46.3 | 50.7 | 32.1 |
| k = 7 × 7 | 13.84 | 46.4 | 50.1 | 31.7 |
| Offset b | Initial λ0 | Params/M | GFLOPs | mAP0.5/% | mAP0.5:0.95/% |
|---|---|---|---|---|---|
| b = −5.0 | ≈6.70 × 10−3 | 13.76 | 46.3 | 49.6 | 31.2 |
| b = −7.0 | ≈9.21 × 10−4 | 13.76 | 46.3 | 50.5 | 32.0 |
| b = −9.0 | ≈1.23 × 10−4 | 13.76 | 46.3 | 51.2 | 32.6 |
| b = −11.0 | ≈2.67 × 10−5 | 13.76 | 46.3 | 49.8 | 31.5 |
| Quantization Scheme | Q/K Precision | V Precision | Peak Memory/MB | GPU Latency /ms | mAP0.5/% | mAP0.5:0.95/% |
|---|---|---|---|---|---|---|
| Full Precision | FP32 | FP32 | 86.50 | 3.033 | 48.0 | 30.1 |
| Uniform 4-bit | INT4 | FP32 | 54.20 | 1.250 | 50.3 | 31.8 |
| Binary-AIFI | 1-bit ([−1, +1]) | FP32 | 40.14 | 0.709 | 51.2 | 32.6 |
| Extreme Low-bit | 1-bit ([−1, +1]) | INT8 | 36.80 | 0.685 | 49.5 | 31.0 |
| Model | Mode | Params/M | Latency/ms | FPS |
|---|---|---|---|---|
| YOLOv8m | FP16 | 25.80 | 21.8 | 45.8 |
| RT-DETR-R18 | FP16 | 19.81 | 24.5 | 40.8 |
| FBR-DETR (Ours) | FP16 | 13.76 | 16.0 | 62.5 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liang, S.; Wang, R.; Wang, X.; Wang, A.; Sun, J.; Liu, H. FBR-DETR: An Efficient End-to-End Network for Real-Time Small-Object Detection in UAV Imagery. Remote Sens. 2026, 18, 2900. https://doi.org/10.3390/rs18172900
Liang S, Wang R, Wang X, Wang A, Sun J, Liu H. FBR-DETR: An Efficient End-to-End Network for Real-Time Small-Object Detection in UAV Imagery. Remote Sensing. 2026; 18(17):2900. https://doi.org/10.3390/rs18172900
Chicago/Turabian StyleLiang, Shuai, Ran Wang, Xiao Wang, Aixue Wang, Jialong Sun, and Hui Liu. 2026. "FBR-DETR: An Efficient End-to-End Network for Real-Time Small-Object Detection in UAV Imagery" Remote Sensing 18, no. 17: 2900. https://doi.org/10.3390/rs18172900
APA StyleLiang, S., Wang, R., Wang, X., Wang, A., Sun, J., & Liu, H. (2026). FBR-DETR: An Efficient End-to-End Network for Real-Time Small-Object Detection in UAV Imagery. Remote Sensing, 18(17), 2900. https://doi.org/10.3390/rs18172900

