MRU-YOLO: Marginal-Utility-Guided Selective Local Re-Observation for Small-Object Detection in UAV Imagery
Highlights
- MRU-YOLO formulates selective local re-observation as a prediction-conditioned marginal-utility ranking problem, enabling a fixed local-processing budget to prioritize regions with the highest expected residual detection gain after a single global pass.
- Across three independent runs, MRU-YOLO improved mean mAP50–95 by 2.32 and 3.28 percentage points over YOLO11n-640 on SeaDronesSee ODv2 and VisDrone2019-DET, respectively.
- The results show that expected marginal detection utility provides an effective regional allocation signal under a fixed local-processing budget.
- By operating entirely at inference time without modifying the detector backbone, neck, or detection head, MRU-YOLO provides a practical strategy for budgeted local computation in UAV small-object detection.
Abstract
1. Introduction
- (1)
- Prediction-conditioned marginal-utility learning. We combine a 26-dimensional candidate state with offline multi-IoU supervision to estimate residual regional value from global predictions. The learned selector ranks regions by the expected detection contribution of another local observation.
- (2)
- Budgeted selective local re-observation. Following one global pass, the policy ranks nine candidates and assigns the Top-2 observation budget to the highest-utility regions using shared detector weights. Online selection requires no annotations, no additional whole-image detector pass, and no modification to the detector architecture.
- (3)
- Source-aware fusion of global and local predictions. Original-coordinate size gating, local-score calibration, stable-global protection, and class-aware suppression preserve valid local recoveries while controlling duplicate boxes, confidence shifts, and crop-induced false positives.
2. Materials and Methods
2.1. Framework Overview
2.2. Candidate Regions and Prediction-Conditioned State
2.3. Offline Marginal-Utility Construction and Learning
2.4. Online Budgeted Local Re-Observation
2.5. Source-Aware Global–Local Fusion
2.6. Datasets and Data Preparation
2.7. Training and Implementation Details
2.8. Evaluation Protocol and Ranking Diagnostics
3. Results
3.1. Overall Detection Performance
3.2. Class-Wise and Qualitative Results
3.3. Selection Performance and Ranking Quality
3.4. Selector Robustness
3.5. Observation-Budget Analysis
3.6. Scene-Density and Computational Analysis
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Mohsan, S.A.H.; Khan, M.A.; Noor, F.; Ullah, I.; Alsharif, M.H. Towards the Unmanned Aerial Vehicles (UAVs): A Comprehensive Review. Drones 2022, 6, 147. [Google Scholar] [CrossRef] [Scilit]
- Yang, T.; Jiang, Z.; Sun, R.; Cheng, N.; Feng, H. Maritime Search and Rescue Based on Group Mobile Computing for Unmanned Aerial Vehicles and Unmanned Surface Vehicles. IEEE Trans. Ind. Inform. 2020, 16, 7700–7708. [Google Scholar] [CrossRef] [Scilit]
- Nikouei, M.; Baroutian, B.; Nabavi, S.; Taraghi, F.; Aghaei, A.; Sajedi, A.; Moghaddam, M.E. Small Object Detection: A Comprehensive Survey on Challenges, Techniques and Real-World Applications. Intell. Syst. Appl. 2025, 27, 200561. [Google Scholar] [CrossRef] [Scilit]
- Yu, X.; Gong, Y.; Jiang, N.; Ye, Q.; Han, Z. Scale Match for Tiny Person Detection. In Proceedings of the 2020 IEEE Winter Conference on Applications of Computer Vision, Snowmass, CO, USA, 1–5 March 2020; pp. 1246–1254. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Wang, A.; Yi, J.; Song, Y.; Chehri, A. Small Object Detection Based on Deep Learning for Remote Sensing: A Comprehensive Review. Remote Sens. 2023, 15, 3265. [Google Scholar] [CrossRef] [Scilit]
- Tang, G.; Ni, J.; Zhao, Y.; Gu, Y.; Cao, W. A Survey of Object Detection for UAVs Based on Deep Learning. Remote Sens. 2024, 16, 149. [Google Scholar] [CrossRef] [Scilit]
- Ni, J.; Zhu, S.; Tang, G.; Ke, C.; Wang, T. A Small-Object Detection Model Based on Improved YOLOv8s for UAV Image Scenarios. Remote Sens. 2024, 16, 2465. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Yang, W.; Guo, H.; Zhang, R.; Xia, G.-S. Tiny Object Detection in Aerial Images. In Proceedings of the 25th International Conference on Pattern Recognition, Milan, Italy, 10–15 January 2021; pp. 3791–3798. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Ye, M.; Zhu, G.; Liu, Y.; Guo, P.; Yan, J. FFCA-YOLO for Small Object Detection in Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5611215. [Google Scholar] [CrossRef] [Scilit]
- Akyon, F.C.; Altinuc, S.O.; Temizel, A. Slicing Aided Hyper Inference and Fine-Tuning for Small Object Detection. In Proceedings of the 2022 IEEE International Conference on Image Processing, Bordeaux, France, 16–19 October 2022; pp. 966–970. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Yang, T.; Zhu, S.; Chen, C.; Guan, S. Density Map Guided Object Detection in Aerial Images. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA, 14–19 June 2020; pp. 737–746. [Google Scholar] [CrossRef] [Scilit]
- Yang, F.; Fan, H.; Chu, P.; Blasch, E.; Ling, H. Clustered Object Detection in Aerial Images. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 8310–8319. [Google Scholar] [CrossRef] [Scilit]
- Najibi, M.; Singh, B.; Davis, L.S. AutoFocus: Efficient Multi-Scale Inference. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9744–9754. [Google Scholar] [CrossRef] [Scilit]
- Yang, C.; Huang, Z.; Wang, N. QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 13658–13667. [Google Scholar] [CrossRef] [Scilit]
- Jiang, L.; Yuan, B.; Du, J.; Chen, B.; Xie, H.; Tian, J.; Yuan, Z. MFFSODNet: Multiscale Feature Fusion Small Object Detection Network for UAV Aerial Images. IEEE Trans. Instrum. Meas. 2024, 73, 5015214. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 936–944. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Li, X.; Chen, J.; Zhou, L.; Guo, L.; He, Z.; Zhou, H.; Zhang, Z. DPH-YOLOv8: Improved YOLOv8 Based on Double Prediction Heads for the UAV Image Object Detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5647715. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path Aggregation Network for Instance Segmentation. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8759–8768. [Google Scholar] [CrossRef] [Scilit]
- Doherty, J.; Gardiner, B.; Kerr, E.; Siddique, N. BiFPN-YOLO: One-Stage Object Detection Integrating Bi-Directional Feature Pyramid Networks. Pattern Recognit. 2025, 160, 111209. [Google Scholar] [CrossRef] [Scilit]
- Xu, S.; Song, L.; Yin, J.; Chen, Q.; Zhan, T.; Huang, W. MFFCI–YOLOv8: A Lightweight Remote Sensing Object Detection Network Based on Multiscale Features Fusion and Context Information. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 19743–19755. [Google Scholar] [CrossRef] [Scilit]
- Zhou, S.; Zhou, H. Detection Based on Semantics and a Detail Infusion Feature Pyramid Network and a Coordinate Adaptive Spatial Feature Fusion Mechanism Remote Sensing Small Object Detector. Remote Sens. 2024, 16, 2416. [Google Scholar] [CrossRef] [Scilit]
- Sun, C.; Zhang, Y.; Ma, S. DFLM-YOLO: A Lightweight YOLO Model with Multiscale Feature Fusion Capabilities for Open Water Aerial Imagery. Drones 2024, 8, 400. [Google Scholar] [CrossRef] [Scilit]
- Jin, Z.; He, T.; Qiao, L.; Duan, J.; Shi, X.; Yan, B.; Guo, C. MES-YOLO: An Efficient Lightweight Maritime Search and Rescue Object Detection Algorithm with Improved Feature Fusion Pyramid Network. J. Vis. Commun. Image Represent. 2025, 109, 104453. [Google Scholar] [CrossRef] [Scilit]
- Xu, J.; Fan, X.; Jian, H.; Xu, C.; Bei, W.; Ge, Q.; Zhao, T. YoloOW: A Spatial Scale Adaptive Real-Time Object Detection Neural Network for Open Water Search and Rescue from UAV Aerial Imagery. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5623115. [Google Scholar] [CrossRef] [Scilit]
- Zhao, B.; Zhou, Y.; Song, R.; Yu, L.; Zhang, X.; Liu, J. Modular YOLOv8 Optimization for Real-Time UAV Maritime Rescue Object Detection. Sci. Rep. 2024, 14, 24492. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ma, S.; Zhang, Y.; Peng, L.; Sun, C.; Ding, B.; Zhu, Y. OWRT-DETR: A Novel Real-Time Transformer Network for Small-Object Detection in Open-Water Search and Rescue from UAV Aerial Imagery. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4205313. [Google Scholar] [CrossRef] [Scilit]
- Liu, Q.; Yu, H.; Zhang, P.; Geng, T.; Yuan, X.; Ji, B.; Zhu, S.; Ma, R. MFEF-YOLO: A Multi-Scale Feature Extraction and Fusion Network for Small Object Detection in Aerial Imagery over Open Water. Remote Sens. 2025, 17, 3996. [Google Scholar] [CrossRef] [Scilit]
- Pang, J.; Li, C.; Shi, J.; Xu, Z.; Feng, H. R2-CNN: Fast Tiny Object Detection in Large-Scale Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2019, 57, 5512–5524. [Google Scholar] [CrossRef] [Scilit]
- Garza, J.E.; Islam, M.F. Enhanced YOLOv12 through Sliced Contrastive Supervision and Full-Scene Fine-Tuning. IEEE Access 2025, 13, 138813–138819. [Google Scholar] [CrossRef] [Scilit]
- Khorsand, H.; Arezoomandan, S.; Han, D.K. Enhanced Long-Range UAV Detection: Leveraging Slicing Aided Hyper Inference with YOLOv8. In Proceedings of the 2025 IEEE International Conference on Consumer Electronics, Las Vegas, NV, USA, 11–14 January 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Huang, K.; Yang, Y.; Jiang, Y.; Zhang, X.; Li, Z.A. AFSDet: Video Small Object Detection Based on Adaptive Focused Slicing. In Proceedings of the 2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, Macau, China, 3–6 December 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Ge, L.; Dou, L. Non-Maximum Suppression for Rotated Object Detection during Merging Slices of High-Resolution Images. IEEE Access 2024, 12, 149999–150007. [Google Scholar] [CrossRef] [Scilit]
- Meethal, A.; Granger, E.; Pedersoli, M. Cascaded Zoom-In Detector for High Resolution Aerial Images. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, BC, Canada, 17–24 June 2023; pp. 2046–2055. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Chen, Z.; Yang, B.; Guo, Q.; Wang, H.; Zeng, X. CAMS-AI: A Coarse-to-Fine Framework for Efficient Small Object Detection in High-Resolution Images. Remote Sens. 2026, 18, 259. [Google Scholar] [CrossRef] [Scilit]
- Kwon, S.; Lim, G.; Han, Y. SPAR-Det: Segmentation-Guided and Prior-Aided Routing for Small Object Detection. In Proceedings of the 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, Tucson, AZ, USA, 6–10 March 2026; pp. 2146–2155. [Google Scholar] [CrossRef] [Scilit]
- Burges, M.; Zambanini, S.; Sablatnig, R. Interactive Object Detection for Tiny Objects in Large Remotely Sensed Images. In Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision, Tucson, AZ, USA, 28 February–4 March 2025; pp. 4704–4713. [Google Scholar] [CrossRef] [Scilit]
- Wu, L.; Feng, Y.; Zhang, H.; Li, Y. Mask-Guided Feature Routing and Adaptive Context Modeling for Wide-FoV UAV Object Detection in IoT Remote Sensing. Remote Sens. 2026, 18, 1753. [Google Scholar] [CrossRef] [Scilit]
- Duan, C.; Wei, Z.; Zhang, C.; Qu, S.; Wang, H. Coarse-Grained Density Map Guided Object Detection in Aerial Images. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision Workshops, Montreal, QC, Canada, 11–17 October 2021; pp. 2789–2798. [Google Scholar] [CrossRef] [Scilit]
- Wang, D.; Sapkota, H.; Yu, Q. Adaptive Important Region Selection with Reinforced Hierarchical Search for Dense Object Detection. In Advances in Neural Information Processing Systems 37; Curran Associates: Red Hook, NY, USA, 2024; pp. 45636–45665. [Google Scholar] [CrossRef] [Scilit]
- Tang, D.; Tang, S.; Wang, Y.; Guan, S.; Jin, Y. A Global Object-Oriented Dynamic Network for Low-Altitude Remote Sensing Object Detection. Sci. Rep. 2025, 15, 19071. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shi, Y.; Li, J.; Jia, Y.; Hong, Q. LDA-DETR: A Lightweight Dynamic Attention-Enhanced DETR for Small Object Detection. PLoS ONE 2026, 21, e0340977. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, M. Dynamic Feature Pyramid Networks for Object Detection. In Proceedings of the Fifteenth International Conference on Signal Processing Systems, Xi’an, China, 17–19 November 2023; SPIE: Bellingham, WA, USA, 2024; Volume 13059, p. 1305903. [Google Scholar] [CrossRef] [Scilit]
- Qian, K.; Liu, W.; Chen, M.; Wang, X.; Yuan, X. RPUDet: Learning Relational Prior and Uncertainty for Robust Aerial Object Detection. In Proceedings of the 2025 International Conference on Multimedia Retrieval, Chicago, IL, USA, 30 June–3 July 2025; pp. 1109–1117. [Google Scholar] [CrossRef] [Scilit]
- Fuller, A.; Yassin, Y.; Wen, J.; Ibrahim, T.; Kyrollos, D.G.; Green, J.; Shelhamer, E. LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision. In Advances in Neural Information Processing Systems 38; Curran Associates: Red Hook, NY, USA, 2025. [Google Scholar]
- Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; Hsieh, C.-J. DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification. In Advances in Neural Information Processing Systems 34; Curran Associates: Red Hook, NY, USA, 2021; pp. 13937–13949. [Google Scholar]
- Zhan, Z.; Kong, Z.; Gong, Y.; Wu, Y.; Meng, Z.; Zheng, H.; Shen, X.; Ioannidis, S.; Niu, W.; Zhao, P.; et al. Exploring Token Pruning in Vision State Space Models. In Advances in Neural Information Processing Systems 37; Curran Associates: Red Hook, NY, USA, 2024. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar] [CrossRef] [Scilit]
- Neubeck, A.; Van Gool, L. Efficient Non-Maximum Suppression. In Proceedings of the 18th International Conference on Pattern Recognition, Hong Kong, China, 20–24 August 2006; pp. 850–855. [Google Scholar] [CrossRef] [Scilit]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
- Varga, L.A.; Kiefer, B.; Messmer, M.; Zell, A. SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water. In Proceedings of the 2022 IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 3–8 January 2022; pp. 3686–3696. [Google Scholar] [CrossRef] [Scilit]
- Zhu, P.; Wen, L.; Du, D.; Bian, X.; Fan, H.; Hu, Q.; Ling, H. Detection and Tracking Meet Drones Challenge. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 7380–7399. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Järvelin, K.; Kekäläinen, J. Cumulated Gain-Based Evaluation of IR Techniques. ACM Trans. Inf. Syst. 2002, 20, 422–446. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Computer Vision–ECCV 2014; Springer: Cham, Switzerland, 2014; Volume 8693, pp. 740–755. [Google Scholar] [CrossRef] [Scilit]










| Decision Property | Uniform Slicing | Cluster/Focus-Based | Density-Guided | MRU-YOLO |
|---|---|---|---|---|
| Representative references | SAHI [10] | ClusDet, AutoFocus [12,13] | DMNet, CGD [11,38] | This work |
| Candidate regions | Fixed tiles | Generated regions | Candidate regions | Nine expanded regions |
| Selection signal | Spatial coverage | Objectness or aggregation | Target density | Marginal utility |
| Prediction-conditioned | No | Proposal- or response-conditioned | Feature- or detection-conditioned | Conditioned on global predictions |
| Selective budget | No | Window- or proposal-dependent | Yes | Fixed Top-K |
| Incremental-value supervision | No | Task-specific focus objective | Density- or count- based objective | Marginal-utility target |
| Group | No. | Feature | Definition |
|---|---|---|---|
| Candidate geometry | 1 | Candidate index | Zero-based index for region . |
| 2 | Grid size | Grid side length; 3 in the formal configuration. | |
| 3 | Horizontal grid index | q mod 3. | |
| 4 | Vertical grid index | . | |
| 5 | Normalized left bound | . | |
| 6 | Normalized top bound | . | |
| 7 | Normalized right bound | . | |
| 8 | Normalized bottom bound | . | |
| 9 | Candidate area fraction | . | |
| 10 | Candidate aspect ratio | . | |
| 11 | Center distance | Euclidean distance from the normalized candidate center to . | |
| Image context | 12 | Image aspect ratio | . |
| 13 | Image width | Original width W (px). | |
| 14 | Image height | Original height H (px). | |
| 15 | Global prediction count | Predictions in the complete global set. | |
| Detection statistics | 16 | Candidate prediction count | Prediction centers inside the candidate, including the boundary. |
| 17 | Candidate prediction density | Candidate count per original-image candidate pixels. | |
| 18 | Mean confidence | Mean confidence of candidate-member predictions. | |
| 19 | Maximum confidence | Maximum confidence of candidate-member predictions. | |
| 20 | Confidence dispersion | Population standard deviation (ddof=0); zero for one prediction. | |
| 21 | Small-prediction count | Predictions with equivalent side px. | |
| 22 | Tiny-prediction count | Predictions with equivalent side px. | |
| 23 | Mean equivalent side | Mean (px). | |
| 24 | Maximum equivalent side | Maximum equivalent side s (px). | |
| Class statistics | 25 | Class entropy | , using natural logarithms. |
| 26 | Number of predicted classes | Distinct predicted classes inside the candidate. |
| (a) Dataset Characteristics | ||||||
| Dataset | Scenario | Train | Validation | Classes | Median GT/Image | Median Side at 640 px |
| SeaDronesSee ODv2 | Maritime | 8930 | 1547 | 5 | 6 | 8.5 |
| VisDrone2019-DET | Urban | 6471 | 548 | 10 | 65 | 11.3 |
| (b) Common Evaluation Protocol | ||||||
| Setting | Value | Setting | Value | |||
| Global input | 640 px | Local input | 640 px | |||
| Candidates | Expansion ratio | |||||
| Local budget | Top-2 | Confidence threshold | 0.001 | |||
| NMS IoU | 0.70 | Maximum detections | 1000 | |||
| Detector precision | FP16 | Final fusion | Source-aware | |||
| Online selection | One global pass; no GT | Hardware | RTX 3090 | |||
| Dataset | Method | P (%) | R (%) | mAP50 (%) | mAP50–95 (%, Mean ± SD) | Local Crops/ Image | ΔmAP50–95 (pp) |
|---|---|---|---|---|---|---|---|
| SeaDronesSee ODv2 | YOLO11n-640 (global) | 83.69 | 68.06 | 69.55 | 0 | 0.00 | |
| Predicted density (Top-2) | 75.43 | 70.07 | 72.73 | 2 | +2.29 | ||
| MRU-YOLO (Top-2) | 75.40 | 70.10 | 72.86 | 41.82 ± 0.34 | 2 | +2.32 | |
| VisDrone2019-DET | YOLO11n-640 (global) | 45.19 | 33.93 | 32.87 | 0 | 0.00 | |
| Predicted density (Top-2) | 46.05 | 39.62 | 37.71 | 2 | +3.24 | ||
| MRU-YOLO (Top-2) | 46.07 | 39.77 | 37.78 | 21.72 ± 0.07 | 2 | +3.28 |
| (a) Detection Performance | ||||
| Dataset | Method | mAP50–95 (%) | vs. Global Only (pp) | vs. Predicted Density (pp) |
| SeaDronesSee ODv2 | Global only | 39.496 | – | – |
| Predicted density (Top-2) | 41.784 | +2.289 | – | |
| Learned utility (Top-2) | 41.819 | +2.323 | +0.035 | |
| VisDrone2019-DET | Global only | 18.446 | – | – |
| Predicted density (Top-2) | 21.690 | +3.245 | – | |
| Learned utility (Top-2) | 21.725 | +3.279 | +0.035 | |
| (b) Offline Ranking Diagnostics | ||||
| Dataset | Selection Policy | Utility Capture (%) | NDCG@2 (%) | Evaluation Subset (Oracle Top-2 Utility > 0) |
| SeaDronesSee ODv2 | Predicted density | 83.74 | 80.16 | 1141/1547 images |
| Learned utility | 84.94 | 81.13 | ||
| VisDrone2019-DET | Predicted density | 84.19 | 83.04 | 532/548 images |
| Learned utility | 89.36 | 88.79 | ||
| (a) Controlled Ablations | |||||
| Component | Configuration | SeaDronesSee | VisDrone | ||
| mAP50–95 (%) | Δ vs. Ref. (pp) | mAP50–95 (%) | Δ vs. Ref. (pp) | ||
| State Representation | Full 26-dimensional state | 41.68 | 0.00 | 21.80 | 0.00 |
| Without detection- statistics features | 41.37 | −0.31 | 21.20 | −0.60 | |
| Prediction-only state | 41.79 | +0.11 | 21.77 | −0.04 | |
| Utility Target | Multi-IoU mean matched count | 41.68 | 0.00 | 21.80 | 0.00 |
| Matched count at IoU | 41.76 | +0.08 | 21.84 | +0.04 | |
| Label-Construction Fusion | Plain fusion | 41.71 | 0.00 | 21.82 | 0.00 |
| Source-aware fusion | 41.72 | +0.01 | 21.81 | −0.01 | |
| Inference-Time Fusion | Plain direct fusion | 39.90 | 0.00 | 21.47 | 0.00 |
| Joint source-aware fusion | 41.68 | +1.78 | 21.80 | +0.33 | |
| (b) Hyperparameter Sensitivity | |||||
| Dataset | Within 0.5 pp | Total Tested | Most Sensitive Parameter | Tested Value | ΔmAP50–95 (pp) |
| SeaDronesSee ODv2 | 15 | 18 | Fusion NMS IoU threshold | 0.90 | −1.37 |
| VisDrone2019-DET | 14 | 18 | Fusion NMS IoU threshold | 0.90 | −1.17 |
| Observation Budget | Local Crops/ Image | SeaDronesSee mAP50–95 (%) | VisDrone mAP50–95 (%) |
|---|---|---|---|
| Top-1 | 1 | 41.18 | 21.25 |
| Top-2 | 2 | 41.68 | 21.80 |
| Top-3 | 3 | 41.69 | 22.04 |
| Top-4 | 4 | 41.82 | 22.09 |
| All 9 candidates | 9 | 41.72 | 21.54 |
| Dataset | Density Stratum | Images | Mean GT/Image | ΔmAP50–95 (pp) |
|---|---|---|---|---|
| SeaDronesSee | Low | 515 | 2.13 | |
| SeaDronesSee | Mid | 516 | 6.18 | +3.20 |
| SeaDronesSee | High | 516 | 10.35 | +1.42 |
| VisDrone | Low | 182 | 26.71 | +2.11 |
| VisDrone | Mid | 183 | 65.81 | +3.15 |
| VisDrone | High | 183 | 119.42 | +3.89 |
| Dataset | Route | mAP50–95 (%) | Detector Inputs/Image | Nominal Detector GFLOPs | End-to-End Latency (ms) | FPS | Peak CUDA Memory (MB) |
|---|---|---|---|---|---|---|---|
| SeaDronesSee ODv2 | YOLO11n-640 (global) | 39.50 | 1.00 | 6.32 | 86.40 | 48.13 | |
| MRU-YOLO (Top-2) | 41.82 | 3.00 | 18.95 | 30.86 | 61.07 | ||
| SAHI (NMS/IoU) | 34.62 | 33.64 | 212.53 | 1.70 | 58.70 | ||
| VisDrone2019-DET | YOLO11n-640 (global) | 18.45 | 1.00 | 6.32 | 96.83 | 48.12 | |
| MRU-YOLO (Top-2) | 21.72 | 3.00 | 18.95 | 33.24 | 61.19 | ||
| SAHI (NMS/IoU) | 21.50 | 6.19 | 39.11 | 5.99 | 58.71 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chen, J.; He, J.; Wang, Y.; Lu, P.; Li, H. MRU-YOLO: Marginal-Utility-Guided Selective Local Re-Observation for Small-Object Detection in UAV Imagery. Remote Sens. 2026, 18, 2680. https://doi.org/10.3390/rs18162680
Chen J, He J, Wang Y, Lu P, Li H. MRU-YOLO: Marginal-Utility-Guided Selective Local Re-Observation for Small-Object Detection in UAV Imagery. Remote Sensing. 2026; 18(16):2680. https://doi.org/10.3390/rs18162680
Chicago/Turabian StyleChen, Jiajun, Jinxin He, Yongzhi Wang, Peng Lu, and Hengshuo Li. 2026. "MRU-YOLO: Marginal-Utility-Guided Selective Local Re-Observation for Small-Object Detection in UAV Imagery" Remote Sensing 18, no. 16: 2680. https://doi.org/10.3390/rs18162680
APA StyleChen, J., He, J., Wang, Y., Lu, P., & Li, H. (2026). MRU-YOLO: Marginal-Utility-Guided Selective Local Re-Observation for Small-Object Detection in UAV Imagery. Remote Sensing, 18(16), 2680. https://doi.org/10.3390/rs18162680

