UAV Visual Localization Method Based on Token-Level Local Matching Reranking and Neighborhood-Consistent Position Fusion
Highlights
- A token-level local matching reranking module improves the ranking reliability of top-K candidates by modeling token-level local matching evidence, geometric priors, and candidate set relationships.
- A neighborhood-consistent position fusion module reduces meter-level localization error under same-area settings by adaptively fusing multiple candidate tile centers and learning residual offsets in a local coordinate system.
- The proposed global-to-local framework bridges the gap between discrete satellite tile retrieval and continuous geographic position estimation for UAV visual localization.
- Experiments on GTA-UAV and UAV-VisLoc datasets validate the method’s effectiveness. Over five independent runs under the GTA-UAV same-area setting, the proposed method achieves an average Recall@1 (R@1) gain of 2.89 percentage points and an average Dis@1 reduction of 60.95 m relative to Global Retrieval. In a single evaluation run under the UAV-VisLoc same-area setting, the proposed method achieves an R@1 gain of 3.03 percentage points and a Dis@1 reduction of 39.53 m relative to Global Retrieval.
Abstract
1. Introduction
2. Related Work
2.1. Datasets, Task Settings, and Evaluation Protocols
2.2. Global Retrieval Methods for Cross-View Geo-Localization
2.3. Candidate Reranking Methods
2.4. Coordinate Refinement and Position Fusion
3. Methodology
3.1. Global Retrieval
3.2. Token-Level Local Matching Reranking
3.2.1. Feature Preparation
3.2.2. Local Match Pair Construction
3.2.3. Pair Encoding and Aggregation
3.2.4. CL-LMM Inference
3.2.5. Shortlist Score Calibration
3.3. Neighborhood-Consistent Position Fusion
3.3.1. Local Support Set Construction
3.3.2. Learnable Position Fusion Network
3.3.3. Training Objective
3.4. Summary of Notation
4. Results
4.1. Experimental Datasets
4.1.1. GTA-UAV Dataset
4.1.2. UAV-VisLoc Dataset
4.2. Experimental Setup
4.3. Main Experimental Results Across Multiple Scenarios
4.3.1. GTA-UAV Same-Area Scenario Experiments
| Method | R@1 (%) | R@5 (%) | AP (%) | SDM@1 (%) | SDM@3 (%) | Dis@1 (m) | Dis@3 (m) |
|---|---|---|---|---|---|---|---|
| GlobalRetrieval | 85.18 ± 0.23 | 97.61 ± 0.09 | 90.50 ± 0.15 | 90.17 ± 0.12 | 88.95 ± 0.11 | 145.85 ± 2.90 | 177.20 ± 3.30 |
| +CL-LMM | 87.45 ± 0.44 | 98.49 ± 0.08 | 92.04 ± 0.23 | 91.25 ± 0.16 | 90.01 ± 0.14 | 140.95 ± 2.60 | 169.40 ± 3.00 |
| +CL-LMM +Calibration | 88.07 ± 0.61 † | 98.77 ± 0.07 | 92.38 ± 0.25 | 91.47 ± 0.18 | 89.69 ± 0.15 | 139.30 ± 2.40 | 167.95 ± 2.80 |
| +CL-LMM + Calibration + LocFusion | 88.07 ± 0.61 | 98.77 ± 0.07 | 92.38 ± 0.25 | 93.84 ± 0.16 | 90.66 ± 0.14 | 84.90 ± 2.30 | 145.90 ± 2.70 |
4.3.2. GTA-UAV Cross-Area Scene Experiments
| Method | R@1 (%) | R@5 (%) | AP (%) | SDM@1 (%) | SDM@3 (%) | Dis@1 (m) | Dis@3 (m) |
|---|---|---|---|---|---|---|---|
| GlobalRetrieval | 57.82 ± 0.43 | 83.80 ± 0.28 | 68.68 ± 0.35 | 81.55 ± 0.24 | 75.70 ± 0.26 | 340.20 ± 9.20 | 507.10 ± 11.20 |
| +CL-LMM | 58.96 ± 0.70 | 83.86 ± 0.30 | 69.40 ± 0.43 | 81.72 ± 0.27 | 75.75 ± 0.28 | 322.10 ± 8.70 | 500.40 ± 10.60 |
| +CL-LMM +Calibration | 59.17 ± 0.82 † | 83.87 ± 0.31 | 69.57 ± 0.46 | 81.78 ± 0.28 | 75.71 ± 0.29 | 316.10 ± 8.30 | 512.00 ± 10.90 |
| +CL-LMM + Calibration + LocFusion | 59.17 ± 0.82 | 83.87 ± 0.31 | 69.57 ± 0.46 | 83.48 ± 0.31 | 76.82 ± 0.30 | 244.00 ± 7.20 | 464.50 ± 9.50 |
4.4. Comparative Experiments
4.4.1. Backbone and Loss Comparison
4.4.2. Comparison with Adapted Post-Retrieval Reranking Methods
4.4.3. Comparison with Adapted Localization Refinement Strategies
4.5. Ablation and Efficiency Analysis
4.6. Parameter Sensitivity Analysis
4.6.1. Sensitivity Analysis of Reranking Candidate Pool Size
4.6.2. Parameter Sensitivity of the LocFusion Residual Correction Scale
4.7. Generalization Experiments on the UAV-VisLoc Dataset
4.7.1. UAV-VisLoc Same-Area Results
4.7.2. UAV-VisLoc Cross-Area Results
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Workman, S.; Souvenir, R.; Jacobs, N. Wide-Area Image Geolocalization with Aerial Reference Imagery. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 3961–3969. [Google Scholar] [CrossRef] [Scilit]
- Tian, Y.; Chen, C.; Shah, M. Cross-View Image Matching for Geo-Localization in Urban Environments. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1998–2006. [Google Scholar] [CrossRef] [Scilit]
- Hu, S.; Feng, M.; Nguyen, R.M.H.; Lee, G.H. CVM-Net: Cross-View Matching Network for Image-Based Ground-to-Aerial Geo-Localization. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 7258–7267. [Google Scholar] [CrossRef] [Scilit]
- Zhu, S.; Shah, M.; Chen, C. TransGeo: Transformer Is All You Need for Cross-View Image Geo-Localization. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 1152–1161. [Google Scholar]
- Deuser, F.; Habel, K.; Oswald, N. Sample4Geo: Hard Negative Sampling For Cross-View Geo-Localisation. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 16847–16856. [Google Scholar]
- Zhu, S.; Yang, L.; Chen, C.; Shah, M.; Shen, X.; Wang, H. R2Former: Unified Retrieval and Reranking Transformer for Place Recognition. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 19370–19380. [Google Scholar]
- Toker, A.; Zhou, Q.; Maximov, M.; Leal-Taixe, L. Coming Down to Earth: Satellite-to-Street View Synthesis for Geo-Localization. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 6488–6497. [Google Scholar] [CrossRef] [Scilit]
- Shi, Y.; Wu, F.; Perincherry, A.; Vora, A.; Li, H. Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View Transformer. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 21516–21526. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Z.; Wei, Y.; Yang, Y. University-1652: A Multi-View Multi-Source Benchmark for Drone-Based Geo-Localization. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 1395–1403. [Google Scholar] [CrossRef] [Scilit]
- Dai, M.; Zheng, E.; Feng, Z.; Qi, L.; Zhuang, J.; Yang, W. Vision-Based UAV Self-Positioning in Low-Altitude Urban Environments. IEEE Trans. Image Process. 2024, 33, 493–508. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, R.; Yin, L.; Yang, M.; Wu, F.; Yang, Y.; Hu, W. SUES-200: A Multi-Height Multi-Scene Cross-View Image Benchmark Across Drone and Satellite. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 4825–4839. [Google Scholar] [CrossRef] [Scilit]
- Zhu, S.; Yang, T.; Chen, C. VIGOR: Cross-View Image Geo-Localization Beyond One-to-One Retrieval. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 19–25 June 2021; pp. 5316–5325. [Google Scholar] [CrossRef] [Scilit]
- Ji, Y.; He, B.; Tan, Z.; Wu, L. Game4Loc: A UAV Geo-Localization Benchmark from Game Data. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), Philadelphia, PA, USA, 25 February–4 March 2025; Volume 39, pp. 3913–3921. [Google Scholar] [CrossRef] [Scilit]
- Xu, W.; Yao, Y.; Cao, J.; Wei, Z.; Liu, C.; Wang, J.; Peng, M. UAV-VisLoc: A Large-Scale Dataset for UAV Visual Localization. arXiv 2024, arXiv:2405.11936. [Google Scholar]
- Ding, L.; Zhou, J.; Meng, L.; Long, Z. A Practical Cross-View Image Matching Method between UAV and Satellite for UAV-Based Geo-Localization. Remote Sens. 2021, 13, 47. [Google Scholar] [CrossRef] [Scilit]
- Zhuang, J.; Dai, M.; Chen, X.; Zheng, E. A Faster and More Effective Cross-View Matching Method of UAV and Satellite Images for UAV Geolocalization. Remote Sens. 2021, 13, 3979. [Google Scholar] [CrossRef] [Scilit]
- Xia, P.; Yu, L.; Wan, Y.; Wu, Q.; Chen, P.; Zhong, L.; Yao, Y.; Wei, D.; Liu, X.; Ru, L.; et al. Cross-View Geo-Localization with Panoramic Street-View and VHR Satellite Imagery in Decentrality Settings. ISPRS J. Photogramm. Remote Sens. 2025, 227, 1–11. [Google Scholar] [CrossRef] [Scilit]
- Yang, H.; Lu, X.; Zhu, Y. Cross-View Geo-Localization with Layer-to-Layer Transformer. In Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Virtual, 6–14 December 2021; pp. 29009–29020. [Google Scholar]
- Liu, L.; Li, H. Lending Orientation to Neural Networks for Cross-View Geo-Localization. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; pp. 5624–5633. [Google Scholar] [CrossRef] [Scilit]
- Lin, J.; Zheng, Z.; Zhong, Z.; Luo, Z.; Li, S.; Yang, Y.; Sebe, N. Joint Representation Learning and Keypoint Detection for Cross-View Geo-Localization. IEEE Trans. Image Process. 2022, 31, 3780–3792. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, Q.; Wang, T.; Yang, Z.; Li, H.; Lu, R.; Sun, Y.; Zheng, B.; Yan, C. SDPL: Shifting-Dense Partition Learning for UAV-View Geo-Localization. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 11810–11824. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Y.; Yang, H.; Lu, Y.; Huang, Q. Simple, Effective and General: A New Backbone for Cross-view Image Geo-localization. arXiv 2023, arXiv:2302.01572. [Google Scholar]
- Yan, Y.; Wang, M.; Su, N.; Hou, W.; Zhao, C.; Wang, W. IML-Net: A Framework for Cross-View Geo-Localization with Multi-Domain Remote Sensing Data. Remote Sens. 2024, 16, 1249. [Google Scholar] [CrossRef] [Scilit]
- Wu, Q.; Wan, Y.; Zheng, Z.; Zhang, Y.; Wang, G.; Zhao, Z. CAMP: A Cross-View Geo-Localization Method Using Contrastive Attributes Mining and Position-Aware Partitioning. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5637614. [Google Scholar] [CrossRef] [Scilit]
- Zhao, H.; Ren, K.; Yue, T.; Zhang, C.; Yuan, S. TransFG: A Cross-View Geo-Localization of Satellite and UAVs Imagery Pipeline Using Transformer-Based Feature Aggregation and Gradient Guidance. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4700912. [Google Scholar] [CrossRef] [Scilit]
- Hu, Y.; Liu, Y.; Hui, B. Combining OpenStreetMap with Satellite Imagery to Enhance Cross-View Geo-Localization. Sensors 2025, 25, 44. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, Z.; Shi, D.; Qiu, C.; Jin, S.; Li, T.; Qiao, Z.; Chen, Y. VecMapLocNet: Vision-Based UAV Localization Using Vector Maps in GNSS-Denied Environments. ISPRS J. Photogramm. Remote Sens. 2025, 225, 362–381. [Google Scholar] [CrossRef] [Scilit]
- Zhong, Z.; Zheng, L.; Cao, D.; Li, S. Re-Ranking Person Re-Identification with k-Reciprocal Encoding. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 3652–3661. [Google Scholar] [CrossRef] [Scilit]
- Kannan, S.S.; Min, B.-C. PlaceFormer: Transformer-Based Visual Place Recognition Using Multi-Scale Patch Selection and Fusion. IEEE Robot. Autom. Lett. 2024, 9, 6552–6559. [Google Scholar] [CrossRef] [Scilit]
- Hu, B.; Chen, L.; Chen, R.; Bu, S.; Han, P.; Li, H. CurriculumLoc: Enhancing Cross-Domain Geolocalization Through Multistage Refinement. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5615914. [Google Scholar] [CrossRef] [Scilit]
- Dagda, B.; Awais, M.; Fallah, S. GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching. arXiv 2025, arXiv:2505.13669. [Google Scholar]
- Zhang, X.; Shore, T.; Chen, C.; Mendez, O.; Hadfield, S.; Wshah, S. VICI: VLM-Instructed Cross-view Image-localisation. In Proceedings of the 3rd International Workshop on UAVs in Multimedia: Capturing the World from a New Perspective (UAVM ’25), Dublin, Ireland, 27–31 October 2025. [Google Scholar] [CrossRef] [Scilit]
- Xiao, Z.; Suma, P.; Sachdeva, A.; Wang, H.-J.; Kordopatis-Zilos, G.; Tolias, G.; Ordonez, V. LOCORE: Image Re-ranking with Long-Context Sequence Modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 9580–9590. [Google Scholar]
- Lentsch, T.; Xia, Z.; Caesar, H.; Kooij, J.F.P. SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 17225–17234. [Google Scholar]
- Wang, X.; Xu, R.; Cui, Z.; Wan, Z.; Zhang, Y. Fine-Grained Cross-View Geo-Localization Using a Correlation-Aware Homography Estimator. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
- Zhang, X.; Sultani, W.; Wshah, S. Cross-View Image Sequence Geo-Localization. In Proceedings of the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–7 January 2023; pp. 2914–2923. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Wang, J.; Wei, Z.; Xu, W. Jointly Optimized Global-Local Visual Localization of UAVs. arXiv 2023, arXiv:2310.08082. [Google Scholar] [CrossRef] [Scilit]
- Ye, Q.; Luo, J.; Lin, Y. A Coarse-to-Fine Visual Geo-Localization Method for GNSS-Denied UAV with Oblique-View Imagery. ISPRS J. Photogramm. Remote Sens. 2024, 212, 306–322. [Google Scholar] [CrossRef] [Scilit]
- Shetty, A.; Gao, G.X. UAV Pose Estimation Using Cross-View Geolocalization with Satellite Imagery. In Proceedings of the 2019 IEEE International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20–24 May 2019; pp. 1827–1833. [Google Scholar] [CrossRef] [Scilit]






| Symbol | Description |
|---|---|
| The i-th drone-view query image and the j-th satellite candidate. | |
| Scaled cosine similarity between the i-th query descriptor and the j-th satellite descriptor. | |
| Learnable temperature coefficient used to rescale cosine similarities. | |
| Ground footprints of the query field of view and the satellite reference tile. | |
| Overlap weight of the i-th query–reference pair, computed from the IoU between and | |
| Number of retrieved candidates and the corresponding Top-K candidate set. | |
| Number of retained local tokens and the number of matched candidate tokens selected for each query token. | |
| Local patch-token features of the query token and candidate token. | |
| Token importance scores of the query token and candidate token. | |
| Normalized 2D coordinates of the query token and candidate token in the feature grid, computed from their row and column indices. | |
| Dimension-reduced local token representations. | |
| Fused token weights of the query token and candidate token, obtained by combining token importance scores with learnable gating scores. | |
| Token-level cosine similarity between the m-th query token and the n-th token in the j-th candidate. | |
| The r-th matched candidate-token position selected for the m-th query token. | |
| Relative displacement between the query token and its matched candidate token. | |
| Local match-pair feature constructed from token coordinates, displacement, saliency, and similarity information. | |
| Same-scale and cross-scale neighborhood sets of the j-th candidate. | |
| Probabilities of local matching, strict-positive tendency, structural prior, and fusion branches. | |
| Number of query–candidate pairs in the current batch for CL-LMM training losses. | |
| Query-level shortlist context vector computed from the Top-K shortlist. | |
| Score adjustment predicted by the shortlist score calibration network. | |
| Final refined ranking score after shortlist score calibration. | |
| Local support set and the number of support candidates used in LocFusion. | |
| Local coordinates of the m-th support candidate, target query location, and predicted location. | |
| Learnable residual correction in LocFusion. | |
| Local coordinate normalization scale factor. | |
| Haversine distance function between two geographic coordinates. |
| Module | Hyperparameter Description | Symbol | Value |
|---|---|---|---|
| Global Retrieval | Adaptive smoothing coefficient (Weighted-InfoNCE) | 5 | |
| CL-LMM | Retained local tokens per image | 8 | |
| Matched candidate tokens per query token | 80 | ||
| Loss balancing weights | 1, 1, 1 | ||
| Shortlist Calibration | Baseline score prior weights | 0.60, 0.28, 0.08, 0.04 | |
| Residual scaling coefficient (score adjustment) | 0.18 | ||
| Ordering margins (Pairwise ranking) | |||
| Anchor threshold & Anchor protection margin | 0.03, 0.01 | ||
| Loss balancing weights (Equation (72)) | 1, 1.0, 0.5, 0.3, 0.02 | ||
| LocFusion | Support-set size & Local coord normalization scale | 3, 200 | |
| Distance decay parameter & Easy-sample threshold | 80, 60 | ||
| Residual scaling coefficient (coordinate correction) | 40 | ||
| Loss balancing weights | 1, 0.35, 0.02, 0.005, 0.1 |
| Method | R@1 ↑ | R@5 ↑ | AP ↑ | SDM@3 ↑ | Dis@1/m ↓ |
|---|---|---|---|---|---|
| ResNet-101 + Weighted-InfoNCE | 58.10% | -- | 69.98% | 82.64% | 371.78 |
| ViT-Base/16 + InfoNCE | 65.89% | 93.09% | 77.84% | 86.52% | 196.59 |
| SwinV2-Base + Weighted-InfoNCE | 81.73% | -- | 88.32% | 87.35% | 196.06 |
| ConvNeXt-Base + Weighted-InfoNCE | 83.94% | -- | 89.54% | 87.98% | 160.49 |
| ViT-Base/16 + Weighted-InfoNCE | 85.23% | 97.62% | 90.52% | 88.96% | 145.64 |
| Ours | 88.13% | 98.79% | 92.41% | 90.68% | 84.50 |
| Method | R@1 ↑ | R@5 ↑ | AP ↑ | SDM@3 ↑ | Dis@1/m ↓ |
|---|---|---|---|---|---|
| Global Retrieval | 85.23% | 97.62% | 90.52% | 88.96% | 145.64 |
| k-reciprocal Reranking | 84.68% | 97.35% | 90.06% | 88.53% | 156.80 |
| R2Former-adapted | 85.74% | 97.86% | 90.86% | 89.05% | 146.20 |
| LOCORE-adapted | 86.08% | 98.03% | 91.10% | 89.22% | 144.35 |
| Ours: CL-LMM + Calibration | 88.13% | 98.79% | 92.41% | 89.71% | 138.96 |
| Ours: Full Method | 88.13% | 98.79% | 92.41% | 90.68% | 84.50 |
| Method | Type | SDM@1 ↑ | SDM@3 ↑ | Dis@1/m ↓ | Dis@3/m ↓ |
|---|---|---|---|---|---|
| Top-1 tile center | Post-reranking tile-center baseline | 91.52% | 89.71% | 138.96 | 167.74 |
| SliceMatch-adapted | Geometry-guided pose refinement | 92.54% | 90.22% | 113.80 | 153.40 |
| HC-Net-adapted | Homography-based coordinate refinement | 92.82% | 90.36% | 107.60 | 150.80 |
| GLVL-adapted | Global-local fine-grained matching | 93.06% | 90.49% | 101.40 | 148.90 |
| LocFusion w/o residual head | Candidate-set residual correction | 93.18% | 90.47% | 98.80 | 148.50 |
| LocFusion w/o weight head | Candidate-set residual correction | 93.04% | 90.39% | 102.30 | 150.10 |
| Ours: Full LocFusion | Candidate-set residual correction | 93.88% | 90.68% | 84.50 | 145.62 |
| Configuration | Extra Params | Online Latency | Peak Memory | R@1 | R@5 | AP | SDM@1 | Dis@1/m |
|---|---|---|---|---|---|---|---|---|
| Global Retrieval | 0 | 13.1 ms | 1.80 GB | 85.23% | 97.62% | 90.52% | 90.19% | 145.64 |
| +Matching head only | 0.92 M | 35.6 ms | 2.45 GB | 86.18% | 98.06% | 91.15% | 90.71% | 143.50 |
| + Matching + Strict-positive | 0.96 M | 35.8 ms | 2.47 GB | 86.74% | 98.24% | 91.55% | 90.96% | 142.20 |
| + Matching + Strict-positive + Prior | 1.05 M | 36.2 ms | 2.52 GB | 87.18% | 98.36% | 91.83% | 91.12% | 141.30 |
| Full CL-LMM | 1.15 M | 36.6 ms | 2.56 GB | 87.51% | 98.51% | 92.07% | 91.28% | 140.80 |
| Full CL-LMM +Calibration | 1.38 M | 37.6 ms | 2.61 GB | 88.13% | 98.79% | 92.41% | 91.52% | 138.96 |
| Full Method +LocFusion | 1.72 M | 38.3 ms | 2.67 GB | 88.13% | 98.79% | 92.41% | 93.88% | 84.50 |
| K | R@1 | R@5 | AP | SDM@1 | SDM@3 | Dis@1 | Dis@3 |
|---|---|---|---|---|---|---|---|
| 20 | 87.78% | 98.55% | 91.48% | 91.28% | 89.45% | 141.85 | 169.80 |
| 50 | 88.02% | 98.71% | 92.04% | 91.43% | 89.62% | 140.18 | 168.60 |
| 100 | 88.13% | 98.79% | 92.41% | 91.52% | 89.71% | 138.96 | 167.74 |
| 150 | 88.10% | 98.77% | 92.38% | 91.50% | 89.69% | 139.35 | 168.10 |
| 200 | 88.05% | 98.77% | 92.34% | 91.47% | 89.67% | 139.78 | 168.45 |
| S | SDM@1 | SDM@3 | Dis@1/m | Dis@3/m |
|---|---|---|---|---|
| 10 | 92.86% | 90.14% | 95.32 | 148.96 |
| 20 | 93.41% | 90.49% | 89.58 | 146.78 |
| 30 | 93.75% | 90.57% | 85.34 | 145.76 |
| 40 | 93.88% | 90.68% | 84.50 | 145.62 |
| 50 | 93.81% | 90.63% | 85.04 | 146.15 |
| 60 | 93.56% | 90.45% | 87.68 | 147.42 |
| 80 | 92.88% | 90.05% | 94.66 | 151.05 |
| Method | R@1 | R@5 | AP | SDM@1 | SDM@3 | Dis@1/m | Dis@3/m |
|---|---|---|---|---|---|---|---|
| GlobalRetrieval | 80.30% | 93.94% | 86.44% | 87.69% | 81.11% | 169.00 | 323.99 |
| +CL-LMM | 82.83% | 94.95% | 88.20% | 88.56% | 82.27% | 145.76 | 306.42 |
| +CL-LMM + Calibration | 83.33% | 95.20% | 88.51% | 88.79% | 81.63% | 143.65 | 305.47 |
| +CL-LMM + Calibration + LocFusion | 83.33% | 95.20% | 88.51% | 90.05% | 82.26% | 129.47 | 300.98 |
| Method | R@1 | R@5 | AP | SDM@1 | SDM@3 | Dis@1/m | Dis@3/m |
|---|---|---|---|---|---|---|---|
| GlobalRetrieval | 43.18% | 66.42% | 52.65% | 63.85% | 54.10% | 965.40 | 1336.20 |
| +CL-LMM | 44.46% | 66.80% | 53.47% | 64.42% | 54.68% | 895.30 | 1284.60 |
| +CL-LMM + Calibration | 44.66% | 66.88% | 53.64% | 64.58% | 54.41% | 872.80 | 1301.50 |
| +CL-LMM + Calibration + LocFusion | 44.66% | 66.88% | 53.64% | 65.92% | 55.36% | 768.50 | 1216.40 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, J.; Liu, Y.; Li, Q.; Xu, D.; Wang, T. UAV Visual Localization Method Based on Token-Level Local Matching Reranking and Neighborhood-Consistent Position Fusion. Remote Sens. 2026, 18, 3016. https://doi.org/10.3390/rs18173016
Liu J, Liu Y, Li Q, Xu D, Wang T. UAV Visual Localization Method Based on Token-Level Local Matching Reranking and Neighborhood-Consistent Position Fusion. Remote Sensing. 2026; 18(17):3016. https://doi.org/10.3390/rs18173016
Chicago/Turabian StyleLiu, Jiaxin, Yunqing Liu, Qi Li, Dongpo Xu, and Tao Wang. 2026. "UAV Visual Localization Method Based on Token-Level Local Matching Reranking and Neighborhood-Consistent Position Fusion" Remote Sensing 18, no. 17: 3016. https://doi.org/10.3390/rs18173016
APA StyleLiu, J., Liu, Y., Li, Q., Xu, D., & Wang, T. (2026). UAV Visual Localization Method Based on Token-Level Local Matching Reranking and Neighborhood-Consistent Position Fusion. Remote Sensing, 18(17), 3016. https://doi.org/10.3390/rs18173016

