Next Article in Journal
Surface Subsidence Monitoring and Interpretable Factor Analysis in Coal Mining Areas of Henan Province Based on SBAS-InSAR
Previous Article in Journal
A Coordinate-Based Framework for Sea Surface Wind Speed Reconstruction from Sparse Multi-Source Observations
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Vision Transformer with Dynamic Masking and Cross-Modal Semantic Learning for Remote Sensing Scene Classification

1
School of Automation, Nanjing University of Science and Technology, Nanjing 210094, China
2
Beijing Microelectronics Technology Institute, Beijing 100029, China
3
National Key Laboratory of Information Systems Engineering, Nanjing Research Institute of Electronic Engineering, Nanjing 210007, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(16), 2710; https://doi.org/10.3390/rs18162710
Submission received: 29 June 2026 / Revised: 23 July 2026 / Accepted: 31 July 2026 / Published: 12 August 2026

Abstract

Remote sensing scene classification plays a vital role in various Earth observation applications. Although supervised learning remains the dominant paradigm, vast quantities of unlabeled imagery remain significantly underutilized. To leverage these unlabeled resources and enhance categorization accuracy, we propose a novel framework based on the Vision Transformer (ViT) that integrates a dynamic masking strategy with a cross-modal semantic learning mechanism. Specifically, a dynamic masking strategy guided by a smooth reconstruction loss is designed to learn robust feature representations from unlabeled samples prior to downstream fine-tuning. Furthermore, we incorporate cross-modal learning to enrich semantic information, thereby addressing the inherent supervisory limitations of conventional one-hot labels. Comprehensive experiments demonstrate that the proposed method significantly improves classification accuracy while maintaining high pre-training efficiency and strong generalization capabilities.
Keywords: remote sensing image; scene classification; transformer; masked autoencoder; cross-modal semantic learning remote sensing image; scene classification; transformer; masked autoencoder; cross-modal semantic learning

Share and Cite

MDPI and ACS Style

Ni, F.; Liu, Y.; Dai, S.; Chen, L.; Feng, C.; Zhang, F.; Wu, X.; Bo, Y. A Vision Transformer with Dynamic Masking and Cross-Modal Semantic Learning for Remote Sensing Scene Classification. Remote Sens. 2026, 18, 2710. https://doi.org/10.3390/rs18162710

AMA Style

Ni F, Liu Y, Dai S, Chen L, Feng C, Zhang F, Wu X, Bo Y. A Vision Transformer with Dynamic Masking and Cross-Modal Semantic Learning for Remote Sensing Scene Classification. Remote Sensing. 2026; 18(16):2710. https://doi.org/10.3390/rs18162710

Chicago/Turabian Style

Ni, Feng, Yi Liu, Shibo Dai, Lei Chen, Changlei Feng, Fan Zhang, Xiang Wu, and Yuming Bo. 2026. "A Vision Transformer with Dynamic Masking and Cross-Modal Semantic Learning for Remote Sensing Scene Classification" Remote Sensing 18, no. 16: 2710. https://doi.org/10.3390/rs18162710

APA Style

Ni, F., Liu, Y., Dai, S., Chen, L., Feng, C., Zhang, F., Wu, X., & Bo, Y. (2026). A Vision Transformer with Dynamic Masking and Cross-Modal Semantic Learning for Remote Sensing Scene Classification. Remote Sensing, 18(16), 2710. https://doi.org/10.3390/rs18162710

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop