1.1. Research Background and Significance
Land use/land cover change (LUCC) represents the most direct and profound manifestation of human activities acting on the Earth’s surface system and constitutes a fundamental component of global change research [
1,
2,
3]. Since the beginning of the 21st century, rapid population growth, industrialization, urbanization, and intensified energy resource exploitation have continuously increased the intensity of land resource utilization worldwide. Consequently, natural ecosystems, including forests, grasslands, wetlands, and croplands, have undergone extensive transformations, leading to substantial changes in ecosystem structure, ecological processes, and ecosystem service functions [
2,
4]. Numerous studies have demonstrated that LUCC not only directly alters land cover patterns but also exerts profound impacts on regional and global climate change, biodiversity conservation, and ecological security by modifying vegetation composition, surface albedo, water and energy exchanges, carbon cycling, and biogeochemical cycles. As a major branch of global environmental change research, Land Change Science (LCS) has gradually developed into a relatively comprehensive theoretical framework in recent decades. Turner et al. proposed that land-use change is not only the result of interactions between natural processes and human activities but also serves as a critical link between socioeconomic development and ecological environmental evolution. Accordingly, the focus of land-change research has gradually shifted from simply describing land-use changes to revealing their driving mechanisms, ecological environmental effects, and sustainable management strategies. Therefore, accurately characterizing the spatiotemporal dynamics of land use and examining their associations with ecological environmental conditions have become important research priorities in geography, ecology, and remote sensing.
With increasingly severe global ecological and environmental challenges, land-use change is no longer regarded as merely a local environmental issue but is widely recognized as one of the major drivers of global environmental change. Foley et al. pointed out that human activities have altered nearly half of the Earth’s terrestrial surface, making land-use change a major driving force affecting global ecosystem services, food security, freshwater availability, and climate regulation [
2]. Meanwhile, large-scale land development has resulted in continuous degradation of natural ecosystems and declining ecosystem resilience, thereby intensifying the conflict between land resource utilization and ecological environmental protection. Consequently, under the framework of global sustainable development and climate change mitigation, establishing high-accuracy, long-term land-use monitoring systems and systematically evaluating the ecological environmental effects induced by land-use change have become major research priorities in the international remote sensing community [
5].
Resource-based cities serve as important bases for energy and mineral resource exploitation and play an irreplaceable role in ensuring national energy security and promoting regional economic development [
6,
7,
8]. Resource exploitation is typically characterized by high intensity, long duration, and spatial concentration [
7,
8,
9]. While facilitating rapid industrialization and urbanization, it has also profoundly reshaped regional land-use patterns and ecological environmental structures. Coal resource-based cities, in particular, have experienced extensive occupation of natural ecological space due to large-scale open-pit mining, underground coal extraction, waste dump construction, and industrial infrastructure expansion. These activities have intensified landscape fragmentation, reduced vegetation cover, increased water consumption, and weakened ecosystem service functions, thereby exacerbating the conflict between resource exploitation and ecological conservation. In recent years, with the accelerating global transition toward green and low-carbon energy systems, achieving ecological restoration and high-quality development while maintaining energy security has become a major scientific issue attracting widespread attention from both the international academic community and policymakers [
6,
10].
China is the world’s largest producer and consumer of coal, with coal remaining the dominant source of national energy supply for decades [
10]. According to the National Sustainable Development Plan for Resource-Based Cities (2013–2020) and the Implementation Plan for Promoting High-Quality Development of Resource-Based Regions during the 14th Five-Year Plan Period, China has 262 resource-based cities [
6], nearly half of which are coal resource-based cities. These cities are primarily distributed across northern China’s energy bases, the ecological conservation region of the Yellow River Basin, and other nationally strategic energy development areas. Although long-term coal resource exploitation has provided strong support for China’s rapid socioeconomic development, it has also caused a series of environmental problems, including land subsidence, mining area expansion, groundwater depletion, vegetation degradation, and ecosystem impairment, which have increasingly constrained regional sustainable development [
7,
9]. Particularly under China’s “Dual Carbon” goals, coordinating resource exploitation, territorial spatial optimization, and ecological restoration has become a central task for the green transformation of resource-based cities.
Ordos City, located in southwestern Inner Mongolia Autonomous Region, is one of China’s most important coal production bases and an integral part of the National Modern Energy Economy Demonstration Zone [
10]. The region possesses approximately one-sixth of China’s predicted coal reserves and has developed an industrial system centered on coal mining, coal chemical industries, and energy equipment manufacturing, making it strategically important for safeguarding national energy security [
11]. However, the combined effects of long-term intensive coal exploitation and rapid urbanization have substantially reshaped regional land-use patterns. Large areas of natural grassland and unused land have been converted into mining areas, industrial land, and built-up land. In some areas, coal mining subsidence, groundwater drainage, and waste dump construction have reduced ecosystem stability, resulting in increasingly prominent ecological and environmental problems such as land degradation, water body shrinkage, and vegetation fragmentation. On the other hand, the continuous implementation of major ecological restoration programs, including the Grain for Green Program, natural forest conservation, green mine construction, and mine ecological restoration, has contributed to a gradual recovery of regional ecological environmental quality in recent years [
12,
13]. Therefore, Ordos not only represents a typical case of rapid land-use transformation driven by resource exploitation but also reflects the process of ecosystem reconstruction under the combined effects of ecological restoration and natural recovery. It thus provides an ideal study area for investigating the relationship between resource exploitation and ecological environmental responses, as well as for evaluating the applicability of land-use classification models in complex mining environments.
1.2. Progress in Deep Learning-Based Land-Use Classification Research
Previous studies have conducted extensive research on land-use change, ecological environmental quality assessment, and ecosystem services in Ordos City, yielding a substantial body of results [
11,
12,
13]. However, most of these studies rely on national-scale land-use products for analysis [
14]. Although such datasets can effectively capture the overall regional trends of land-use change, their classification schemes typically merge mining areas into built-up land, making it difficult to accurately characterize the expansion of mining areas and its associated ecological environmental effects under coal resource exploitation. Meanwhile, mining areas, bare land, built-up land, and unused land often exhibit highly similar spectral characteristics, and complex mining environments are commonly affected by the phenomena of “spectral confusion and intraclass spectral variability” [
15,
16]. As a result, traditional land-use products still have limitations in accurately identifying mining areas. Therefore, developing a refined land-use classification system tailored to resource-based cities and establishing high-precision remote sensing classification methods suitable for complex mining environments has become an important direction in current land-use remote sensing research.
With the rapid development of Earth observation technologies, the acquisition capability of multi-source remote sensing data—represented by Landsat, Sentinel, Gaofen (GF) series, and commercial high-resolution satellites—has been continuously enhanced, providing rich data support for regional-scale dynamic land-use monitoring [
17,
18]. In particular, since its launch in 1972, the Landsat program has generated the longest continuous time-series of global remote sensing observations, offering a unified data foundation for long-term LUCC studies [
17,
19,
20]. In recent years, the development of cloud computing platforms such as Google Earth Engine has further improved the capability of processing remote sensing big data, making long time-series land-use monitoring an increasingly important direction in global land change research [
21]. However, while spatial, temporal, and spectral resolutions of remote sensing data continue to improve, traditional classification methods are facing increasingly significant challenges in complex land surface environments [
15,
22,
23].
Early land-use classification primarily relied on traditional statistical classification methods such as maximum likelihood classification and ISODATA. These approaches mainly utilize pixel-based spectral features for discrimination and require relatively strict assumptions regarding sample distribution and data quality [
24]. When spectral overlap among different land-cover types is significant, classification accuracy tends to decrease [
25,
26]. Subsequently, machine learning methods such as support vector machines and random forests were gradually introduced into remote sensing image classification. Compared with traditional statistical approaches, machine learning methods can better exploit multi-dimensional spectral information and texture features, achieving improved performance in land-cover classification tasks with low to moderate complexity. As a result, they once became the dominant technical paradigm in land-use classification research. However, these methods still rely on manually designed features and have limited ability to utilize spatial contextual information and high-level semantic features [
27]. In complex environments such as mining areas, urban fringes, and heterogeneous bare land, they often suffer from a pronounced “salt-and-pepper effect,” making it difficult to meet the requirements of fine-scale land-use mapping.
In recent years, the development of deep learning has driven remote sensing image classification into the era of semantic segmentation [
28,
29,
30]. Long et al. first proposed the Fully Convolutional Network (FCN), enabling end-to-end pixel-wise classification and overcoming the limitation of traditional convolutional neural networks that can only perform image-level classification, thereby laying the theoretical foundation for semantic segmentation in remote sensing imagery [
27]. Subsequently, Badrinarayanan et al. proposed the SegNet model, which improves spatial information recovery capability through an encoder–decoder architecture and has achieved promising performance in tasks such as road extraction and scene segmentation [
31]. Ronneberger et al. further introduced the U-Net model, which incorporates skip connections to effectively fuse low-level spatial details with high-level semantic features [
32]. Even under limited training samples, it maintains high classification accuracy, and thus has rapidly become one of the most widely used classical models in medical imaging and remote sensing semantic segmentation.
With further advances in research, many improved U-Net-based models have been proposed. Zhou et al. introduced UNet++, which redesigns the skip connection pathways to enhance feature fusion efficiency between the encoder and decoder, thereby effectively improving object boundary delineation capability [
33]. Oktay et al. proposed Attention U-Net, which incorporates an attention gating mechanism during feature fusion, enabling the network to focus more on target regions, suppress background noise interference, and improve small-object recognition accuracy in complex scenes [
16]. Diakogiannis et al. developed the ResUNet-a model by integrating residual connections, multi-scale convolutions, and atrous (dilated) convolutions, further enhancing the model’s ability to represent complex spatial structures and multi-scale objects, and achieving high accuracy in remote sensing land-cover classification tasks [
34,
35]. Meanwhile, the DeepLab series has also undergone continuous development. Chen et al. proposed DeepLabV3+, which employs atrous spatial pyramid pooling (ASPP) to capture multi-scale contextual information and further integrates an encoder–decoder structure to improve object boundary refinement, making it one of the representative models in remote sensing semantic segmentation [
36].
In recent years, the development of Transformer architectures has further advanced remote sensing intelligent interpretation techniques. Dosovitskiy et al. first proposed the Vision Transformer (ViT) [
37,
38,
39], which leverages a self-attention mechanism to establish global spatial dependencies, effectively overcoming the limitation of local receptive fields in convolutional neural networks. Subsequently, Liu et al. introduced the Swin Transformer, which achieves hierarchical feature extraction through a shifted window-based self-attention mechanism [
40]. While maintaining computational efficiency, it significantly improves the performance of remote sensing image classification and object recognition, and has become an important research direction in high-resolution remote sensing interpretation.
A large body of studies has demonstrated that deep learning-based methods, particularly those combining CNNs and Transformers, have become a major trend in current remote sensing land-use classification research, exhibiting significantly superior classification accuracy and generalization capability compared with traditional machine learning methods [
23,
40,
41].
Although deep learning has significantly improved the accuracy of land-use classification, complex mining environments remain a major challenge in remote sensing semantic segmentation. On the one hand, mining areas exhibit highly complex surface compositions, where open pits, waste dumps, industrial facilities, roads, bare surfaces, and restored vegetation are interwoven, resulting in pronounced spatial heterogeneity. On the other hand, mining areas, built-up land, bare land, and unused land share similar spectral characteristics, while different mining stages can lead to distinct spectral responses for the same land-cover type, forming typical cases of “spectral ambiguity for different objects” and “spectral heterogeneity for identical objects,” which often cause class confusion in traditional deep learning models. In addition, mining areas are characterized by complex boundary geometries and large variations in object scales, placing higher demands on feature extraction capability. Therefore, relying solely on the classical U-Net model is often insufficient to fully capture discriminative features in complex industrial and mining environments.
As shown in
Table 1, previous studies have examined mining-area monitoring, land-use change, ecological quality assessment, and deep learning-based land-cover classification. However, relatively few studies have explicitly identified mining areas as an independent class when examining long-term land-use change patterns in coal resource-based cities. When mining areas are merged with built-up land or unused land, mining-related land conversions may be obscured, limiting the interpretation of long-term mining expansion, reclamation, and their association with land-cover-based ecological quality changes.
Recent studies have begun to explore the application of deep learning to mining-area land-use classification and ecological environmental monitoring. For example, some studies have employed models such as DeepLab, UNet++, and HRNet for land-cover mapping in mining areas, which have improved the accuracy of boundary delineation to a certain extent [
33,
43,
44,
45]. Other studies have combined high-resolution UAV imagery to conduct land-use classification in mining regions, providing new technical approaches for ecological restoration assessment in these areas. However, most existing studies primarily focus on improving the accuracy of classification models themselves, while paying insufficient attention to the relationships among classification results, land-use change analysis, and ecological environmental quality assessment. Meanwhile, widely used global land-cover products such as Globeland30 and ESA WorldCover adopt unified classification systems that typically merge mining areas into built-up land [
12]. Although these products can meet the needs of global or national-scale land-cover mapping, they fail to accurately represent the expansion process of mining areas and its associated land-cover-based ecological quality patterns in coal resource-based cities. In addition, most existing studies rely on relatively short time-series data and lack long-term dynamic analyses spanning more than two decades, resulting in limited understanding of the long-term “trade-off” process between resource exploitation and ecological restoration.
In summary, current research on land-use classification still has several limitations. (1) Existing publicly available land-cover products usually adopt relatively coarse classification systems and lack refined representations that independently distinguish mining areas in resource-based cities. Mining areas are often merged with built-up land, bare land, or unused land, making it difficult to accurately quantify mining-related land-use change. (2) Many deep learning studies focus mainly on improving classification accuracy itself, while less attention has been paid to whether the classification results can support subsequent land-use transition analysis and ecological quality assessment in mining regions. (3) Land-use classification, land-use change analysis, and ecological environmental assessment are often treated as separate research components, and the link between fine-grained mining-area classification, long-term land-use evolution, and land-cover-based ecological quality change remains insufficiently explored [
46].
This study addresses the question of how explicitly mapping mining areas as a separate class, rather than merging them with built-up land or unused land, changes the interpretation of long-term land-use change patterns in Ordos. Using land-cover maps from six time points between 2000 and 2025, we trace land-use transitions involving mining areas and examine their associated land-cover-based ecological quality patterns using the ecological environmental quality index (EQI) and ecological contribution index (LEI). The U-Net-based framework is used to distinguish mining areas and other spectrally confused land-cover types; it is not intended as a new general-purpose segmentation architecture.