Next Article in Journal
DC4Former: Orientation-Stable UAV Disaster Image Segmentation via Diagonal-Complemented C4 Consistency
Previous Article in Journal
Adaptive Fusion of Multiple Land-Cover Products for Improved Spatial Representation of Key Land Classes in Central Asia
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Benchmarking Open-Access Building Footprints: A Multi-Dimensional Assessment with High-Fidelity References

1
School of Remote Sensing and Information Engineering, Wuhan University, Wuhan 430079, China
2
School of Geography and Information Engineering, China University of Geosciences (Wuhan), Wuhan 430074, China
3
Joint Laboratory of Spatial Intelligent Perception and Large Model Application, Nanjing University of Aeronautics and Astronautics, Nanjing 210016, China
4
Huantian Wisdom Technology Co., Ltd., Meishan 620010, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(17), 2895; https://doi.org/10.3390/rs18172895
Submission received: 6 July 2026 / Revised: 21 August 2026 / Accepted: 22 August 2026 / Published: 27 August 2026

Highlights

What are the main findings?
  • Open-access building footprint datasets generally underestimate building count and area; object recovery and spatial agreement remain limited, while matched-pair morphology fidelity varies substantially.
  • No product led all three tiers: GABLE ranked highest in statistical consistency, Microsoft Global ML Building Footprints in object recovery and spatial fidelity, and Google Research Open Buildings in matched-pair morphology fidelity within the evaluated coverage.
What are the implications of the main findings?
  • Local validation is advisable when reliable completeness, positional accuracy, or boundary fidelity is required.
  • Tier-specific rankings support selection of the most suitable available product for a given region and application.

Abstract

The rapid proliferation of open-access vector building footprint datasets has created substantial opportunities for large-scale urban analysis, yet their reliability remains difficult to assess effectively and consistently. This challenge arises primarily from the scarcity of large-scale, high-fidelity reference benchmarks across diverse urban environments. To address these limitations, we construct a benchmark-grade reference dataset comprising over 200,000 manually annotated building footprints across 10 purposively selected cities across diverse global regions through an exhaustive curation campaign and strict manual-automated quality validation. Using this dataset, we develop a three-tier evaluation framework that assesses statistical consistency, object recovery and spatial accuracy, and morphology fidelity among matched building pairs. Evaluating 11 mainstream datasets reveals pervasive underestimation in both building area (average bias: −14.9%) and count (−34.0%). Average city-level intersection-over-union values range from 29.1% to 58.3%, indicating low-to-moderate spatial overlap. For matched building pairs, the datasets demonstrate relatively high geometric fidelity, with morphological similarity scores of 0.495 to 0.803. Pronounced regional performance variations indicate that no single dataset is universally optimal. Specifically, Microsoft Global ML Building Footprints offers the most balanced multi-city performance, Google Research Open Buildings is recommended for boundary shape fidelity, and regional datasets remain competitive for local analyses.

1. Introduction

Driven by advances in remote sensing and crowd-sourced data collection, the volume of open geospatial data has expanded exponentially, facilitating extensive application across diverse disciplines [1,2]. Among these resources, building footprint vectors, spanning regional to global scales, have emerged as foundational datasets for modern urban studies and applications [3]. These datasets are utilized extensively in research concerning pedestrian accessibility assessment [4,5], urban form analysis [6], rooftop photovoltaic potential estimation [7], and the monitoring of Sustainable Development Goals [8].
However, the exponential growth of data volume [9], combined with the accelerated rate of urban transformation, has resulted in substantial variability concerning the quality of existing public building vector datasets [10,11]. Consequently, identifying the dataset that optimally aligns with a specific task or application context remains a significant challenge for researchers and practitioners [12,13]. At present, comprehensive evaluation of public building vector datasets remains limited in scope. Therefore, a unified and systematic reliability assessment of mainstream public building vector datasets would substantially assist users in making informed dataset selections tailored to specific research or application requirements [13].
A major challenge in conducting comprehensive evaluations is the acquisition of high-quality reference data [14]. Currently, the most reliable approach for establishing ground truth relies heavily on manual annotation across extensive geographic areas [15]. However, constrained by prohibitive time and budget limitations, available reference datasets typically present an inherent trade-off, falling short in either broad spatial scope or surveyor-level geometric accuracy [16]. Consequently, high-precision reference data produced at a scale sufficient to evaluate diverse urban morphological contexts remains critically scarce [17].
Beyond the limitations of reference data, existing evaluation methodologies often adopt a simplified assessment paradigm that does not differentiate the multi-faceted quality requirements of downstream tasks [18]. Current paradigms predominantly rely on generic overlap-based metrics from computer vision and semantic segmentation [19]. While these indicators provide basic statistics for spatial alignment, they struggle to capture the distinct aspects of data quality—such as macro-scale capacity or micro-scale morphology—that are critical to varying urban analysis tasks [18]. Consequently, relying solely on overlap metrics is often insufficient to fully characterize the specific quality attributes demanded by diverse downstream urban applications [17].
To address these gaps, this study constructs a high-quality, manually annotated reference dataset and evaluates 11 mainstream public building footprint datasets across 10 representative cities worldwide. These cities span diverse national contexts, development stages, and urban morphological characteristics. Using this reference dataset, we develop a three-tier evaluation framework integrating statistical consistency, object recovery and spatial accuracy, and morphology fidelity among matched building pairs, thereby supporting task-specific dataset selection and targeted quality improvement.
Quantitatively, the evaluated datasets show a general underestimation in both building area and building count, with average area and count biases of −14.9% and −34.0%, respectively; ONEGEO exhibits the strongest area underestimation (−38.2%), while Google Research Open Buildings slightly overestimates area (9.60%). In terms of spatial accuracy, the mean object-level intersection-over-union (IoU) ranges from 45.5% to 62.4%, and the centroid root mean square error (RMSE) ranges from 4.85 m to 9.02 m, indicating low-to-moderate footprint overlap. Overall, despite systematic biases in area and quantity, the datasets demonstrate relatively high geometric fidelity for successfully matched building pairs, yielding morphological similarity scores between 0.495 and 0.888. The results reveal pronounced regional performance variations, and clarify the task-specific suitability of different products. Within the sampled coverage, GABLE achieved the highest composite score, Google Research Open Buildings had the largest observed matched-pair morphology score within its evaluated city subset, and regional products remained competitive where their coverage matched the study area.
This study makes three primary contributions:
(1) We establish a high-precision, expert-delineated reference dataset comprising over 200,000 building footprints across 10 purposively selected cities across diverse global regions, serving as a reliable benchmark-grade ground truth for validating large-scale building products.
(2) We develop a reproducible three-tier framework that integrates macro-scale statistical consistency, object recovery and spatial fidelity, and morphology fidelity among matched building pairs, while retaining the component scores needed for task-specific interpretation.
(3) We publicly release the benchmarking code and detailed evaluation results, providing reusable tools and transparent evidence for dataset selection, comparative analysis, and the future evaluation of additional building-footprint products.

2. Related Works

2.1. Public Vector Building Datasets and Reference Data

Over the past decade, advances in Earth observation technologies and geographic information systems have substantially expanded the availability of open-access vector datasets [1,2]. Specifically, Microsoft Global ML Building Footprints (MBF) [20] and Google Research Open Buildings (GOB) [21] provide immense data volumes and high-resolution outlines across broad geographic extents; however, these products remain subject to algorithmic omissions and topological inconsistencies in complex environments. Crowdsourced and commercial platforms, including OpenStreetMap (OSM) [22], Mapbox [23], and ONEGEO [24], offer continuous global updates. Nevertheless, heterogeneous annotation standards and practices result in considerable regional differences in spatial accuracy. Furthermore, regional datasets like CMAB [25], GABLE [26], 3D-GloBFP [27], East Asian Buildings [28], and China 90Cities Rooftop Data [29] (China-90C) deliver exceptional localized geometric fidelity, though their limited geographic scope constrains global-scale urban applications. Table 1 summarizes the 11 open-access building footprint datasets selected for benchmarking in this study, along with their key attributes.
In terms of dataset evaluation, reliability depends fundamentally on the reference data used for comparison. Based on the reference sources and validation designs reported in previous studies, existing assessment strategies can be broadly grouped into three categories: indirect or proxy-based assessment, evaluation against authoritative reference data, and evaluation using purpose-built manual annotations [8,31,32,33,34,35,36]. Indirect assessment does not rely on independently verified building-level reference data, instead using ancillary indicators or cross-product comparisons. For example, Herfort et al. estimated OSM building completeness across 13,189 urban agglomerations using a machine-learning framework based on ancillary covariates rather than verified footprint geometries [8]. Chamberlain et al. identified substantial cross-product differences in building counts, area, coverage, and completeness, underscoring the need for independent evaluation of large-scale building footprint datasets [31]. Although these approaches are well suited to broad geographical analysis, they generally provide indirect rather than direct evidence of geometric and morphological agreement at the individual-footprint level.
Authoritative records, such as the cadastral and government mapping products used by Minghini et al. [32] and Brovelli and Zamboni [33] for continental and regional comparisons—offer practical convenience, yet remain subject to heterogeneous quality standards and varying degrees of temporal consistency with the evaluated products. Purpose-built manual annotations provide greater control over reference definitions and geometric quality but have generally been restricted to narrower geographical areas. Nyandwi et al. constructed reference data within Rwanda [34]; Okyere et al. manually delineated footprints across five metropolitan areas [35]; and Fan et al. examined localized urban samples through multiscale shape-similarity analysis [36]. While these efforts yield valuable localized insights, their limited spatial coverage restricts the generalizability of conclusions across diverse urban morphologies and development contexts.
Across these three reference strategies, few studies explicitly document rigorous quality-control procedures, such as independent review, repeated inspection, or systematic inter-annotator consistency checks [37]. This omission makes it difficult to assess the reliability of the reference standard itself. Table 2 summarizes the reference data characteristics of representative building footprint evaluation studies. Compared with the representative studies summarized in Table 2, this study combines expert-delineated annotations from 10 cities across diverse geographical and urban contexts, more than 200,000 reference footprints, and a systematic visual–algorithmic quality-assurance procedure.

2.2. Vector Dataset Quality Assessment

To comprehensively characterize the quality of vector spatial data, existing evaluation paradigms employ various methodological lenses, spanning macro-scale aggregate statistics, object-level spatial agreement, and boundary-level morphological assessment [32,38].
Regarding macro-scale evaluation, prior studies predominantly used aggregate metrics to quantify total building area or building count within administrative units or analysis grids [39]. Foundational work by Hecht et al. established the measurement of building footprint completeness using unit-based and grid-based counting of total area and building numbers [40]. Building on this statistical framework, subsequent studies have quantified structural completeness and mapped macroscopic density distributions to identify systematic dataset biases across large-scale urban extents [41,42].
At the object level, localized alignment is commonly evaluated using region-based overlap metrics and boundary-distance measures. Precision, recall, and the F1 score quantify detection performance but do not fully describe positional or geometric agreement. To advance instance-level evaluation, Chen et al. employed IoU variants to robustly quantify overall spatial coverage and classification accuracy [43]. Object-based comparison can also reveal inconsistent segmentation in dense urban blocks, where adjacent buildings may be merged or individual structures divided into multiple polygons. Concurrently, boundary-level distance measurements are utilized to quantify spatial divergence between predicted contours and ground truth polygons. Foundational work by Huttenlocher et al. established point-to-point deviation assessment using the Hausdorff distance [44], while other representative assessments frequently utilize conventional RMSE to measure overall positional shifts. Zhang et al. proposed BANet to improve hierarchical feature integration and scene adaptability, highlighting the importance of multiscale structural and boundary information for the geometric quality evaluation [45].
Beyond spatial displacement, the evaluation of structural shape fidelity employs complex geometric descriptors to assess morphological consistency in the frequency and sequence domains. Addressing complex local deformations, Avbelj et al. introduced the Polygon-Line Spatial (PoLiS) distance, which offers a robust and symmetric evaluation of subtle contour shifts [38]. Furthermore, Arkin et al. and Ai et al. utilized turning functions and Fourier descriptors, respectively, enabling robust, scale-invariant shape similarity comparisons [46]. Additionally, Basaraner and Cetinkaya incorporated a comprehensive suite of geometric shape indices, including compactness-related measures, to capture the holistic morphological attributes of geographic entities [47].
Although these metrics effectively quantify specific geometric properties, existing validation paradigms generally deploy them in isolation. A more comprehensive evaluation would substantially benefit from systematically integrating statistical consistency, object recovery and spatial accuracy, and matched-pair morphology assessment within a unified three-tier framework.

3. Materials and Methods

3.1. Construction of the Reference Dataset

To construct a reliable reference dataset for quantitative evaluation, trained professional interpreters manually delineated building footprints following standardized annotation and quality-control procedures. This process produced a high-fidelity vector dataset comprising over 200,000 building footprints across 10 purposively selected cities across diverse global regions. The workflow included source-image screening, interpreter training and calibration, manual delineation under shared object and boundary rules, automated geometric and topological validation, blind peer review, and supervisory batch audit, as illustrated in Figure 1 and detailed below.
The source data comprised sub-meter GaoFen-02 (GF-02), GaoFen-07 (GF-07), Pleiades, and aerial imagery. Before delineation, each image was screened for cloud cover, heavy shadow, blur, and insufficient building visibility. Images in which building boundaries could not be interpreted reliably were not used. All interpreters received the same cartographic specifications and training before production.
During delineation, interpreters traced the visible outer boundary of each building and corrected identifiable roof-to-ground displacement. The specifications required consistent vertex placement, preservation of recognizable boundary detail, orthogonal regularization of rectilinear structures, and valid polygon topology. Adjacent structures were separated only when a visible boundary allowed them to be distinguished consistently. Small or partially occluded structures were included only when their outlines remained interpretable. The minimum mapping unit (MMU) was defined operationally by reliable boundary interpretation rather than by a single area threshold applied to all cities.
The digitized footprints underwent three quality-control stages. First, automated validation identified duplicate vertices, unclosed rings, self-intersections, unintended overlaps, and sliver polygons. Second, another interpreter independently checked building extent, local alignment, roof-offset correction, and topology against the source imagery. Nonconforming footprints were returned for redigitization. Third, senior supervisors inspected 20% of each submitted batch. The batch error rate was calculated as the proportion of inspected footprints containing at least one boundary, alignment, displacement, or topology error. A low error tolerance was necessary because the reference footprints supported object matching and boundary-based metrics, both of which are sensitive to delineation, alignment, and topology errors. Accordingly, batches with an error rate above 2% were corrected and reinspected before acceptance. The 2% cutoff served as an internal batch-rejection criterion for maintaining consistent production quality.
Following manual digitization and multi-stage quality assurance, the accepted reference labels were compiled into a building-footprint reference dataset comprising over 200,000 polygons across 10 cities. As shown in Figure 2, the study areas comprise Chongqing, Delingha, Durban, Edmonton, Cayenne, Hangzhou, Melbourne, Shanghai, St. Gallen, and Xuzhou. These cities were purposively selected to capture variation in geographical setting, urban morphology, architectural pattern, and topographic condition, thereby supporting evaluation across heterogeneous urban contexts.
The annotated areas were deliberately selected from established built-up zones rather than rapidly expanding urban fringes. To strictly mitigate temporal inconsistencies between our reference data and the varied timestamps of the evaluated products, we conducted visual cross-checks using historical Google Earth imagery spanning the respective temporal gaps. Any localized blocks exhibiting building construction or demolition were explicitly excluded from the reference extent. Consequently, the selected mature urban fabrics ensure highly stable building configurations, guaranteeing that the observed performance discrepancies stem from dataset quality rather than temporal mismatch. Table 3 summarizes the characteristics of the ten selected urban regions and their corresponding manually annotated building footprint reference data, and the number of evaluated open-access building footprint datasets covering each city.
As shown in Table 3, the dataset contains in total 201,511 building instances, with the number of reference buildings varying substantially across cities, from 3971 in Delingha to 50,485 in Melbourne. Furthermore, the datasets demonstrate substantial heterogeneity across the study areas in both urban morphology and dataset coverage, which is essential for a balanced evaluation. The study areas span mountainous megacities, river delta hubs, arid plateau settlements, and alpine cities.
Figure 3 demonstrates the spatial intersection and data availability between the 11 evaluated open-access datasets and the 10 reference benchmark cities. While global-scale vector building datasets (i.e., OSM) exhibit ubiquitous coverage across nearly all target sites, regional or crowdsourced alternatives present localized availability or distinct spatial gaps, necessitating that subsequent metrics be aggregated only over valid spatial pairs.

3.2. Multi-Dimensional Evaluation Framework

To systematically evaluate the reliability of open-access building footprint datasets, we formulate a multi-dimensional assessment framework. The framework evaluates datasets across three complementary tiers: statistical consistency, object recovery and spatial fidelity, and morphology fidelity among matched building pairs. For each dataset d , the three dataset-level tier scores contributed equally to the overall score:
S d = S 1 , d + S 2 , d + S 3 , d 3
where S d denotes the overall score of dataset d , whereas S 1 , d , S 2 , d , and S 3 , d denote its dataset-level scores for Tiers 1, 2, and 3, respectively. Each dataset-level tier score was obtained by averaging the corresponding city-level tier scores over all valid cities within the evaluated coverage. Missing city–dataset combinations were excluded from the averaging rather than assigned a score of zero.
Before equal-weight aggregation, the component indicators were converted to higher-is-better scores between 0 and 1 according to their original directions and scales. Indicators already expressed on a direction-aligned 0–1 scale were retained without further transformation. Bounded indicators with the opposite direction were reversed, whereas the remaining lower-is-better indicators were rescaled relative to the benchmark range using the following transformation:
s ( x ) = x m a x x x m a x x m i n
where x m i n and x m a x denote the benchmark-wide minimum and maximum values for the relevant indicator; the observations used to determine these bounds are specified in the corresponding tier subsection. All calculations retained full numerical precision, with rounding applied only to values presented in tables and figures. The resulting scores are therefore comparative summaries defined relative to the fixed benchmark population.

3.2.1. Statistical Consistency

Statistical consistency evaluates aggregate agreement in building count and footprint area. This tier is relevant for downstream applications such as urban population estimation, macroeconomic modeling, and regional building stock assessment, which rely on spatial capacity and structural density rather than precise boundary alignments [8,48,49].
To examine the spatial heterogeneity of statistical consistency, the study areas were further divided into local analysis grids. These grids were delineated primarily according to streets and natural spatial boundaries, rather than by a rigid regular grid, to better preserve the integrity of urban grids. All grids were manually inspected to ensure that no boundary intersected the ground-truth building footprints. For analysis grid u in city c and dataset d , signed count bias ( Δ N u c d ) and area bias ( Δ A u c d ) are calculated as [50]:
Δ N u c d = N O S , u c d N G T , u c d N G T , u c d Δ A u c d = A O S , u c d A G T , u c d A G T , u c d
where the subscripts G T and O S denote the reference and evaluated open-access data, respectively. N O S , u c d and N G T , u c d denote the numbers of evaluated and reference building objects, respectively, whereas A O S , u c d and A G T , u c d denote the corresponding total footprint areas within analysis grid u . A value of zero indicates exact aggregate agreement, whereas negative and positive values indicate underestimation and overestimation, respectively.
Because deviations in either direction indicate disagreement, the signed biases were converted to nonnegative error magnitudes at the grid level:
E N , u c d = | Δ N u c d | E A , u c d = | Δ A u c d |
For each city–dataset combination, the absolute count and area biases were averaged separately across all valid analysis grids to obtain the corresponding city-level mean errors. Because the absolute values were taken before averaging, local overestimation and underestimation could not cancel each other. The original signed biases were retained for descriptive maps and distributional analyses.
The city-level mean count and area errors were then converted to higher-is-better scores using Equation (2), yielding s N , c d and s A , c d , respectively. For each indicator, the minimum and maximum in Equation (2) were determined separately from the pooled grid-level absolute biases across all valid analysis grids and all city–dataset combinations in the benchmark. The same benchmark-wide bounds were applied to every city–dataset combination, ensuring comparability across cities and datasets within the fixed benchmark. The two component scores were assigned equal weights to obtain the Tier 1 score for city c and dataset d:
S 1 , c d = s N , c d + s A , c d 2
Here, S 1 , c d is the city-level Tier 1 score; averaging it over all valid cities for dataset d yields S 1 , d , the dataset-level Tier 1 score used in Equation (1).

3.2.2. Object Recovery and Spatial Fidelity

This tier evaluates both object recovery over the complete building inventory and spatial fidelity at the full-set and matched-object levels. This metric group is critical for precision-dependent applications such as autonomous navigation, cadastral mapping, and disaster emergency response [51,52,53].
For building-level evaluation, spatial entity matching was performed to establish one-to-one correspondences between individual reference and evaluated footprints. Candidate pairs were first identified by spatial intersection, and pairwise pixel-based IoU was then used to establish one-to-one matches, with 0.30 adopted as the primary threshold. Split and merged configurations were not treated as one-to-many or many-to-one-correspondences. Under the one-to-one matching rule, unmatched evaluated-footprints arising from split cases were counted as false positives, whereas unmatched-reference footprints associated with-merged cases were counted as false negatives, and were therefore penalized through-Object F1. Sensitivity to this choice is examined in Section 4.2.4. To better account for both matched and unmatched objects, object F 1 was calculated as:
F 1 o b j , c d = 2 T P c d 2 T P c d + F P c d + F N c d
where T P c d denotes the number of true positives (accepted one-to-one matches), and F P c d and F N c d denote the numbers of false positives (evaluated footprints without an accepted reference match) and false negatives (reference footprints without an accepted evaluated match), respectively. Object F 1 summarizes object recovery under the specified matching protocol by penalizing both unmatched reference footprints and unmatched evaluated footprints.
Spatial overlap was evaluated at two levels on a common raster grid. City-level IoU measured full-set agreement between the unions of all reference and evaluated footprints, whereas mean object-level IoU summarized pairwise agreement across accepted one-to-one matches. Both were calculated as follows [54]:
I o U c i t y , c d = | C c d G T C c d O S | | C c d G T C c d O S | I o U o b j , c d = 1 T P c d k = 1 T P c d | B k G T B k O S | | B k G T B k O S |
where C c d G T and C c d O S denote the pixel sets collectively occupied by all reference and evaluated footprints, respectively, in city c for dataset d ; B k G T and B k O S denote the pixel sets occupied by the two footprints in the k -th accepted pair; | · | denotes the number of pixels.
Beyond area overlap, positional disagreement was quantified using the RMSE between the centroids of accepted building pairs in city c for dataset d [50]:
R M S E c d = 1 T P c d k = 1 T P c d [ ( X k G T X k O S ) 2 + ( Y k G T Y k O S ) 2 ]
where ( X k G T , Y k G T ) and ( X k O S , Y k O S ) denote the centroid coordinates of the reference and evaluated footprints, respectively, in the k -th accepted pair.
Object F1, city-level IoU, and mean object-level IoU are higher-is-better indicators bounded between 0 and 1 and were used directly in the aggregation. The mean object-level IoU was obtained by averaging Equation (7) over all accepted matches in each city–dataset combination. Centroid RMSE is lower-is-better and was converted to a higher-is-better score s R M S E , c d using Equation (2). Its minimum and maximum were determined from the centroid RMSE values of all valid city–dataset combinations in the benchmark, and the same bounds were applied throughout. The four component scores were assigned equal weights to obtain the Tier 2 score for city c and dataset d :
S 2 , c d = F 1 o b j , c d + I o U c i t y , c d + I o U o b j , c d + s R M S E , c d 4
The score in Equation (9) is defined at the city–dataset level; averaging it over all valid cities for dataset d yields the dataset-level Tier 2 score used in Equation (1).

3.2.3. Morphology Fidelity of Matched Building Pairs

This tier evaluates morphology fidelity among building pairs for which an accepted one-to-one correspondence has been established. These structural attributes are indispensable for emerging domains such as high-fidelity 3D urban reconstruction, digital twin modeling, and microclimate simulations [49,55,56].
In practical building footprint extraction, morphological degradation commonly occurs in several forms, including boundary displacement, angular-structure distortion, compactness inconsistency, and high-frequency contour noise. Figure 4 schematically illustrates these typical degradation patterns. Based on these characteristics, we employ four complementary descriptors to quantify complementary aspects of matched-pair morphology fidelity: PoLiS distance, turning function distance, compactness deviation, and Fourier descriptors [18,38,57,58].
As illustrated in Figure 4, boundary displacement and contour deformation between the two footprints in a matched polygon pair, denoted by A and B , were quantified using the PoLiS distance, a symmetric measure based on the average shortest distances from the vertices of each polygon to the boundary of the other [38]:
P o L i S ( A , B ) = 1 2 q a j A min b B a j b + 1 2 r b k B min a A b k a
where a j and b k are vertices of A and B ; q and r are their corresponding vertex counts; and A and B are their continuous boundaries.
Complementing the boundary distance, the turning function distance quantifies morphological similarity independent of spatial translation or scale [18]. It evaluates the cumulative angle of the tangent ( Θ ) as a function of the normalized arc length ( l ) along the polygon boundary, effectively capturing the preservation of rigid, orthogonal architectural corners:
d T F ( A , B ) = min θ , t 0 1 | Θ A ( λ ) Θ B ( λ + t ) + θ | | d λ
where t denotes a cyclic shift in the boundary starting point, with λ + t interpreted cyclically, and Θ denotes relative rotation.
Additionally, we evaluate the global structural complexity and fragmentation of building footprints using the compactness deviation ( Δ C ) [59]:
Δ C = | 4 π × A r e a O S ( P e r i m e t e r O S ) 2 4 π × A r e a G T ( P e r i m e t e r G T ) 2 |
where A r e a D and P e r i m e t e r D denote the planar area and boundary length of footprint for D { O S , G T } ; Δ C is the absolute difference between their compactness values. This metric determines if the automated extraction overly smoothed or artificially fragmented the true morphology. Finally, Fourier descriptors are utilized to transform the spatial coordinate sequence of a closed boundary into the frequency domain. By comparing the low- and high-frequency coefficients, Fourier descriptors effectively assess global contour similarity and the retention of fine-grained architectural details [57,58].
For each valid city–dataset combination, PoLiS distance, turning function distance, and Fourier-descriptor distance were averaged separately across all accepted matched pairs to obtain city-level mean distances. Each mean distance was converted to a higher-is-better score using Equation (2), with metric-specific benchmark-wide bounds derived from all accepted matched-pair values and applied throughout. Compactness required no benchmark-range rescaling: for each matched pair, its compactness score was calculated directly as 1     Δ C , and these pair-level scores were averaged to obtain the city-level compactness score. The four component scores were assigned equal weights to obtain the Tier 3 score for city c and dataset d :
S 3 , c d = s P o L i S , c d + s T F , c d + s c o m p a c t n e s s , c d + s F D , c d 4
where s P o L i S , c d , s T F , c d , s c o m p a c t n e s s , c d , and s F D , c d denote the direction-aligned city-level scores for PoLiS distance, turning function distance, compactness, and Fourier-descriptor distance, respectively. The score in Equation (13) is defined at the city–dataset level; averaging it over all valid cities for dataset d yields the dataset-level Tier 3 score used in Equation (1). Tier 3 is conditional on the accepted one-to-one matches and does not independently quantify unmatched reference or evaluated objects.

3.3. Data Standardization and Evaluation Implementation

Before metric computation, each evaluated dataset and the corresponding reference footprints were transformed to a common, locally appropriate projected coordinate reference system. This avoided mixing native coordinate systems and ensured that distance and area metrics were calculated in consistent linear units.
The evaluation framework was implemented using Python (v3.9) in the Visual Studio Code (v2026) environment. The calculations were conducted on a workstation running Windows 11 X64, equipped with a 13th Generation Intel Core i5 CPU (2.50 GHz) and 16 GB of RAM. The open-access vector datasets were obtained from official GitHub (accessed on 30 June 2026) repositories and data platforms through programmatic extraction or direct downloads.
Furthermore, we developed an open-source Skill package, available at https://github.com/Tykuinn/building-footprint-eval (accessed on 16 August 2026), to facilitate practical reuse of the benchmark framework. The Skill organizes the detailed accuracy reports and dataset metadata generated in this study into a structured knowledge module. By referencing these embedded evaluation results, it supports interactive dataset recommendation by matching dataset strengths to user-defined selection criteria and application scenarios.

4. Results and Discussion

4.1. Overall Scoring Results

We report coverage-conditional product summaries by averaging each tier score across the cities covered by the corresponding dataset. The overall score is calculated as the equal-weight mean of the three tier scores. Table 4 reports the number of covered cities together with the resulting tier and overall scores. The corresponding city-level component scores and aggregation inputs are provided in Appendix A Table A1, Table A2 and Table A3.
As shown in Table 4, the three tiers exhibited distinct score distributions under the defined scoring configuration. Tier 1 scores ranged from 0.419 to 0.853, Tier 2 scores from 0.343 to 0.587, and Tier 3 scores from 0.495 to 0.803. Tier 2 remained below Tier 3 for every product, whereas Tier 3 exceeded Tier 1 for eight of the eleven products. China-90C, GABLE, and MBF were the exceptions, with Tier 1 scores of 0.609, 0.853, and 0.751 and Tier 3 scores of 0.495, 0.759, and 0.709, respectively. Because the tiers contain different indicators and normalization procedures, their numerical levels should not be interpreted as direct evidence that one quality attribute is inherently stronger than another. Instead, the separation reflects differences in the properties assessed by each tier.
Aggregate agreement was not consistently accompanied by object recovery and spatial fidelity. For example, Mapbox, East Asian buildings, and GABLE produced Tier 1 scores of 0.610, 0.631, and 0.853, respectively, but their Tier 2 scores were 0.423, 0.431, and 0.502. GOB showed a smaller separation, with scores of 0.641 and 0.561 for Tier 1 and Tier 2, respectively. These profiles indicate that agreement in aggregate building count and footprint area can coexist with omissions, commissions, positional displacement, or incomplete object correspondence. Tier 1 indicators alone are therefore insufficient to characterize object-level dataset quality.
Tier 2 and Tier 3 also captured different aspects of performance. Across the products, Tier 3 exceeded Tier 2 by 0.042–0.296. GOB, East Asian buildings, and CMAB, for example, produced Tier 3 scores of 0.803, 0.694, and 0.620, compared with Tier 2 scores of 0.561, 0.431, and 0.343. This separation is consistent with their different evaluation scopes. Tier 2 incorporates object recovery, regional overlap, matched-object overlap, and centroid displacement, whereas Tier 3 describes morphology fidelity only among accepted one-to-one matches. Consequently, well-matched objects may retain their contour characteristics even when full-set object recovery or spatial correspondence remains limited. Tier 3 should therefore be interpreted jointly with Tier 2 rather than as an independent measure of complete-set quality.
Similar overall scores also concealed different tier profiles. GABLE and GOB produced overall scores of 0.705 and 0.668, respectively. GABLE’s summary contained a larger Tier 1 contribution, whereas the GOB summary contained larger Tier 2 and Tier 3 contributions within its two-city subset. A comparable contrast occurred between China-90C and ONEGEO, whose overall scores were 0.519 and 0.497. China-90C had a larger Tier 1 score, while ONEGEO had a larger Tier 3 score. In addition, 3D-GloBFP, East Asian buildings, Mapbox, and Tencent Map fell within a narrow overall-score interval of 0.567–0.585 despite being evaluated over different numbers and combinations of cities. The overall score thus provides a concise synthesis but cannot substitute for inspection of the contributing tier scores.
Figure 5 presents city–dataset heatmaps for the nine evaluation components. In each panel, rows represent the evaluated datasets and columns represent the ten study cities. Arrows adjacent to the color bars indicate the direction of improving performance; for the signed count and area biases, the arrows point toward the optimal value of zero.
As shown in Figure 5, a more detailed view of city-level variation reveals how these performance patterns manifest across different urban contexts. Overall, MBF and GABLE maintain relatively stable performance across multiple cities, with achieving consistently strong overall scores, whereas GOB shows more pronounced advantages in spatial accuracy and morphology-sensitive applications, with particularly high Tier 2 and Tier 3 scores. In contrast, several middle-performing datasets exhibit clear dimension-specific strengths, suggesting that their suitability depends strongly on the target application and local urban morphology. Together, these observations indicate that dataset performance is shaped by the interplay among statistical consistency, spatial alignment, and morphological fidelity, rather than by any single metric alone.

4.2. Tier-Specific Performance Patterns

4.2.1. Statistical Bias Patterns

While the overall scores establish a macro-level baseline, understanding the root causes of these performance variations requires dissecting the fundamental statistical consistency of the extracted features. Figure 6 and Figure 7 present grid-level maps of count bias and area change across representative city–dataset pairs. Grid cells are colored by the count and area bias respectively.
As shown in Figure 6 and Figure 7, statistical inconsistency is not uniformly distributed across the study areas but is spatially clustered within specific urban fabrics. Count bias is more pronounced in dense or morphologically complex grids, where adjacent buildings are more likely to be merged or omitted. In contrast, area change exhibits a partly different spatial pattern, indicating that errors in building quantity and errors in areal representation are not always coupled. Some grids with substantial count underestimation still show moderate area deviation, suggesting that merged polygons may preserve built-up extent while losing object-level separability. These patterns provide direct evidence that statistical consistency is shaped by both dataset generation strategy and local urban morphology.
Figure 8 summarizes the grid-level distributions of building count bias and area change for each evaluated dataset. The box represents the interquartile range, while the central line indicates the average value. It represents the relative count bias per 100 reference buildings and the areal change per 100 m2 of reference building area.
Figure 8 reveals systematic discrepancies in both building count estimation and areal representation across datasets. For building count, most datasets exhibit negative biases, indicating that omission and under-segmentation remain common issues in open-access building footprint products. This pattern is particularly evident in densely built-up areas, where multiple adjacent buildings are frequently merged into a single polygon, thereby reducing the apparent number of detected building instances. In contrast, GABLE and MBF show comparatively smaller count biases, suggesting stronger building-level separability and more stable segmentation behavior in moderately dense urban contexts.
GOB presents a different pattern, with a tendency toward higher building counts in several grid cells. This overestimation is likely associated with finer-grained segmentation of contiguous or semi-connected structures, where built-up blocks represented as single buildings in the reference data may be decomposed into multiple smaller footprint entities. This behavior improves local structural detail in some cases but may also introduce inconsistency when building boundaries are ambiguous.
The area-change boxplots further indicate that low count bias does not necessarily correspond to stable areal representation. MBF and Tencent Map exhibit relatively compact distributions, implying lower grid-level variability and more consistent areal estimation across spatial samples. By contrast, OSM and ONEGEO show broader interquartile ranges and more pronounced dispersion, reflecting spatially uneven data quality. For OSM, this variability is consistent with heterogeneous crowdsourced mapping practices, whereas for ONEGEO it is likely related to multi-source data fusion and uneven sample coverage. Overall, it demonstrates that statistical inconsistency is not only expressed as systematic overestimation or underestimation, but also as substantial spatial dispersion across local grid units.

4.2.2. Object Recovery and Spatial Agreement Patterns

The Tier 2 analysis considers full-set areal agreement, object recovery, and positional fidelity among accepted matches. As shown in Figure 5e, city-level IoU varies substantially across both cities and datasets, indicating that spatial agreement is strongly context dependent. MBF generally achieves higher city-level IoU values across multiple cities, suggesting strong areal agreement with the reference footprints. GOB also shows high overlap where coverage is available, although its limited spatial coverage constrains broader comparison. In contrast, datasets such as CMAB, ONEGEO, and OSM exhibit more variable or lower city-level IoU values, reflecting less stable spatial coverage or stronger regional dependence.
Figure 9 further characterizes object recovery by jointly plotting object recall and object precision. Positions toward the upper-right indicate that a dataset recovers a larger proportion of reference buildings while limiting unmatched predictions. The upper-left region instead indicates omission-dominated behavior, whereas the lower-right region indicates commission-dominated behavior. The light iso-F1 curves provide a supplementary reference for interpreting the balance between the two components.
The object-level precision–recall plot reveals distinct recovery patterns across datasets. MBF showed median recall and precision of approximately 0.51 and 0.61, respectively, with both omission and commission contributing to unmatched objects. GOB showed higher median recall (0.59) but lower precision (0.32), indicating broader object recovery accompanied by more unmatched predictions within its two-city subset. In contrast, OSM had a median precision of 0.62 but a recall of only 0.06, indicating an omission-dominated pattern. CMAB, Mapbox, ONEGEO, and several regional products were likewise concentrated in the low-recall region, although their precision varied.
The areal precision–recall plot provides a complementary complete-set assessment derived from the same intersection, false-positive, and false-negative areas used to calculate City-IoU. MBF had median areal recall and precision of approximately 0.65 and 0.79, while the corresponding values for GOB were 0.72 and 0.69. Several datasets with low object recall nevertheless retained moderate areal recall. For example, Mapbox increased from an object recall of 0.09 to an areal recall of 0.63, while OSM increased from 0.06 to 0.45. This contrast indicates that substantial footprint area may overlap the reference even when individual buildings are not recovered as distinct one-to-one objects. Such patterns are consistent with differences in object partitioning, including merged or split footprints, rather than boundary displacement alone.
The dispersion of city–dataset observations in both plots further shows that recovery performance varies across geographic contexts. The object-level plot identifies whether omission or commission limits instance recovery, whereas the areal plot characterizes the corresponding imbalance in total footprint coverage. These measures should therefore be interpreted jointly with City-IoU and matched-pair geometry metrics. Because the datasets cover different city subsets, the observed positions are coverage-conditional summaries rather than controlled cross-dataset rankings.
Figure 10 presents box plots illustrating the distribution of RMSE and IoU across ten open-access datasets. The box represents the interquartile range, while the central line indicates the mean value, providing a robust summary of central tendency and dispersion for each evaluation metric.
As shown in Figure 10, the object-level IoU distribution centers around approximately 0.5, while the RMSE values are generally concentrated around 5 m. This pattern suggests that, within the benchmark, current open-access building footprint products achieve a moderate level of geometric overlap with the reference data, while still exhibiting non-negligible positional deviations. Such behavior is likely attributable to inherent limitations in large-scale automated extraction pipelines, including mixed-resolution training imagery, inconsistent vectorization standards, and the difficulty of accurately delineating building boundaries in dense or irregular urban environments.
Among all datasets, MBF exhibits a notably higher average object-level IoU and a substantially lower RMSE compared to other datasets, indicating superior geometric alignment and improved positional accuracy relative to the ground truth. This performance advantage is likely driven by higher-quality training data, more robust model generalization, and more refined post-processing strategies that enhance both footprint completeness and boundary precision.
In contrast, the CMAB dataset shows a significantly higher RMSE than the other datasets, suggesting larger positional discrepancies in footprint localization. This may be attributed to variations in data acquisition sources, heterogeneous annotation standards, or less strict geometric correction procedures during post-processing, which can collectively lead to systematic spatial misalignment and increased positional error.
Figure 11 illustrates the spatial comparison outcomes between the evaluated datasets and the ground truth across representative urban regions.
As observed in Figure 11, the CMAB dataset exhibits a visibly higher frequency of commission and omission errors across the four selected cities compared to 3D-GloBFP and East Asian Buildings. This visual evidence intuitively corroborates the significantly higher RMSE observed for CMAB in the prior statistical analysis. Furthermore, the spatial matching maps for 3D-GloBFP and East Asian Buildings are nearly identical in Delingha and Shanghai, and show only marginal discrepancies in Hangzhou and Chongqing. This striking visual consistency unveils a finding not explicitly detailed earlier: it directly explains their highly similar object-level IoU and RMSE distributions presented in Figure 10, indicating comparable spatial extraction logic and boundary alignment capabilities between these two datasets.

4.2.3. Morphology Score Distributions Among Matched Pairs

To characterize shape fidelity across open-access building footprint datasets, we examined the distribution of footprint instances according to their similarity scores relative to the reference data. To place this conditional analysis in the context of overall object recovery, Figure 12 jointly summarizes object-matching outcomes and morphology-score distributions across datasets. Figure 12a presents the relative composition and pooled counts of TP, FP, and FN, whereas Figure 12b reports the proportions of accepted matched pairs with scores below 0.60, between 0.60 and 0.80, and above 0.80. Higher score bands indicate closer agreement with the reference footprints, but these thresholds serve as descriptive cutoffs for cross-dataset comparison rather than absolute quality standards.
Figure 12a shows that FN constitutes the largest outcome category for every dataset except GOB, indicating that omission is the main limitation of object recovery. GOB has the largest relative TP share, but FP accounts for more than half of its pooled outcomes. OSM records the largest absolute TP count (24,678), yet its much larger FN count (176,833) leaves accepted matches as a relatively small proportion.
Among accepted matches, Figure 12b shows that GOB has the largest share of scores above 0.80 (61.9%), followed by GABLE (53.0%), OSM (48.2%), and MBF (46.2%). China-90C presents the opposite pattern, with 48.1% of scores below 0.60 and only 11.5% above 0.80. The remaining datasets occupy intermediate positions, with substantial proportions concentrated between 0.60 and 0.80.
Read together, the two panels show that object recovery and conditional shape fidelity do not necessarily vary in parallel. GOB combines the largest upper-score share with the highest relative TP share, but also contains many unmatched dataset objects. OSM exhibits relatively strong shape fidelity among accepted matches despite being dominated by unmatched reference objects. The morphology-score distribution should therefore be interpreted as conditional on successful matching rather than as a measure of overall dataset completeness.
To illustrate the shape differences represented by these score bands, Figure 13 presents representative matched pairs from the three intervals. The first column identifies the sample locations, while the remaining columns overlay the reference and evaluated footprints, accompanied by their dataset names and normalized scores.
Figure 13 demonstrates that the proposed score categories correspond well to distinct levels of geometric fidelity. High-score footprints closely match the ground truth, exhibiting accurate boundary delineation with minimal omission or redundant details. Medium-score footprints generally preserve the overall building geometry but show moderate deviations, typically reflected by partial loss of fine-scale details or locally redundant structures. In contrast, low-score footprints exhibit pronounced geometric distortions, including oversimplified outlines, irregular polygon boundaries, jagged edges, and excessive redundant structures.
Representative examples illustrate these characteristic error patterns. For instance, the Mapbox dataset in Melbourne, along with the ONEGEO and Tencent Map datasets in Shanghai, exhibits varying degrees of omission of fine-scale structural details. In contrast, the East Asian Buildings and GOB datasets introduce redundant geometric details in Xuzhou and Cayenne, respectively. More severe degradation is observed in low-score footprints, including oversimplified outlines (MBF in Melbourne), excessive redundant structures (OSM in Cayenne), and jagged edges (3D-GloBFP in Xuzhou).
Overall, the decline in footprint score is associated with increasingly severe geometric representation errors. These errors primarily arise from boundary simplification, over-segmentation, and boundary irregularities introduced during polygon generation or post-processing, ultimately reducing the geometric similarity between predicted footprints and the reference annotations.

4.2.4. Sensitivity to the Object-Matching IoU Threshold

As the IoU matching threshold directly determines which object pairs are accepted, its choice may influence both object-recovery performance and morphological metrics derived from matched pairs. We therefore examined the threshold sensitivity of the metrics underlying Tier 2 and Tier 3. Four thresholds, I o U = 0.20 ,   0.30 ,   0.40 ,   a n d   0.50 , were evaluated. For each dataset and threshold, the plotted value represents the median of its city-level metric values across the cities covered by that dataset as Figure 14 illustrated.
For Tier 2, increasing the matching threshold produced the expected trade-off between object recovery and conditional fidelity (Figure 14a–c). Object F1 decreased consistently, whereas mean object-level IoU increased and centroid RMSE declined. The main product contrasts nevertheless remained recognizable: MBF retained comparatively strong object recovery and low RMSE, while OSM and ONEGEO showed relatively high conditional IoU despite lower F1. This contrast confirms that object recovery and matched-pair fidelity capture different aspects of dataset performance.
The Tier 3 metrics exhibited similar threshold-dependent changes (Figure 14d–g). PoLiS distance, compactness deviation, and Fourier-descriptor distance generally decreased as the threshold increased because stricter matching retained more closely overlapping pairs. Turning-function distance showed no uniform trend, indicating that greater overlap does not necessarily correspond to stronger angular-structure agreement. Persistent product characteristics also remained visible: MBF, GABLE, and GOB generally had smaller PoLiS distances, whereas China-90C retained a distinct morphology profile, particularly in turning-function distance and compactness deviation.
Despite these changes in metric magnitude, the broad product-level patterns remained stable. Relative to the primary threshold of 0.30, seven metrics produced Spearman correlations of 0.827–1.000 across the alternative settings; only Fourier-descriptor distance declined to 0.773 at the threshold of 0.50. Stricter matching narrowed some between-product differences but did not materially alter the comparative conclusions. The evaluation is therefore reasonably robust to the matching threshold within the tested range, while remaining conditional on each product’s city coverage.

4.3. Dataset Adoption Suggestions

In practical applications, dataset selection is constrained by both data quality and geographic availability. For studies requiring broad and consistent spatial coverage, particularly those spanning multiple countries or regions, OSM remains a practical baseline because of its extensive availability. Although OSM does not consistently outperform other products and shows relatively weak statistical consistency, it provides moderate object recovery and spatial fidelity together with comparatively strong morphology fidelity, and can therefore serve as a useful general-purpose or fallback dataset where higher-performing regional products are unavailable. Nevertheless, its quality varies substantially among cities, reflecting the heterogeneous nature of crowdsourced mapping. Therefore, when applications require higher reliability or finer spatial detail, OSM and other open-access products should be supplemented by local validation, manual correction, or authoritative mapping data where available.
For applications primarily concerned with statistical consistency, such as large-scale building-stock characterization, population estimation, or comparative urban analysis, existing open-access datasets can provide a useful approximation of broad spatial patterns. However, the systematic underestimation of building count and area observed in this study indicates that these products should not be interpreted as direct substitutes for authoritative statistics. Their reliability also decreases as the analytical scale becomes finer. At regional or city scales, substantial spatial heterogeneity may emerge, particularly in dense and morphologically complex urban areas where omission and building aggregation are more frequent. Within the evaluated coverage, GABLE exhibits the strongest statistical consistency, with a Tier 1 score of 0.853, while MBF also performs comparatively well, with a Tier 1 score of 0.751. Accordingly, GABLE may be particularly valuable for regional studies in China where its coverage is available. In contrast, the relatively low Tier 1 score of OSM (0.438), together with its pronounced city-to-city variability, suggests that it should be adopted more cautiously for localized statistical analysis and preferably be subjected to prior quality assessment.
For applications emphasizing spatial accuracy, greater caution is required. Across the evaluated datasets, object-level overlap remains moderate and positional discrepancies are generally on the order of several meters, indicating that none of the examined open-access products consistently reaches the accuracy required by precision-sensitive applications. Tasks such as cadastral mapping, high-precision change detection, or other applications requiring accurate building localization should therefore not rely directly on these datasets as final mapping products without additional correction or verification. Among the evaluated products, MBF exhibits the strongest overall object recovery and spatial fidelity, achieving the highest Tier 2 score of 0.587, and should therefore be prioritized when positional agreement is an important consideration. GOB also performs comparatively well in this dimension, with a Tier 2 score of 0.561. For less accuracy-sensitive applications, MBF can provide an effective default option, whereas OSM may still serve as a supplementary alternative where more accurate products are unavailable.
For morphology-sensitive applications, dataset preference differs substantially. GOB demonstrates the strongest morphology fidelity among the evaluated products, with the highest Tier 3 score of 0.803, and should be prioritized where its spatial coverage is available, particularly for applications involving building-shape analysis, three-dimensional reconstruction, digital-twin modeling, or other tasks sensitive to boundary fidelity. Where GOB is unavailable, regional products can provide effective alternatives. GABLE, for example, also exhibits strong morphology fidelity, with a Tier 3 score of 0.759, while East Asian Buildings achieves a score of 0.694 across the evaluated East Asian cities. OSM also provides comparatively strong morphology fidelity, with a Tier 3 score of 0.715, and may therefore remain useful in regions with limited alternatives; however, its performance exhibits clear geographic variability, and substantial positional offsets, boundary simplification, or high-frequency contour errors may occur in individual cities. Consequently, local inspection remains advisable before OSM is used for morphology-dependent analysis.
Overall, the results indicate that no single dataset should be treated as universally optimal. Dataset adoption should instead follow a task-oriented strategy in which geographic availability is considered first, followed by the quality dimension most relevant to the intended application. Broad-scale statistical studies can tolerate moderate geometric inaccuracies but should account for systematic bias, with GABLE showing the strongest statistical consistency within its coverage; spatially precise applications require local verification or correction even when relatively strong products such as MBF are used; and morphology-sensitive applications should preferentially adopt GOB or competitive regional products such as GABLE where available.

4.4. Scope and Limitations

Despite the use of high-quality reference data and a consistent multi-dimensional evaluation framework, several limitations should be acknowledged in this study. The 10 cities were purposively selected to represent diverse geographic and urban conditions rather than to provide a statistically representative global sample. Rapidly changing urban fringes were also excluded to minimize temporal mismatch. In addition, dataset coverage varied substantially among products, meaning that aggregate scores were derived from different subsets of cities. Therefore, the reported rankings and recommendations should be interpreted within the evaluated city–dataset pairs and locally validated before being transferred to unsampled regions.
The reference footprints were manually delineated from high-resolution imagery rather than derived from cadastral surveys. Although standardized annotation rules and multi-stage quality control reduced interpretation inconsistencies, residual uncertainty may remain for occluded, adjoining, or roof-displaced buildings. The operational minimum mapping unit was also determined by visual interpretability rather than a fixed area threshold. These factors should be considered when interpreting the benchmark results and applying them to more demanding mapping scenarios.

5. Conclusions

This study established an expert-delineated reference benchmark comprising more than 200,000 building footprints and used it to evaluate 11 open-access building footprint datasets across 10 cities. The proposed three-tier framework distinguishes statistical consistency, object recovery and spatial fidelity, and morphology fidelity among accepted building pairs. Across the evaluated city–dataset pairs, building count and area showed average biases of −34.0% and −14.9%, respectively, while city-level IoU ranged from 29.1% to 58.3% and positional errors remained on the order of several meters. Aggregate agreement did not necessarily correspond to successful object recovery, and strong morphology fidelity among accepted matches did not imply complete-set quality. Sensitivity analysis further showed that the broad product-level patterns were stable across the tested IoU matching thresholds.
Under the coverage-conditional aggregation, GABLE achieved the highest overall and Tier 1 scores, MBF achieved the highest Tier 2 score, and GOB achieved the highest Tier 3 score. These dimension-specific results confirm that no product is universally optimal and that an overall score cannot substitute for inspection of the individual tiers. Regional products remain competitive within their coverage areas, while OSM provides a broadly available fallback but requires local validation because of its geographic variability. Dataset selection should therefore consider geographic availability and application-specific requirements. Despite limitations related to city sampling, unequal product coverage, and reference-data uncertainty, the framework and accompanying open-source Skill package provide a reproducible basis for task-oriented evaluation and future benchmark expansion.

Author Contributions

Conceptualization, Y.L. and T.Z.; methodology, T.Z. and W.G.; software, T.Z.; validation, Z.C., C.L. and Q.X.; formal analysis, W.G. and Q.X.; investigation, Z.C. and C.L.; resources, G.Z., W.G. and Q.C.; data curation, C.L., G.Z. and Q.X.; writing—original draft preparation, T.Z. and Y.L.; writing—review and editing, P.T., Q.C. and Y.L.; visualization, Z.C. and G.Z.; supervision, Y.L. and P.T.; project administration, Q.C.; funding acquisition, Y.L., P.T. and Q.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by Open Project Funds for the Joint Laboratory of Spatial intelligent Perception and Large Model Application, (Grant No. SIPLMA-2025-YB-10) and Key Laboratory of Smart Earth (No. KF2023ZD02-01).

Data Availability Statement

The datasets presented in this article are not readily available because the data are part of an ongoing study. Further inquiries regarding the datasets can be directed to the corresponding author.

Acknowledgments

The authors gratefully acknowledge the providers and maintainers of the open-access building footprint datasets used in this study, including Microsoft Global ML Building Footprints, Google Research Open Buildings, OpenStreetMap, Mapbox, ONEGEO, Tencent Map, GABLE, East Asian Buildings, 3D-GloBFP, China 90Cities Rooftop Data, and CMAB. The China 90Cities Rooftop Data dataset is provided by the National Tibetan Plateau/Third Pole Environment Data Center (http://data.tpdc.ac.cn). OpenStreetMap data are © OpenStreetMap contributors and are made available under the Open Database License. Where Mapbox map data is used, attribution is provided as © Mapbox and © OpenStreetMap. The authors also acknowledge ONEGEO GmbH and Tencent Location Service for their respective data products.

Conflicts of Interest

Author Qian Xue was employed by the company Huantian Wisdom Technology Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Appendix A

Table A1, Table A2 and Table A3 report the component scores used to calculate the three tier scores for each valid city–dataset combination. All component scores are direction-aligned to the interval [ 0 , 1 ] , with larger values indicating closer agreement with the reference data. Dataset-level tier scores in Table 4 were calculated by averaging the corresponding city-level tier scores over the valid cities within each dataset’s evaluated coverage. The calculations retained full numerical precision, whereas displayed values were rounded only for presentation.
Table A1. Component scores for Tier 1: Statistical Consistency.
Table A1. Component scores for Tier 1: Statistical Consistency.
DatasetCount-Error ScoreArea-Error ScoreTier 1: Statistical Consistency
3D-GloBFP0.4770.7130.595
CMAB0.3250.7210.523
China-90C0.4210.7960.609
East Asian Buildings0.5310.7310.631
GABLE0.7210.9860.853
GOB0.3300.9520.641
Mapbox0.3650.8550.610
MBF0.6660.8360.751
ONEGEO0.2930.5440.419
OSM0.3710.5050.438
Tencent Map0.3390.8070.573
* The count- and area-error scores were derived from the two-sided magnitudes of the corresponding signed biases. The Tier 1 score is the equal-weight arithmetic mean of the two component scores. Calculations used full-precision values; the values shown here were rounded to six decimal places.
Table A2. Component scores for Tier 2: Object Recovery and Spatial Fidelity.
Table A2. Component scores for Tier 2: Object Recovery and Spatial Fidelity.
DatasetObject F1City-Level IoU ScoreObject-Level IoU ScoreCentroid-RMSE ScoreTier 2: Object Recovery and Spatial Fidelity
3D-GloBFP0.1970.6330.4050.5050.435
CMAB0.1380.4380.3250.4700.343
China-90c0.2320.5930.4650.5230.453
East Asian Buildings0.2050.6430.3780.5000.431
GABLE0.2610.7370.4540.5570.502
GOB0.4100.7910.5470.4960.561
Mapbox0.1380.6200.4800.4550.423
MBF0.4390.7560.5830.5710.587
ONEGEO0.1360.5940.2910.5300.388
OSM0.2380.6190.3710.5500.444
Tencent Map0.2140.6840.4550.5190.468
* Object F1 was calculated at an IoU matching threshold of 0.30. The Tier 2 score is the equal-weight arithmetic mean of object F1, the centroid-RMSE score, the object-level IoU score, and the city-level IoU score. None of the four components was multiplied by object F1.
Table A3. Component scores for Tier 3: Matched-Pair Morphology Fidelity.
Table A3. Component scores for Tier 3: Matched-Pair Morphology Fidelity.
DatasetPoLiS ScoreTurning-Function ScoreFourier-Descriptor ScoreCompactness ScoreTier 3: Matched-Pair Morphology Fidelity
3D-GloBFP0.5950.9430.3410.9060.696
CMAB0.3880.9520.2610.8780.620
China-90C0.5930.3480.2320.8090.495
East Asian Buildings0.6110.9730.2880.9020.694
GABLE0.7200.9610.4590.8940.759
GOB0.7110.9920.5610.9480.803
Mapbox0.5130.9580.3070.8890.667
MBF0.6710.7880.4670.9110.709
ONEGEO0.5170.9400.3860.8910.684
OSM0.5500.8760.5280.9070.715
Tencent Map0.5640.9320.4200.8740.698
* All four morphology descriptors were calculated only for accepted one-to-one matches and normalized as lower-is-better indicators. The Tier 3 score is their equal-weight arithmetic. It therefore characterizes morphology fidelity conditional on successful matching and should be interpreted together with the object-recovery statistics reported in Table A2.

References

  1. Sun, Y.; Meng, L.; Camero, A.; Auer, S.; Zhu, X.X. A Deep Dive into OpenStreetMap Research since Its Inception (2008–2024): Contributors, Topics, and Future Trends. Int. J. Geogr. Inf. Sci. 2026, 40, 3013–3067. [Google Scholar] [CrossRef] [Scilit]
  2. Huang, X.; Wang, S.; Lu, T.; Liu, Y.; Serrano-Estrada, L. Crowdsourced Geospatial Data Is Reshaping Urban Sciences. Int. J. Appl. Earth Obs. Geoinf. 2024, 127, 103687. [Google Scholar] [CrossRef] [Scilit]
  3. Biljecki, F.; Chew, L.Z.X.; Milojevic-Dupont, N.; Creutzig, F. Open Government Geospatial Data on Buildings for Planning Sustainable and Resilient Cities. arXiv 2021, arXiv:2107.04023. [Google Scholar] [CrossRef] [Scilit]
  4. Liu, S.; Higgs, C.; Arundel, J.; Boeing, G.; Cerdera, N.; Moctezuma, D.; Cerin, E.; Adlakha, D.; Lowe, M.; Giles-Corti, B. A Generalized Framework for Measuring Pedestrian Accessibility around the World Using Open Data. Geogr. Anal. 2022, 54, 559–582. [Google Scholar] [CrossRef] [Scilit]
  5. Kucukali, A.; Pjeternikaj, R.; Zeka, E.; Hysa, A. Evaluating the Pedestrian Accessibility to Public Services Using Open-Source Geospatial Data and QGIS Software. Nova Geod. 2022, 2, 42. [Google Scholar] [CrossRef] [Scilit]
  6. Boeing, G. Exploring Urban Form Through Openstreetmap Data: A Visual Introduction. arXiv 2020, arXiv:2008.12142. [Google Scholar] [CrossRef] [Scilit]
  7. Chen, Q.; Li, X.; Zhang, Z.; Zhou, C.; Guo, Z.; Liu, Z.; Zhang, H. Remote Sensing of Photovoltaic Scenarios: Techniques, Applications and Future Directions. Appl. Energy 2023, 333, 120579. [Google Scholar] [CrossRef] [Scilit]
  8. Herfort, B.; Lautenbach, S.; Porto de Albuquerque, J.; Anderson, J.; Zipf, A. A Spatio-Temporal Analysis Investigating Completeness and Inequalities of Global Urban Building Data in OpenStreetMap. Nat. Commun. 2023, 14, 3985. [Google Scholar] [CrossRef] [Scilit]
  9. Cao, K.; Zhou, C.; Church, R.; Li, X.; Li, W. Revisiting Spatial Optimization in the Era of Geospatial Big Data and GeoAI. Int. J. Appl. Earth Obs. Geoinf. 2024, 129, 103832. [Google Scholar] [CrossRef] [Scilit]
  10. Goodchild, M.F.; Li, L. Assuring the Quality of Volunteered Geographic Information. Spat. Stat. 2012, 1, 110–120. [Google Scholar] [CrossRef] [Scilit]
  11. Biljecki, F.; Chow, Y.S. Global Building Morphology Indicators. Comput. Environ. Urban Syst. 2022, 95, 101809. [Google Scholar] [CrossRef] [Scilit]
  12. Mohammed, S.; Ehrlinger, L.; Harmouch, H.; Naumann, F.; Srivastava, D. Data Quality Assessment: Challenges and Opportunities. arXiv 2024, arXiv:2403.00526. [Google Scholar] [CrossRef] [Scilit]
  13. Chamberlain, H.R.; Pollard, D.; Winters, A.; Renn, S.; Borkovska, O.; Musuka, C.A.; Membele, G.; Lazar, A.N.; Tatem, A.J. Assessing the Impact of Building Footprint Dataset Choice for Health Programme Planning: A Case Study of Indoor Residual Spraying (IRS) in Zambia. Int. J. Health Geogr. 2025, 24, 13. [Google Scholar] [CrossRef] [Scilit]
  14. Wang, R.; Huang, S.; Yang, H. Building3D: An Urban-Scale Dataset and Benchmarks for Learning Roof Structures from Point Clouds. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2023), Paris, France, 1–6 October 2023; pp. 20076–20086. [Google Scholar]
  15. Yang, G.; Xue, F.; Zhang, Q.; Xie, K.; Fu, C.-W.; Huang, H. UrbanBIS: A Large-Scale Benchmark for Fine-Grained Urban Building Instance Segmentation. In ACM SIGGRAPH 2023 Conference Proceedings; Association for Computing Machinery: New York, NY, USA, 2023; pp. 1–11. [Google Scholar] [CrossRef] [Scilit]
  16. Egger, P.; Rao, S.X.; Papini, S. Building Floorspace in China: A Dataset and Learning Pipeline. arXiv 2023, arXiv:2303.02230. [Google Scholar] [CrossRef] [Scilit]
  17. Li, R.; Sun, T.; Ghaffarian, S.; Tsamados, M.; Ni, G. GLAMOUR: GLobAl Building MOrphology Dataset for URban Hydroclimate Modelling. Sci. Data 2024, 11, 618. [Google Scholar] [CrossRef] [Scilit]
  18. Moradi, M.; Roche, S.; Mostafavi, M.A. A Novel GIS-Based Polygon Shape Similarity Measure Applied to OSM Building Footprints. Adv. Cartogr. GISci. ICA 2025, 5, 22. [Google Scholar] [CrossRef] [Scilit]
  19. Dabove, P.; Daud, M.; Olivotto, L. Revolutionizing Urban Mapping: Deep Learning and Data Fusion Strategies for Accurate Building Footprint Segmentation. Sci. Rep. 2024, 14, 13510. [Google Scholar] [CrossRef] [Scilit]
  20. Microsoft. Microsoft Global ML Building Footprints. Data Repository. 2026. Available online: https://github.com/microsoft/GlobalMLBuildingFootprints (accessed on 30 June 2026).
  21. Sirko, W.; Kashubin, S.; Ritter, M.; Annkah, A.; Bouchareb, Y.S.E.; Dauphin, Y.; Keysers, D.; Neumann, M.; Cisse, M.; Quinn, J. Continental-Scale Building Detection from High Resolution Satellite Imagery. arXiv 2021, arXiv:2107.12283. [Google Scholar] [CrossRef] [Scilit]
  22. OpenStreetMap contributors. OpenStreetMap. OpenStreetMap Foundation. 2026. Available online: https://www.openstreetmap.org/ (accessed on 30 June 2026).
  23. Mapbox. Mapbox Streets v8: Tileset Reference Documentation. Online Documentation. 2026. Available online: https://docs.mapbox.com/data/tilesets/reference/mapbox-streets-v8/ (accessed on 30 June 2026).
  24. ONEGEO. Data Format Documentation 2026. Online Documentation, Version 2026-01. 2026. Available online: https://onegeo.co/documentation/data/ (accessed on 30 June 2026).
  25. Zhang, Y.; Zhao, H.; Long, Y. CMAB: A Multi-Attribute Building Dataset of China. Sci. Data 2025, 12, 430. [Google Scholar] [CrossRef] [Scilit]
  26. Sun, X.; Huang, X.; Mao, Y.; Sheng, T.; Li, J.; Wang, Z.; Lu, X.; Ma, X.; Tang, D.; Chen, K. GABLE: A First Fine-Grained 3D Building Model of China on a National Scale from Very High Resolution Satellite Imagery. Remote Sens. Environ. 2024, 305, 114057. [Google Scholar] [CrossRef] [Scilit]
  27. Che, Y.; Li, X.; Liu, X.; Wang, Y.; Liao, W.; Zheng, X.; Zhang, X.; Xu, X.; Shi, Q.; Zhu, J.; et al. 3D-GloBFP: The First Global Three-Dimensional Building Footprint Dataset. Earth Syst. Sci. Data 2024, 16, 5357–5374. [Google Scholar] [CrossRef] [Scilit]
  28. Shi, Q.; Zhu, J.; Liu, Z.; Guo, H.; Gao, S.; Liu, M.; Liu, Z.; Liu, X. The Last Puzzle of Global Building Footprints—Mapping 280 Million Buildings in East Asia Based on VHR Images. J. Remote Sens. 2024, 4, 0138. [Google Scholar] [CrossRef] [Scilit]
  29. Zhang, Z.; Qian, Z.; Zhong, T.; Chen, M.; Zhang, K.; Yang, Y.; Zhu, R.; Zhang, F.; Zhang, H.; Zhou, F.; et al. Vectorized Rooftop Area Data for 90 Cities in China. Sci. Data 2022, 9, 66. [Google Scholar] [CrossRef] [Scilit]
  30. Tencent Location Service. JavaScript API GL Documentation. Online Documentation. 2026. Available online: https://lbs.qq.com/webApi/javascriptGL/glDoc/docIndexMap (accessed on 30 June 2026).
  31. Chamberlain, H.R.; Darin, E.; Adewole, W.A.; Jochem, W.C.; Lazar, A.N.; Tatem, A.J. Building Footprint Data for Countries in Africa: To What Extent Are Existing Data Products Comparable? Comput. Environ. Urban Syst. 2024, 110, 102104. [Google Scholar] [CrossRef] [Scilit]
  32. Minghini, M.; Thabit Gonzalez, S.; Gabrielli, L. Pan-European Open Building Footprints: Analysis and Comparison in Selected Countries. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2024, XLVIII-4-W12-2024, 97–103. [Google Scholar] [CrossRef] [Scilit]
  33. Brovelli, M.A.; Zamboni, G. A New Method for the Assessment of Spatial Accuracy and Completeness of OpenStreetMap Building Footprints. ISPRS Int. J. Geo-Inf. 2018, 7, 289. [Google Scholar] [CrossRef] [Scilit]
  34. Nyandwi, E.; Gerke, M.; Achanccaray, P. Local Evaluation of Large-Scale Remote Sensing Machine Learning-Generated Building and Road Dataset: The Case of Rwanda. PFG 2024, 92, 705–722. [Google Scholar] [CrossRef] [Scilit]
  35. Okyere, F.; Lu, M.; Brunn, A. Evaluating the Quality of Open Building Datasets for Mapping Urban Inequality: A Comparative Analysis Across 5 Cities. arXiv 2025, arXiv:2508.12872. [Google Scholar] [CrossRef] [Scilit]
  36. Fan, H.; Zhao, Z.; Li, W. Towards Measuring Shape Similarity of Polygons Based on Multiscale Features and Grid Context Descriptors. ISPRS Int. J. Geo-Inf. 2021, 10, 279. [Google Scholar] [CrossRef] [Scilit]
  37. Lin, T.-Y.; Maire, M.; Belongie, S.; Bourdev, L.; Girshick, R.; Hays, J.; Perona, P.; Ramanan, D.; Zitnick, C.L.; Dollár, P. Microsoft COCO: Common Objects in Context. In Computer Vision—ECCV 2014; Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T., Eds.; Springer: Cham, Switzerland, 2014; Volume 8693, pp. 740–755. [Google Scholar] [CrossRef] [Scilit]
  38. Avbelj, J.; Muller, R.; Bamler, R. A Metric for Polygon Comparison and Building Extraction Evaluation. IEEE Geosci. Remote Sens. Lett. 2015, 12, 170–174. [Google Scholar] [CrossRef] [Scilit]
  39. Zhu, X.X.; Chen, S.; Zhang, F.; Shi, Y.; Wang, Y. GlobalBuildingAtlas: An Open Global and Complete Dataset of Building Polygons, Heights and LoD1 3D Models. Earth Syst. Sci. Data 2025, 17, 6647–6668. [Google Scholar] [CrossRef] [Scilit]
  40. Hecht, R.; Kunze, C.; Hahmann, S. Measuring Completeness of Building Footprints in OpenStreetMap over Space and Time. ISPRS Int. J. Geo-Inf. 2013, 2, 1066–1091. [Google Scholar] [CrossRef] [Scilit]
  41. Gevaert, C.; Buunk, T.; Homberg, M.J.C.V.D. Auditing Geospatial Datasets for Biases: Using Global Building Datasets for Disaster Risk Management. TechRxiv 2024. [Google Scholar] [CrossRef] [Scilit]
  42. Zhou, Q.; Zhang, Y.; Chang, K.; Brovelli, M.A. Assessing OSM Building Completeness for Almost 13,000 Cities Globally. Int. J. Digit. Earth 2022, 15, 2400–2421. [Google Scholar] [CrossRef] [Scilit]
  43. Chen, S.; Ogawa, Y.; Zhao, C.; Sekimoto, Y. Enhanced Large-Scale Building Extraction Evaluation: Developing a Two-Level Framework Using Proxy Data and Building Matching. Eur. J. Remote Sens. 2024, 57, 2374844. [Google Scholar] [CrossRef] [Scilit]
  44. Huttenlocher, D.P.; Klanderman, G.A.; Rucklidge, W.J. Comparing Images Using the Hausdorff Distance. IEEE Trans. Pattern Anal. Mach. Intell. 1993, 15, 850–863. [Google Scholar] [CrossRef] [Scilit]
  45. Zhang, Y.; Tu, Z.; Lian, W.; Hu, Y.; Sun, S.; Xiao, Y.; Cheng, Y. BANet: Bidirectional Feature Aggregation and Adaptive Multi-Scene Perception-Based Lane Detection for Autonomous Driving. IEEE Trans. Intell. Transp. Syst. 2026, 27, 9713–9723. [Google Scholar] [CrossRef] [Scilit]
  46. Ai, T.; Cheng, X.; Liu, P.; Yang, M. A Shape Analysis and Template Matching of Building Features by the Fourier Transform Method. Comput. Environ. Urban Syst. 2013, 41, 219–233. [Google Scholar] [CrossRef] [Scilit]
  47. Basaraner, M.; Cetinkaya, S. Performance of Shape Indices and Classification Schemes for Characterising Perceptual Shape Complexity of Building Footprints in GIS. Int. J. Geogr. Inf. Sci. 2017, 31, 1952–1977. [Google Scholar] [CrossRef] [Scilit]
  48. Boo, G.; Darin, E.; Leasure, D.R.; Dooley, C.A.; Chamberlain, H.R.; Lázár, A.N.; Tschirhart, K.; Sinai, C.; Hoff, N.A.; Fuller, T.; et al. High-Resolution Population Estimation Using Household Survey Data and Building Footprints. Nat. Commun. 2022, 13, 1330. [Google Scholar] [CrossRef] [Scilit]
  49. F. de Arruda, H.; Reia, S.M.; Ruan, S.; Atwal, K.S.; Kavak, H.; Anderson, T.; Pfoser, D. An OpenStreetMap Derived Building Classification Dataset for the United States. Sci. Data 2024, 11, 1210. [Google Scholar] [CrossRef] [Scilit]
  50. Zeng, C.; Wang, J.; Lehrbass, B. An Evaluation System for Building Footprint Extraction From Remotely Sensed Data. Sel. Top. Appl. Earth Obs. Remote Sens. IEEE J. 2013, 6, 1640–1652. [Google Scholar] [CrossRef] [Scilit]
  51. Hansch, R.; Arndt, J.; Lunga, D.; Gibb, M.; Pedelose, T.; Boedihardjo, A.; Petrie, D.; Bacastow, T.M. SpaceNet 8—The Detection of Flooded Roads and Buildings. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), New Orleans, LA, USA, 19–20 June 2022; pp. 1471–1479. [Google Scholar]
  52. Florio, P.; Giovando, C.; Goch, K.; Pesaresi, M.; Politis, P.; Martinez, A. Towards A Pan-Eu Building Footprint Map Based On The Hierarchical Conflation Of Open Datasets: The Digital Building Stock Model—DBSM. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2023, XLVIII-4-W7-2023, 47–52. [Google Scholar] [CrossRef] [Scilit]
  53. Yang, H.L.; Yuan, J.; Lunga, D.; Laverdiere, M.; Rose, A.; Bhaduri, B. Building Extraction at Scale Using Convolutional Neural Network: Mapping of the United States. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2018, 11, 2600–2614. [Google Scholar] [CrossRef] [Scilit]
  54. Naumann, A.; Bonerath, A.; Haunert, J.-H. Scalable Many-to-Many Building Footprint Matching. Inf. Fusion 2025, 124, 103360. [Google Scholar] [CrossRef] [Scilit]
  55. Lopez-Cabeza, V.P.; Videras-Rodriguez, M.; Gomez-Melgar, S. An Open-Source Urban Digital Twin for Enhancing Outdoor Thermal Comfort in the City of Huelva (Spain). Smart Cities 2025, 8, 160. [Google Scholar] [CrossRef] [Scilit]
  56. Kamath, H.G.; Singh, M.; Malviya, N.; Martilli, A.; He, L.; Aliaga, D.; He, C.; Chen, F.; Magruder, L.A.; Yang, Z.-L.; et al. GLObal Building Heights for Urban Studies (UT-GLOBUS) for City- and Street-Scale Urban Simulations: Development and First Applications. Sci. Data 2024, 11, 886. [Google Scholar] [CrossRef] [Scilit]
  57. Reinoso-Gordo, J.F.; Romero-Zaliz, R.; León-Robles, C.; Mataix-SanJuan, J.; Nero, M.A. Fourier-Based Automatic Transformation between Mapping Shapes—Cadastral and Land Registry Applications. ISPRS Int. J. Geo-Inf. 2020, 9, 482. [Google Scholar] [CrossRef] [Scilit]
  58. Persoon, E.; Fu, K.-S. Shape Discrimination Using Fourier Descriptors. IEEE Trans. Syst. Man. Cybern. 1977, 7, 170–179. [Google Scholar] [CrossRef] [Scilit]
  59. Ďuračiová, R. An Aggregated Shape Similarity Index: A Case Study of Comparing the Footprints of OpenStreetMap and INSPIRE Buildings. ISPRS Int. J. Geo-Inf. 2023, 12, 495. [Google Scholar] [CrossRef] [Scilit]
Figure 1. The pipeline of the manual annotation workflow.
Figure 1. The pipeline of the manual annotation workflow.
Remotesensing 18 02895 g001
Figure 2. Distribution of the study areas and reference annotations. (a) Locations of the 10 selected cities; (b) study-area imagery, with blue dots marking the examples shown in (c); and (c) manually delineated building-footprint annotations shown as red vector outlines.
Figure 2. Distribution of the study areas and reference annotations. (a) Locations of the 10 selected cities; (b) study-area imagery, with blue dots marking the examples shown in (c); and (c) manually delineated building-footprint annotations shown as red vector outlines.
Remotesensing 18 02895 g002aRemotesensing 18 02895 g002b
Figure 3. Coverage matrix of evaluated datasets across the 10 reference cities.
Figure 3. Coverage matrix of evaluated datasets across the 10 reference cities.
Remotesensing 18 02895 g003
Figure 4. Typical morphology degradation patterns in building footprints, (a) boundary displacement measured by PoLiS distance; (b) angular-structure change measured by the turning function; (c) global shape-complexity change measured by compactness; and (d) high-frequency contour noise represented by Fourier descriptors.
Figure 4. Typical morphology degradation patterns in building footprints, (a) boundary displacement measured by PoLiS distance; (b) angular-structure change measured by the turning function; (c) global shape-complexity change measured by compactness; and (d) high-frequency contour noise represented by Fourier descriptors.
Remotesensing 18 02895 g004
Figure 5. City–dataset heatmaps of the ten evaluation components across the three evaluation tiers: Tier 1, (a) average count bias and (b) average area bias; Tier 2, (c) average object-level IoU, (d) RMSE, (e) city-level IoU, and (f) Object F1; and Tier 3, (g) PoLiS distance, (h) Turning function distance, (i) Compactness, and (j) Fourier-descriptor distance. Panels (c,d,gj) are computed for matched pairs. All panels display the original metric values without further score adjustment or transformation. Arrows indicate the direction of better performance for each metric.
Figure 5. City–dataset heatmaps of the ten evaluation components across the three evaluation tiers: Tier 1, (a) average count bias and (b) average area bias; Tier 2, (c) average object-level IoU, (d) RMSE, (e) city-level IoU, and (f) Object F1; and Tier 3, (g) PoLiS distance, (h) Turning function distance, (i) Compactness, and (j) Fourier-descriptor distance. Panels (c,d,gj) are computed for matched pairs. All panels display the original metric values without further score adjustment or transformation. Arrows indicate the direction of better performance for each metric.
Remotesensing 18 02895 g005aRemotesensing 18 02895 g005b
Figure 6. Grid-level spatial distribution of building count bias across datasets.
Figure 6. Grid-level spatial distribution of building count bias across datasets.
Remotesensing 18 02895 g006
Figure 7. Grid-level spatial distribution of building area bias across datasets.
Figure 7. Grid-level spatial distribution of building area bias across datasets.
Remotesensing 18 02895 g007
Figure 8. Grid-level distributions of relative count and area biases across datasets.
Figure 8. Grid-level distributions of relative count and area biases across datasets.
Remotesensing 18 02895 g008
Figure 9. Object recovery and city-scale areal agreement across the observed city–dataset pairs. Each small marker represents one city–dataset pair in precision–recall space. Larger labeled markers indicate the product-specific median precision and recall, and gray contours denote constant object-F1 values. Product summaries apply only to the cities covered by each dataset.
Figure 9. Object recovery and city-scale areal agreement across the observed city–dataset pairs. Each small marker represents one city–dataset pair in precision–recall space. Larger labeled markers indicate the product-specific median precision and recall, and gray contours denote constant object-F1 values. Product summaries apply only to the cities covered by each dataset.
Remotesensing 18 02895 g009
Figure 10. Distributions of matched-object spatial fidelity metrics.
Figure 10. Distributions of matched-object spatial fidelity metrics.
Remotesensing 18 02895 g010
Figure 11. Object-based evaluation and error visualization of the extracted building footprints.
Figure 11. Object-based evaluation and error visualization of the extracted building footprints.
Remotesensing 18 02895 g011
Figure 12. Object-recovery outcomes and conditional morphology fidelity across datasets. (a) Relative composition and pooled counts of accepted one-to-one matches (TP), dataset objects without an accepted reference match (FP), and reference objects without an accepted dataset match (FN). Bar widths are normalized by (TP+FP+FN), whereas labels report the corresponding object counts. (b) Distribution of morphology-score bands among accepted matched pairs.
Figure 12. Object-recovery outcomes and conditional morphology fidelity across datasets. (a) Relative composition and pooled counts of accepted one-to-one matches (TP), dataset objects without an accepted reference match (FP), and reference objects without an accepted dataset match (FN). Bar widths are normalized by (TP+FP+FN), whereas labels report the corresponding object counts. (b) Distribution of morphology-score bands among accepted matched pairs.
Remotesensing 18 02895 g012
Figure 13. Qualitative comparison of predicted structures across cities.
Figure 13. Qualitative comparison of predicted structures across cities.
Remotesensing 18 02895 g013
Figure 14. Sensitivity of Tier 2 and Tier 3 metrics to the IoU matching threshold. Panels (ac) show the Tier 2 metrics: (a) Object F1, (b) Mean average object-level IoU, and (c) RMSE. Panels (dg) show the Tier 3 metrics: (d) PoLiS distance, (e) Turning function distance, (f) Compactness, and (g) Fourier-descriptor distance. Metrics in panels (bg) are computed for matched pairs. The four lines represent IoU matching thresholds of 0.20, 0.30, 0.40, and 0.50. For each product and threshold, the plotted value is the median of the corresponding city-level metric values across the cities covered by that product.
Figure 14. Sensitivity of Tier 2 and Tier 3 metrics to the IoU matching threshold. Panels (ac) show the Tier 2 metrics: (a) Object F1, (b) Mean average object-level IoU, and (c) RMSE. Panels (dg) show the Tier 3 metrics: (d) PoLiS distance, (e) Turning function distance, (f) Compactness, and (g) Fourier-descriptor distance. Metrics in panels (bg) are computed for matched pairs. The four lines represent IoU matching thresholds of 0.20, 0.30, 0.40, and 0.50. For each product and threshold, the plotted value is the median of the corresponding city-level metric values across the cities covered by that product.
Remotesensing 18 02895 g014
Table 1. Overview of the open-access building footprint datasets.
Table 1. Overview of the open-access building footprint datasets.
DatasetsSpatial ExtentGeneration Modality *Update Status
Google Research Open Buildings (GOB)Global SouthGeoAI on VHR ImageryMay 2023
Microsoft Global ML Building Footprints (MBF)GlobalGeoAI on VHR Imagery3 February 2026
MapboxContinental (Europe)VGI & Commercial AggregationContinuously updated
GABLENational (China)GeoAI on VHR Imagery27 August 2024
East Asian BuildingsContinental (East Asia)GeoAI on VHR Imagery22 July 2023
Tencent Map [30]National (China)Proprietary Survey & GeoAIContinuously updated
3D-GloBFPNational (China)GeoAI22 May 2025
China 90Cities Rooftop Data (China-90C)Regional (90 Cities)GeoAI on VHR Imagery21 October 2022
OpenStreetMap (OSM)GlobalVGI (Crowdsourcing)Continuously updated
CMABNational (China)Multi-source GeoAI11 March 2025
ONEGEOGlobalGlobal Data AggregationContinuously updated
* GeoAI (Geospatial Artificial Intelligence) denotes automated extraction frameworks driven by machine learning or deep learning. VHR stands for very high resolution. VGI represents volunteered geographic information.
Table 2. Comparison of reference data in building footprint evaluation studies.
Table 2. Comparison of reference data in building footprint evaluation studies.
StudyReference Data SourceSpatial CoverageNo. of Cities/RegionsApprox. Reference FootprintsDatasets EvaluatedQuality Check Protocol Documented
Herfort et al. [8]Proxy indicators (machine-learning-inferred)Global (13,189 agglomerations)No direct footprintsOSM onlyNot applicable
Chamberlain et al. [31]Datasets-to-Datasets comparisonAfrica-wide4Not explicitly reported
Minghini et al. [32]Authoritative/cadastral recordsContinental (Europe)54Not explicitly reported
Brovelli & Zamboni [33]Authoritative sourcesRegional (Lombard)12.8 millionOSM onlyNot explicitly reported
Nyandwi et al. [34]Local reference dataNational (Rwanda)1Not explicitly reportedGoogle, MicrosoftNot explicitly reported
Okyere et al. [35]Manual annotation5 metropolises5Google, Microsoft, OSMNot explicitly reported
This studyExpert-delineated annotationGlobal (10 cities across diverse contexts)10>200,00011Visual-algorithmic dual validation
Table 3. Characteristics of the selected areas and the corresponding reference data.
Table 3. Characteristics of the selected areas and the corresponding reference data.
City NameCountry/RegionUrban Morphology CategoryStudy Area Size (km2)Number of Reference BuildingsTemporal BaselineSource ImageryNumber of Covered Datasets
ChongqingChinaMountainous Megacity81.415,9972024-8GF-077
DelinghaChinaArid Plateau Settlement15.839712024-7GF-074
DurbanSouth AfricaCoastal Hilly Urban Area15.510,0182024-10GF-023
EdmontonCanadaHigh-Latitude Grid Layout8.385492024-8Aerial2
CayenneFrench GuianaTropical Forest Urban Fringe10.010,0612024-1GF-073
HangzhouChinaRiver Delta Megacity11.021,5852024-8Aerial7
MelbourneAustraliaPlanned Low-Density Suburban23.250,4852024-4Pleiades3
ShanghaiChinaHigh-Density Global Hub16.435,8172025-4GF-079
St. GallenSwitzerlandEuropean Alpine Morphology3.352802024-7Aerial2
XuzhouChinaIndustrial and Mining City16.939,7482024-9GF-073
Table 4. Coverage-conditional tier and overall scores for the 11 evaluated datasets.
Table 4. Coverage-conditional tier and overall scores for the 11 evaluated datasets.
DatasetNumber of Covered CitiesTier 1Tier 2Tier 3Overall Score
3D-GloBFP50.5950.4350.6960.576
CMAB40.5230.3430.6200.495
China-90C30.6090.4530.4950.519
East Asian buildings50.6310.4310.6940.585
GABLE30.8530.5020.7590.705
GOB20.6410.5610.8030.668
Mapbox20.6100.4230.6670.567
MBF50.7510.5870.7090.682
ONEGEO20.4190.3880.6840.497
OSM100.4380.4440.7150.532
Tencent Map20.5730.4680.6980.580
* Tiers 1–3 denote Statistical Consistency, Object Recovery and Spatial Fidelity, and Matched-Pair Morphology Fidelity, respectively. Tier 3 is calculated only for accepted one-to-one matches. Boldface and underlining indicate the largest and second-largest values, respectively, within each score column.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhou, T.; Liao, Y.; Gan, W.; Chen, Z.; Li, C.; Zhao, G.; Chen, Q.; Xue, Q.; Tao, P. Benchmarking Open-Access Building Footprints: A Multi-Dimensional Assessment with High-Fidelity References. Remote Sens. 2026, 18, 2895. https://doi.org/10.3390/rs18172895

AMA Style

Zhou T, Liao Y, Gan W, Chen Z, Li C, Zhao G, Chen Q, Xue Q, Tao P. Benchmarking Open-Access Building Footprints: A Multi-Dimensional Assessment with High-Fidelity References. Remote Sensing. 2026; 18(17):2895. https://doi.org/10.3390/rs18172895

Chicago/Turabian Style

Zhou, Taiqing, Yifan Liao, Wenxiang Gan, Zhijie Chen, Chuchu Li, Geyi Zhao, Qi Chen, Qian Xue, and Pengjie Tao. 2026. "Benchmarking Open-Access Building Footprints: A Multi-Dimensional Assessment with High-Fidelity References" Remote Sensing 18, no. 17: 2895. https://doi.org/10.3390/rs18172895

APA Style

Zhou, T., Liao, Y., Gan, W., Chen, Z., Li, C., Zhao, G., Chen, Q., Xue, Q., & Tao, P. (2026). Benchmarking Open-Access Building Footprints: A Multi-Dimensional Assessment with High-Fidelity References. Remote Sensing, 18(17), 2895. https://doi.org/10.3390/rs18172895

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop