iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision
Highlights
- Automated selection methods that read a model’s outputs (uncertainty sampling, self-training, label propagation, and random selection) plateau well below dense supervision on ISPRS Vaihingen, even when an entropy oracle receives ground truth answers at up to 100 times the budget.
- An expert clicking confidently incorrect pixels, at most one per class per frame per iteration and with no auxiliary machinery, reaches 99.6% and 96.9% of the dense mIoU on ISPRS Vaihingen and BsB Aerial.
- Automated acquisition rules cannot find a model’s confident errors in its own outputs; a human inspecting the prediction overlay supplies exactly this missing signal, making expert clicks a sufficient replacement for the label-expansion machinery of prior pipelines.
- Segmentation models can be trained and maintained at a small fraction of the dense annotation effort (14 to 24 times fewer pointing actions), supported by the open-source iSAGE platform.
Abstract
1. Introduction
- A blind spot of output-reading acquisition. We show that four families of methods built on output-reading acquisition functions (uncertainty sampling, self-training, label propagation, and random selection) have a ceiling well below dense supervision, plateauing at 62.25 to 69.35 mIoU on ISPRS Vaihingen against 76.93.
- The human solution. We propose a human-in-the-loop framework that targets this blind spot directly, with no additional machinery and a restrictive budget of at most one labeled pixel per class per frame per iteration, reaching near-dense performance (99.6% and 96.9% of the dense mIoU on ISPRS Vaihingen and BsB Aerial). To the best of our knowledge, no prior work has run this experiment, and among the 35 methods surveyed, this is the only iterative human-in-the-loop configuration with no output-reading mechanism.
- iSAGE. The open-source platform couples inspection, annotation, record-keeping, and retraining in one environment, which no existing tool provides and without which the loop above cannot run. It extends to new domains through preprocessing and encoder choices, and every experiment in this paper was produced on it.
2. Related Work
2.1. Sparse and Weakly Supervised Semantic Segmentation
2.2. Active Learning for Semantic Segmentation
2.3. Interactive Human-in-the-Loop Frameworks for Semantic Segmentation
3. iSAGE Framework
3.1. Sparse Annotations
3.2. Iterative Error-Driven Refinement
| Algorithm 1 Iterative Error-Driven Refinement for Sparse Annotations |
|
3.3. Error-Weighted Dice Loss
Formulation
3.4. Software Platform
- Annotation interface: Hosts prediction overlay and click-driven annotation.
- Record storage: Persists every decision as JSON.
- Session layout: Packages each iteration as a reproducible snapshot.
- Training backend: Closes the loop with EWDL-supervised retraining.
4. Experiments
4.1. Experimental Setup
4.2. Datasets
4.2.1. BsB Aerial
- Small discrete objects (cars): Rigid geometry, sharp boundaries, low spatial frequency.
- Linear connected structures (roads): Elongated, with mixed boundary regimes and frequent occlusion.
- Large polygonal structures (buildings): Rectilinear, with sharp boundaries and rich interior texture.
- Amorphous smooth-boundary regions (permeable areas): Irregular, with mixed textures and high intra-class variance.
4.2.2. ISPRS Vaihingen
4.3. BsB Aerial Experiments
- Binary experiments. Four independent tasks, one per class, to isolate per-class convergence dynamics from class competition and observe how each visual category evolves across iterations. Each task runs for 5 iSAGE iterations after the initial seed.
- Multiclass experiment. Joint training over all four classes plus background, under the same 5-iteration protocol, to test whether multiclass supervision changes per-class behavior relative to the binary baselines.
- Dense supervision upper bound. Models trained with dense ground truth masks under the same architecture, providing the performance ceiling iSAGE is compared against.
- Random selection baseline. The iterative protocol with randomly selected sparse annotations in place of error-driven selection, controlling for whether iSAGE’s gain comes from sparsity per se or from error-targeting. This baseline is the without-expert-guidance control: it shares the pipeline, budget, loss, and training schedule, differing only in who selects the pixels.
- Alternative loss functions. EWDL compared against Binary Cross-Entropy (BCE), Focal, and Dice losses on the binary tasks and against Cross-Entropy (CE), Focal, and Dice on the multiclass task, on the iter-5 final annotation set, testing whether EWDL’s error-weighting actually contributes to iSAGE’s performance or a standard loss suffices.
- EWDL hyperparameter sensitivity. The error-penalty factor varied over {1, 2, 5, 10, 20} on the iter-5 set, multi-seed, to probe the loss’s stability around the chosen .
- Cross-architecture validation. Four model configurations (the primary U-Net + EfficientNet-B7 plus U-Net + ResNet-101, DeepLabV3+ + ResNet-50, and SegFormer + MiT-B2), each paired with a matched dense-supervision baseline and trained under the identical five-seed protocol on the iteration-5 annotation set, covering three decoder families (encoder-decoder, atrous spatial pyramid, hierarchical transformer), to test whether the annotation record transfers across backbones.
4.4. ISPRS Vaihingen Experiments
4.4.1. External Benchmarking
- iSAGE on Vaihingen. The complete iSAGE pipeline (sparse seed annotations, error-driven refinement over 5 iSAGE iterations, EWDL training) is applied to Vaihingen under the same per-iteration budget used throughout (at most one labeled pixel per class per frame), accumulating to at most six labeled pixels per class per frame over the five iterations plus the seed. Performance is compared against published methods on EasySeg’s [29] 17-tile test partition.
4.4.2. Output-Reading Baselines
- Oracle entropy [57,69]: At each iteration, the highest-entropy pixel per predicted class per frame is selected, with ground truth labels assigned at the selected coordinates. Tests whether the strongest possible uncertainty-driven acquisition, given ground truth answers at every query, can match iSAGE. The budget sweep at 10×, 50×, and 100× labels per class per frame reaches up to 0.95% of training pixels and tests whether budget alone closes the gap to iSAGE.
- Pseudo-labeling [25,26]: At each iteration, pixels whose predicted-class confidence exceeds a fixed threshold are converted to pseudo-labels and added to the supervision set. Tests whether self-confidence can expand sparse seeds into adequate supervision without human input. The threshold sweep at 0.90, 0.95, and 0.99 tests whether confidence calibration is the lever that closes the gap.
- CRF-based label propagation [95]: At each iteration, the model’s softmax outputs are refined by a DenseCRF with a Gaussian spatial pairwise term and a bilateral color-position pairwise term (pydensecrf defaults: Gaussian = 3, compatibility 3; bilateral = 80, = 13, compatibility 10; five mean-field iterations) before a confidence threshold matching the strongest pseudo-labeling configuration in the threshold sweep above is applied. Tests whether spatial smoothing of model outputs adds value on top of the best expansion baseline, isolating smoothing from threshold calibration.
- Uniform random: A control baseline that selects one pixel per ground truth class per frame, uniformly at random within that class’s region and independently of the model’s predictions. Tests whether any acquisition structure outperforms the simplest sampling.
5. Results
5.1. BsB Aerial
5.2. ISPRS Vaihingen
5.2.1. External Benchmarking
5.2.2. Output-Reading Baselines
6. Discussion
6.1. iSAGE in the Landscape and Why This Position Is Justified
6.1.1. The Output-Reading Limit
6.1.2. The Only Framework Without Auxiliary Machinery
6.1.3. Theoretical and Empirical Support
6.2. Findings
6.2.1. Match-with-Dense Behavior
6.2.2. Per-Class Dynamics
6.2.3. Annotation-Record Transfer Across Architectures
6.2.4. The Loss Helps Slightly, the Strategy Drives the Result
6.3. Operational and Cognitive Properties of the Workflow
6.3.1. Incremental Deployment and Maintenance
6.3.2. An Auditable, Versionable Record
6.3.3. Directed Rather than Exhaustive Attention
6.3.4. A Model That Is Always Usable and Never Final
6.4. Limitations and Future Work
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Computer Vision—ECCV 2014. Lecture Notes in Computer Science; Fleet, D., Tomas, P., Schiele, B., Tuytelaars, T., Eds.; Springer: Cham/Zurich, Switzerland, 2014; Volume 8693, pp. 740–755. [Google Scholar] [CrossRef] [Scilit]
- Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 3213–3223. [Google Scholar] [CrossRef] [Scilit]
- Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 248–255. [Google Scholar] [CrossRef] [Scilit]
- Whang, S.E.; Roh, Y.; Song, H.; Lee, J.G. Data collection and quality challenges in deep learning: A data-centric AI perspective. VLDB J. 2023, 32, 791–813. [Google Scholar] [CrossRef] [Scilit]
- Zha, D.; Bhat, Z.P.; Lai, K.H.; Yang, F.; Jiang, Z.; Zhong, S.; Hu, X. Data-centric artificial intelligence: A survey. ACM Comput. Surv. 2025, 57, 1–42. [Google Scholar] [CrossRef] [Scilit]
- Sun, C.; Shrivastava, A.; Singh, S.; Gupta, A. Revisiting Unreasonable Effectiveness of Data in Deep Learning Era. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 843–852. [Google Scholar] [CrossRef] [Scilit]
- Rottensteiner, F.; Sohn, G.; Gerke, M.; Wegner, J.D.; Breitkopf, U.; Jung, J. Results of the ISPRS benchmark on urban object detection and 3D building reconstruction. ISPRS J. Photogramm. Remote Sens. 2014, 93, 256–271. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Zheng, Z.; Ma, A.; Lu, X.; Zhong, Y. LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Virtual, 6–14 December 2021. [Google Scholar]
- de Carvalho, O.L.F.; de Carvalho Júnior, O.A.; Silva, C.R.E.; de Albuquerque, A.O.; Santana, N.C.; Borges, D.L.; Gomes, R.A.T.; Guimarães, R.F. Panoptic Segmentation Meets Remote Sensing. Remote Sens. 2022, 14, 965. [Google Scholar] [CrossRef] [Scilit]
- Rahnemoonfar, M.; Chowdhury, T.; Murphy, R. RescueNet: A High Resolution UAV Semantic Segmentation Dataset for Natural Disaster Damage Assessment. Sci. Data 2023, 10, 913. [Google Scholar] [CrossRef] [Scilit]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 3992–4003. [Google Scholar] [CrossRef] [Scilit]
- Lüddecke, T.; Ecker, A. Image Segmentation Using Text and Image Prompts. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 7076–7086. [Google Scholar] [CrossRef] [Scilit]
- Huang, Y.; Yang, X.; Liu, L.; Zhou, H.; Chang, A.; Zhou, X.; Chen, R.; Yu, J.; Chen, J.; Chen, C.; et al. Segment Anything Model for Medical Images? Med. Image Anal. 2024, 92, 103061. [Google Scholar] [CrossRef] [Scilit]
- Ke, L.; Ye, M.; Danelljan, M.; Liu, Y.; Tai, Y.W.; Tang, C.K.; Yu, F. Segment Anything in High Quality. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
- Osco, L.P.; Wu, Q.; de Lemos, E.L.; Gonçalves, W.N.; Ramos, A.P.M.; Li, J.; Junior, J.M. The segment anything model (sam) for remote sensing applications: From zero to one shot. Int. J. Appl. Earth Obs. Geoinf. 2023, 124, 103540. [Google Scholar] [CrossRef] [Scilit]
- Ravi, N.; Gabeur, V.; Hu, Y.T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; Rädle, R.; Rolland, C.; Gustafson, L.; et al. SAM 2: Segment Anything in Images and Videos. In Proceedings of the International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025. [Google Scholar]
- Shaban, A.; Bansal, S.; Liu, Z.; Essa, I.; Boots, B. One-Shot Learning for Semantic Segmentation. In Proceedings of the British Machine Vision Conference (BMVC), London, UK, 4–7 September 2017. [Google Scholar]
- Boudiaf, M.; Kervadec, H.; Masud, Z.I.; Piantanida, P.; Ayed, I.B.; Dolz, J. Few-Shot Segmentation Without Meta-Learning: A Good Transductive Inference Is All You Need? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual, 19–25 June 2021; pp. 13974–13983. [Google Scholar] [CrossRef] [Scilit]
- Bearman, A.; Russakovsky, O.; Ferrari, V.; Fei-Fei, L. What’s the point: Semantic segmentation with point supervision. In Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); Springer: Berlin/Heidelberg, Germany, 2016; Volume 9911, pp. 549–565. [Google Scholar] [CrossRef] [Scilit]
- Hua, Y.; Marcos, D.; Mou, L.; Zhu, X.X.; Tuia, D. Semantic segmentation of remote sensing images with sparse annotations. IEEE Geosci. Remote Sens. Lett. 2021, 19, 1–5. [Google Scholar] [CrossRef] [Scilit]
- Khoreva, A.; Benenson, R.; Hosang, J.; Hein, M.; Schiele, B. Simple Does It: Weakly Supervised Instance and Semantic Segmentation. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1665–1674. [Google Scholar] [CrossRef] [Scilit]
- Sener, O.; Savarese, S. Active learning for convolutional neural networks: A core-set approach. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Kellenberger, B.; Marcos, D.; Lobry, S.; Tuia, D. Half a percent of labels is enough: Efficient animal detection in UAV imagery using deep CNNs and active learning. IEEE Trans. Geosci. Remote Sens. 2019, 57, 9524–9533. [Google Scholar] [CrossRef] [Scilit]
- Mackowiak, R.; Lenz, P.; Ghori, O.; Diego, F.; Lange, O.; Rother, C. CEREALS-Cost-Effective REgion-based Active Learning for Semantic Segmentation. In Proceedings of the British Machine Vision Conference (BMVC), Newcastle upon Tyne, UK, 3–6 September 2018. [Google Scholar]
- Lee, D.H. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Proceedings of the Workshop on Challenges in Representation Learning, ICML, Atlanta, GA, USA, 21 June 2013; Volume 3, p. 896. [Google Scholar]
- Arazo, E.; Ortego, D.; Albert, P.; O’Connor, N.E.; McGuinness, K. Pseudo-Labeling and Confirmation Bias in Deep Semi-Supervised Learning. In Proceedings of the International Joint Conference on Neural Networks (IJCNN), Virtual, 19–24 July 2020; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
- Lenczner, G.; Chan-Hon-Tong, A.; Le Saux, B.; Luminari, N.; Le Besnerais, G. DIAL: Deep interactive and active learning for semantic segmentation in remote sensing. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 3376–3389. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Xia, M.; Jiao, J.; Zhou, S.; Chang, C.; Wang, Y.; Guo, Y. HAL-IA: A Hybrid Active Learning framework using Interactive Annotation for medical image segmentation. Med. Image Anal. 2023, 88, 102862. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Chen, H.; Yang, A.; Li, J. EasySeg: An Error-Aware Domain Adaptation Framework for Remote Sensing Imagery Semantic Segmentation via Interactive Learning and Active Learning. IEEE Trans. Geosci. Remote Sens. 2024, 62, 1–18. [Google Scholar] [CrossRef] [Scilit]
- Liu, P.; Liu, J. When Confidence Fails: Revisiting Pseudo-Label Selection in Semi-supervised Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, HI, USA, 19–23 October 2025. [Google Scholar]
- Jin, Q.; Yuan, M.; Li, S.; Wang, H.; Wang, M.; Song, Z. Cold-start active learning for image classification. Inf. Sci. 2022, 616, 16–36. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Lian, Y.; Bai, J.; Zhang, J.; Xiao, Z.; Hou, B. Weakly Supervised Semantic Segmentation of Remote Sensing Images Using Siamese Affinity Network. Remote Sens. 2025, 17, 808. [Google Scholar] [CrossRef] [Scilit]
- Teng, Y.; Wang, L. Structured Sparse R-CNN for Direct Scene Graph Generation. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 19415–19424. [Google Scholar] [CrossRef] [Scilit]
- Lin, D.; Dai, J.; Jia, J.; He, K.; Sun, J. ScribbleSup: Scribble-Supervised Convolutional Networks for Semantic Segmentation. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 3159–3167. [Google Scholar] [CrossRef] [Scilit]
- Çiçek, Ö.; Abdulkadir, A.; Lienkamp, S.S.; Brox, T.; Ronneberger, O. 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2016; Ourselin, S., Joskowicz, L., Sabuncu, M., Unal, G., Wells, W., Eds.; Springer: Cham, Switzerland, 2016; pp. 424–432. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Qi, X.; Fu, C.W. One Thing One Click: A Self-Training Approach for Weakly Supervised 3D Semantic Segmentation. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2021; pp. 1726–1736. [Google Scholar] [CrossRef] [Scilit]
- Gao, F.; Hu, M.; Zhong, M.E.; Feng, S.; Tian, X.; Meng, X.; Huang, Z.; Lv, M.; Song, T.; Zhang, X.; et al. Segmentation only uses sparse annotations: Unified weakly and semi-supervised learning in medical images. Med. Image Anal. 2022, 80, 102515. [Google Scholar] [CrossRef] [Scilit]
- Kervadec, H.; Dolz, J.; Tang, M.; Granger, E.; Boykov, Y.; Ben Ayed, I. Constrained-CNN losses for weakly supervised segmentation. Med. Image Anal. 2019, 54, 88–99. [Google Scholar] [CrossRef] [Scilit]
- Yang, G.; Wang, C.; Yang, J.; Chen, Y.; Tang, L.; Shao, P.; Dillenseger, J.L.; Shu, H.; Luo, L. Weakly-supervised convolutional neural networks of renal tumor segmentation in abdominal CTA images. BMC Med. Imaging 2020, 20, 37. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Liu, Q.; Zhang, Y.; Wang, M.; Tang, J. TSSK-Net: Weakly supervised biomarker localization and segmentation with image-level annotation in retinal OCT images. Comput. Biol. Med. 2023, 153, 106467. [Google Scholar] [CrossRef] [Scilit]
- Maggiolo, L.; Marcos, D.; Moser, G.; Serpico, S.B.; Tuia, D. A Semisupervised CRF Model for CNN-Based Semantic Segmentation with Sparse Ground Truth. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Mazhar, S.; Sun, G.; Bilal, A.; Hassan, B.; Li, Y.; Zhang, J.; Lin, Y.; Khan, A.; Ahmed, R.; Hassan, T. AUnet: A Deep Learning Framework for Surface Water Channel Mapping Using Large-Coverage Remote Sensing Images and Sparse Scribble Annotations from OSM Data. Remote Sens. 2022, 14, 3283. [Google Scholar] [CrossRef] [Scilit]
- Gbodjo, Y.J.E.; Montet, O.; Ienco, D.; Gaetano, R.; Dupuy, S. Multisensor Land Cover Classification With Sparsely Annotated Data Based on Convolutional Neural Networks and Self-Distillation. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2021, 14, 11485–11499. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.C.; Papandreou, G.; Schroff, F.; Adam, H. Rethinking Atrous Convolution for Semantic Image Segmentation. arXiv 2017, arXiv:1706.05587. [Google Scholar]
- Can, Y.B.; Chaitanya, K.; Mustafa, B.; Koch, L.M.; Konukoglu, E.; Baumgartner, C.F. Learning to Segment Medical Images with Scribble-Supervision Alone. In Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); Stoyanov, D., Taylor, Z., Carneiro, G., Syeda-Mahmood, T., Martel, A., Maier-Hein, L., Tavares, J.M.R., Bradley, A., Papa, J.P., Belagiannis, V., et al., Eds.; Springer International Publishing: Cham, Switzerland, 2018; Volume 11045, pp. 236–244. [Google Scholar] [CrossRef] [Scilit]
- Arnab, A.; Zheng, S.; Jayasumana, S.; Romera-Paredes, B.; Larsson, M.; Kirillov, A.; Savchynskyy, B.; Rother, C.; Kahl, F.; Torr, P.H. Conditional Random Fields Meet Deep Neural Networks for Semantic Segmentation: Combining Probabilistic Graphical Models with Deep Learning for Structured Prediction. IEEE Signal Process. Mag. 2018, 35, 37–52. [Google Scholar] [CrossRef] [Scilit]
- Liang, Z.; Wang, T.; Zhang, X.; Sun, J.; Shen, J. Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 16886–16895. [Google Scholar] [CrossRef] [Scilit]
- Wang, G.; Luo, X.; Gu, R.; Yang, S.; Qu, Y.; Zhai, S.; Zhao, Q.; Li, K.; Zhang, S. PyMIC: A deep learning toolkit for annotation-efficient medical image segmentation. Comput. Methods Programs Biomed. 2023, 231, 107398. [Google Scholar] [CrossRef] [Scilit]
- Belharbi, S.; Ayed, I.B.; McCaffrey, L.; Granger, E. Deep Active Learning for Joint Classification & Segmentation with Weak Annotator. In Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2021; pp. 3337–3346. [Google Scholar] [CrossRef] [Scilit]
- Ren, Q.; Zhang, H.; Zhang, D.; Zhao, X.; Yan, L.; Rui, J.; Zeng, F.; Zhu, X. A framework of active learning and semi-supervised learning for lithology identification based on improved naive Bayes. Expert Syst. Appl. 2022, 202, 117278. [Google Scholar] [CrossRef] [Scilit]
- Desai, S.; Ghose, D. Active Learning for Improved Semi-Supervised Semantic Segmentation in Satellite Images. In Proceedings of the 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2022; pp. 1485–1495. [Google Scholar] [CrossRef] [Scilit]
- Alonso, I.; Yuval, M.; Eyal, G.; Treibitz, T.; Murillo, A.C. CoralSeg: Learning coral segmentation from sparse annotations. J. Field Robot. 2019, 36, 1456–1477. [Google Scholar] [CrossRef] [Scilit]
- Lee, H.; Jeong, W.K. Scribble2Label: Scribble-Supervised Cell Segmentation via Self-generating Pseudo-Labels with Consistency. In Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); Martel, A.L., Abolmaesumi, P., Stoyanov, D., Mateus, D., Zuluaga, M.A., Zhou, S.K., Racoceanu, D., Joskowicz, L., Eds.; Springer International Publishing: Cham, Switzerland, 2020; Volume 12261, pp. 14–23. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Jia, M.; Sun, G.; Zhang, A. PAMSNet: A point annotation-driven multi-source network for remote sensing semantic segmentation. ISPRS J. Photogramm. Remote Sens. 2025, 229, 1–16. [Google Scholar] [CrossRef] [Scilit]
- Chan, S.; Zhou, W.; Lei, Y.; Li, C.; Hu, J.; Hong, F. Sparse point annotations for remote sensing image segmentation. Sci. Rep. 2025, 15, 27347. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Huang, X.; Weng, Q. A SAM-adapted weakly-supervised semantic segmentation method constrained by uncertainty and transformation consistency. Int. J. Appl. Earth Obs. Geoinf. 2025, 137, 104440. [Google Scholar] [CrossRef] [Scilit]
- Ren, P.; Xiao, Y.; Chang, X.; Huang, P.Y.; Li, Z.; Gupta, B.B.; Chen, X.; Wang, X. A Survey of Deep Active Learning. ACM Comput. Surv. 2021, 54, 1–40. [Google Scholar] [CrossRef] [Scilit]
- Yoo, D.; Kweon, I.S. Learning Loss for Active Learning. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 93–102. [Google Scholar] [CrossRef] [Scilit]
- Yuan, T.; Wan, F.; Fu, M.; Liu, J.; Xu, S.; Ji, X.; Ye, Q. Multiple Instance Active Learning for Object Detection. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2021; pp. 5326–5335. [Google Scholar] [CrossRef] [Scilit]
- Yamani, A.; Alyami, A.; Luqman, H.; Ghanem, B.; Giancola, S. Active Learning for Single-Stage Object Detection in UAV Images. In Proceedings of the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2024; pp. 1849–1858. [Google Scholar] [CrossRef] [Scilit]
- Lai, Z.; Wang, C.; Oliveira, L.C.; Dugger, B.N.; Cheung, S.C.; Chuah, C.N. Joint Semi-supervised and Active Learning for Segmentation of Gigapixel Pathology Images with Cost-Effective Labeling. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); IEEE: New York, NY, USA, 2021; pp. 591–600. [Google Scholar] [CrossRef] [Scilit]
- Rangnekar, A.; Kanan, C.; Hoffman, M. Semantic Segmentation with Active Semi-Supervised Learning. In Proceedings of the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2023; pp. 5955–5966. [Google Scholar] [CrossRef] [Scilit]
- Fan, C.; Wu, Q.; Zhao, Y.; Mo, L. Integrating active learning and semi-supervised learning for improved data-driven HVAC fault diagnosis performance. Appl. Energy 2024, 356, 122356. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Ma, B.; Cui, H.; Xia, Y. Think Twice Before Selection: Federated Evidential Active Learning for Medical Image Analysis with Domain Shifts. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 11439–11449. [Google Scholar] [CrossRef] [Scilit]
- Ge, J.; Zhang, Z.; Phan, M.H.; Zhang, B.; Liu, A.; Zhao, Y.; Zhao, S. ESA: Annotation-Efficient Active Learning for Semantic Segmentation. In Proceedings of the Advanced Intelligent Computing Technology and Applications (ICIC); Springer: Berlin/Heidelberg, Germany, 2025; pp. 141–152. [Google Scholar] [CrossRef] [Scilit]
- Didari, S.; Hu, W.; Woo, J.O.; Hao, H.; Moon, H.; Min, S. Bayesian Active Learning for Semantic Segmentation. arXiv 2024, arXiv:2408.01694. [Google Scholar]
- Siddiqui, Y.; Valentin, J.; Niessner, M. ViewAL: Active Learning With Viewpoint Entropy for Semantic Segmentation. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–18 June 2020; pp. 9430–9440. [Google Scholar] [CrossRef] [Scilit]
- Wang, P.; Yao, W.; Shao, J. One class one click: Quasi scene-level weakly supervised point cloud semantic segmentation with active learning. ISPRS J. Photogramm. Remote Sens. 2023, 204, 89–104. [Google Scholar] [CrossRef] [Scilit]
- Mukhoti, J.; Gal, Y. Evaluating Bayesian Deep Learning Methods for Semantic Segmentation. arXiv 2018, arXiv:1811.12709. [Google Scholar]
- Gustafsson, F.K.; Danelljan, M.; Schön, T.B. Evaluating Scalable Bayesian Deep Learning Methods for Robust Computer Vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Virtual, 14–19 June 2020. [Google Scholar] [CrossRef] [Scilit]
- Mittal, S.; Niemeijer, J.; Çiçek, Ö.; Tatarchenko, M.; Ehrhardt, J.; Schäfer, J.P.; Handels, H.; Brox, T. Realistic Evaluation of Deep Active Learning for Image Classification and Semantic Segmentation. Int. J. Comput. Vis. 2025, 133, 4294–4316. [Google Scholar] [CrossRef] [Scilit]
- Amershi, S.; Cakmak, M.; Knox, W.B.; Kulesza, T. Power to the People: The Role of Humans in Interactive Machine Learning. AI Mag. 2014, 35, 105–120. [Google Scholar] [CrossRef] [Scilit]
- Mosqueira-Rey, E.; Hernández-Pereira, E.; Alonso-Ríos, D.; Bobes-Bascarán, J.; Fernández-Leal, Á. Human-in-the-loop machine learning: A state of the art. Artif. Intell. Rev. 2023, 56, 3005–3054. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Xiao, L.; Sun, Y.; Zhang, J.; Ma, T.; He, L. A Survey of Human-in-the-loop for Machine Learning. Future Gener. Comput. Syst. 2022, 135, 364–381. [Google Scholar] [CrossRef] [Scilit]
- Xu, N.; Price, B.; Cohen, S.; Yang, J.; Huang, T.S. Deep Interactive Object Selection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 373–381. [Google Scholar] [CrossRef] [Scilit]
- Maninis, K.K.; Caelles, S.; Pont-Tuset, J.; Van Gool, L. Deep Extreme Cut: From Extreme Points to Object Segmentation. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 616–625. [Google Scholar] [CrossRef] [Scilit]
- Sofiiuk, K.; Petrov, I.A.; Konushin, A. Reviving Iterative Training with Mask Guidance for Interactive Segmentation. In Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP); IEEE: New York, NY, USA, 2022; pp. 3141–3145. [Google Scholar] [CrossRef] [Scilit]
- Smith, A.G.; Han, E.; Petersen, J.; Olsen, N.A.F.; Giese, C.; Athmann, M.; Dresbøll, D.B.; Thorup-Kristensen, K. RootPainter: Deep learning segmentation of biological images with corrective annotation. New Phytol. 2022, 236, 774–791. [Google Scholar] [CrossRef] [Scilit]
- Ho, D.J.; Agaram, N.P.; Schüffler, P.J.; Vanderbilt, C.M.; Jean, M.H.; Hameed, M.R.; Fuchs, T.J. Deep Interactive Learning: An Efficient Labeling Approach for Deep Learning-Based Osteosarcoma Treatment Response Assessment. In Proceedings of the Medical Image Computing and Computer Assisted Intervention (MICCAI); Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2020; Volume 12265, pp. 540–549. [Google Scholar] [CrossRef] [Scilit]
- Kim, H.; Hwang, S.; Kwak, S.; Ok, J. Active Label Correction for Semantic Segmentation with Foundation Models. In Proceedings of the 41st International Conference on Machine Learning (ICML), Vienna, Austria, 21–27 July 2024. [Google Scholar]
- Jeon, Y.; Cho, K.; Woo, S.; Kim, E. A2LC: Active and Automated Label Correction for Semantic Segmentation. Proc. Aaai Conf. Artif. Intell. 2026, 40, 5296–5304. [Google Scholar] [CrossRef] [Scilit]
- Liu, N.; Xu, X.; Su, Y.; Zhang, H.; Li, H.C. PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Carvalho, O.L.F.D.; Carvalho Júnior, O.A.D.; Albuquerque, A.O.D.; Silva, D.G.E. Remote SAMsing: From Segment Anything to Segment Everything. Int. J. Appl. Earth Obs. Geoinf. 2026, 153, 105528. [Google Scholar] [CrossRef] [Scilit]
- Iakubovskii, P. Segmentation Models Pytorch. 2019. Available online: https://github.com/qubvel/segmentation_models.pytorch (accessed on 27 August 2026).
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019; Volume 32. [Google Scholar]
- Kluyver, T.; Ragan-Kelley, B.; Pérez, F.; Granger, B.; Bussonnier, M.; Frederic, J.; Kelley, K.; Hamrick, J.; Grout, J.; Corlay, S.; et al. Jupyter Notebooks—A Publishing Format for Reproducible Computational Workflows. In Proceedings of the Positioning and Power in Academic Publishing: Players, Agents and Agendas; IOS Press: Amsterdam, The Netherlands, 2016; pp. 87–90. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-assisted Intervention; Springer: Berlin/Heidelberg, Germany, 2015; pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Le, Q. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning; Chaudhuri, K., Salakhutdinov, R., Eds.; PMLR: London, UK, 2019; Volume 97, pp. 6105–6114. Available online: http://arxiv.org/abs/1905.11946 (accessed on 27 August 2026).
- Xie, B.; Yuan, L.; Li, S.; Liu, C.H.; Cheng, X. Towards Fewer Annotations: Active Learning via Region Impurity and Prediction Uncertainty for Domain Adaptive Semantic Segmentation. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 8058–8068. [Google Scholar] [CrossRef] [Scilit]
- Wu, T.H.; Liou, Y.S.; Yuan, S.J.; Lee, H.Y.; Chen, T.I.; Huang, K.C.; Hsu, W.H. D2ADA: Dynamic Density-aware Active Domain Adaptation for Semantic Segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 449–467. [Google Scholar] [CrossRef] [Scilit]
- Guan, L.; Yuan, X. Iterative Loop Method Combining Active and Semi-supervised Learning for Domain Adaptive Semantic Segmentation. arXiv 2023, arXiv:2301.13361. [Google Scholar]
- Douglas, D.H.; Peucker, T.K. Algorithms for the Reduction of the Number of Points Required to Represent a Digitized Line or its Caricature. Cartogr. Int. J. Geogr. Inf. Geovis. 1973, 10, 112–122. [Google Scholar] [CrossRef] [Scilit]
- de Carvalho, O.L.F.; de Carvalho Júnior, O.A.; de Albuquerque, A.O.; Santana, N.C.; Guimarães, R.F.; Gomes, R.A.T.; Borges, D.L. Bounding box-free instance segmentation using semi-supervised iterative learning for vehicle detection. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 3403–3420. [Google Scholar] [CrossRef] [Scilit]
- de Carvalho, O.L.F.; de Carvalho Júnior, O.A.; de Albuquerque, A.O.; Santana, N.C.; Borges, D.L. Rethinking Panoptic Segmentation in Remote Sensing: A Hybrid Approach Using Semantic Segmentation and Non-Learning Methods. IEEE Geosci. Remote Sens. Lett. 2022, 19, 1–5. [Google Scholar] [CrossRef] [Scilit]
- Krähenbühl, P.; Koltun, V. Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials. Adv. Neural Inf. Process. Syst. 2011, 24, 109–117. [Google Scholar]
- Vu, T.H.; Jain, H.; Bucher, M.; Cord, M.; Pérez, P. ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; pp. 2512–2521. [Google Scholar] [CrossRef] [Scilit]
- Luo, Y.; Zheng, L.; Guan, T.; Yu, J.; Yang, Y. Taking a Closer Look at Domain Shift: Category-Level Adversaries for Semantics Consistent Domain Adaptation. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; pp. 2502–2511. [Google Scholar] [CrossRef] [Scilit]
- Ning, M.; Lu, D.; Xie, Y.; Chen, D.; Wei, D.; Zheng, Y.; Tian, Y.; Yan, S.; Yuan, L. MADAv2: Advanced Multi-Anchor Based Active Domain Adaptation Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13553–13566. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Zheng, Y.; Shan, D.; Yang, S.; Li, Q.; Wang, B.; Zhang, Y.; Hong, Q.; Shen, D. ScribFormer: Transformer Makes CNN Work Better for Scribble-based Medical Image Segmentation. IEEE Trans. Med. Imaging 2024, 43, 2254–2265. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Zhao, Z.; Zhao, H. UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 3031–3048. [Google Scholar] [CrossRef] [Scilit]
- Kim, H.; Oh, M.; Hwang, S.; Kwak, S.; Ok, J. Adaptive Superpixel for Active Learning in Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 943–953. [Google Scholar] [CrossRef] [Scilit]
- Wang, K.; Liew, J.H.; Zou, Y.; Zhou, D.; Feng, J. PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9196–9205. [Google Scholar] [CrossRef] [Scilit]
- Zou, X.; Yang, J.; Zhang, H.; Li, F.; Li, L.; Wang, J.; Wang, L.; Gao, J.; Lee, Y.J. Segment Everything Everywhere All at Once. Adv. Neural Inf. Process. Syst. 2023, 36, 19769–19782. [Google Scholar] [CrossRef] [Scilit]
- Bickford Smith, F.; Kossen, J.; Trollope, E.; van der Wilk, M.; Foster, A.; Rainforth, T. Rethinking Aleatoric and Epistemic Uncertainty. In Proceedings of the 42nd International Conference on Machine Learning (ICML), Vancouver, BC, Canada, 13–19 July 2025. [Google Scholar]
- Kosarevych, R.; Lutsyk, O.; Rusyn, B.; Pits, N.; Maksymyuk, T.; Volosin, M. Adaptive Patch Reshaping for Edge-Based Semantic Segmentation in Remote Sensing. IEEE Access 2026, 14, 38951–38964. [Google Scholar] [CrossRef] [Scilit]
- Zhang, E.; Lyngaas, I.; Chen, P.; Wang, X.; Igarashi, J.; Huo, Y.; Munetomo, M.; Wahib, M. Adaptive Patching for High-resolution Image Segmentation with Transformers. In Proceedings of the SC24: International Conference for High Performance Computing, Networking, Storage and Analysis; IEEE: New York, NY, USA, 2024; pp. 1–16. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Shi, S.; Wang, J.; Zhong, Y. Seeing Beyond the Patch: Scale-Adaptive Semantic Segmentation of High-resolution Remote Sensing Imagery based on Reinforcement Learning. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 16822–16832. [Google Scholar] [CrossRef] [Scilit]
- de Carvalho, O.L.F. ISAGE: Iterative Sparse Annotation Guided by Expert (v1.0.0); Zenodo. 2026. Available online: https://zenodo.org/records/20596186 (accessed on 27 August 2026). [CrossRef]
- de Carvalho, O.L.F.; de Carvalho Júnior, O.A.; de Albuquerque, A.O.; Guerreiro e Silva, D. BsB Aerial Dataset; Zenodo. 2026. Available online: https://zenodo.org/records/20635237 (accessed on 27 August 2026). [CrossRef]







| Setting | Value |
|---|---|
| Architecture | U-Net [87] |
| Encoder | EfficientNet-B7 [88] |
| Optimizer | Adam (, ) |
| Learning rate | |
| Batch size | 10 |
| Epochs per iteration | 100 |
| Checkpoint selection | Final epoch |
| Data augmentation | Horizontal and vertical flips |
| Setting | Car | Road | Building | Permeable Area | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| IoU | Prec. | Recall | F1 | Effort | IoU | Prec. | Recall | F1 | Effort | IoU | Prec. | Recall | F1 | Effort | IoU | Prec. | Recall | F1 | Effort | |
| iSAGE (EWDL) | ||||||||||||||||||||
| Iter 0 | 27.44 | 30.69 | 72.20 | 43.07 | 0.0029 | 68.22 | 80.37 | 81.85 | 81.10 | 0.0030 | 54.98 | 60.52 | 85.71 | 70.95 | 0.0024 | 81.73 | 88.80 | 91.12 | 89.94 | 0.0030 |
| Iter 1 | 39.12 | 44.04 | 77.79 | 56.24 | 0.0058 | 78.27 | 87.74 | 87.88 | 87.81 | 0.0059 | 67.99 | 73.26 | 90.43 | 80.95 | 0.0048 | 87.03 | 91.91 | 94.25 | 93.07 | 0.0059 |
| Iter 2 | 61.09 | 73.32 | 78.55 | 75.85 | 0.0087 | 80.30 | 89.82 | 88.34 | 89.07 | 0.0086 | 76.70 | 82.40 | 91.73 | 86.81 | 0.0072 | 87.53 | 92.43 | 94.29 | 93.35 | 0.0085 |
| Iter 3 | 71.38 | 84.15 | 82.46 | 83.30 | 0.0116 | 81.47 | 90.37 | 89.22 | 89.79 | 0.0110 | 80.26 | 87.50 | 90.65 | 89.05 | 0.0096 | 88.03 | 92.14 | 95.18 | 93.63 | 0.0106 |
| Iter 4 | 73.12 | 86.44 | 82.59 | 84.47 | 0.0145 | 81.67 | 91.01 | 88.84 | 89.91 | 0.0128 | 80.66 | 86.43 | 92.36 | 89.30 | 0.0121 | 88.09 | 92.02 | 95.38 | 93.67 | 0.0122 |
| Iter 5 | 74.05 | 86.02 | 84.18 | 85.09 | 0.0174 | 81.88 | 91.24 | 88.86 | 90.04 | 0.0140 | 81.48 | 88.62 | 91.01 | 89.80 | 0.0145 | 88.17 | 92.48 | 94.98 | 93.71 | 0.0130 |
| Iterative Random Selection | ||||||||||||||||||||
| Iter 0 | 38.52 | 39.92 | 91.64 | 55.62 | 0.0029 | 71.46 | 77.95 | 89.56 | 83.35 | 0.0030 | 68.91 | 73.93 | 91.03 | 81.59 | 0.0024 | 84.92 | 88.86 | 95.04 | 91.85 | 0.0030 |
| Iter 1 | 43.03 | 44.26 | 93.93 | 60.17 | 0.0058 | 72.04 | 77.98 | 90.44 | 83.75 | 0.0059 | 70.67 | 75.56 | 91.61 | 82.82 | 0.0050 | 85.94 | 90.48 | 94.48 | 92.44 | 0.0060 |
| Iter 2 | 43.51 | 44.61 | 94.60 | 60.63 | 0.0087 | 74.22 | 80.78 | 90.14 | 85.20 | 0.0089 | 71.80 | 76.50 | 92.13 | 83.59 | 0.0075 | 85.96 | 90.91 | 94.04 | 92.45 | 0.0091 |
| Iter 3 | 45.12 | 46.13 | 95.37 | 62.19 | 0.0116 | 74.59 | 80.47 | 91.07 | 85.45 | 0.0118 | 72.96 | 77.83 | 92.10 | 84.37 | 0.0100 | 86.41 | 90.02 | 95.57 | 92.71 | 0.0121 |
| Iter 4 | 44.84 | 45.76 | 95.68 | 61.92 | 0.0145 | 75.11 | 80.57 | 91.72 | 85.79 | 0.0148 | 73.29 | 78.63 | 91.52 | 84.59 | 0.0126 | 86.72 | 91.18 | 94.66 | 92.89 | 0.0151 |
| Iter 5 | 48.47 | 49.85 | 94.61 | 65.30 | 0.0174 | 75.67 | 80.48 | 92.68 | 86.15 | 0.0177 | 73.82 | 78.91 | 91.95 | 84.94 | 0.0151 | 86.69 | 91.87 | 93.89 | 92.87 | 0.0181 |
| Alternative Loss Functions (Final Iteration | ||||||||||||||||||||
| BCE | 74.49 | 87.25 | 83.58 | 85.38 | 0.0174 | 81.71 | 91.70 | 88.24 | 89.94 | 0.0140 | 81.53 | 89.23 | 90.43 | 89.83 | 0.0145 | 87.85 | 90.90 | 96.33 | 93.53 | 0.0130 |
| Focal | 74.01 | 86.37 | 83.80 | 85.07 | 0.0174 | 81.59 | 92.02 | 87.80 | 89.86 | 0.0140 | 80.98 | 88.31 | 90.70 | 89.49 | 0.0145 | 87.91 | 91.42 | 95.82 | 93.57 | 0.0130 |
| Dice | 74.38 | 86.40 | 84.25 | 85.31 | 0.0174 | 81.75 | 91.55 | 88.42 | 89.96 | 0.0140 | 81.37 | 88.83 | 90.65 | 89.73 | 0.0145 | 88.28 | 92.91 | 94.66 | 93.78 | 0.0130 |
| Dense Supervision | ||||||||||||||||||||
| Dense | 77.93 | 86.30 | 88.94 | 87.60 | 100.0 | 84.07 | 89.95 | 92.79 | 91.35 | 100.0 | 85.71 | 92.08 | 92.53 | 92.30 | 100.0 | 90.80 | 93.52 | 96.90 | 95.18 | 100.0 |
| Setting | Macro Metrics (%) | IoU per Class (%) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| mIoU | mPrecision | mRecall | mF1 | Effort | Background | Car | Road | Building | Perm. Area | |
| iSAGE (EWDL) | ||||||||||
| Iter 0 | 58.38 | 71.63 | 77.08 | 72.97 | 0.0068 | 38.46 | 47.08 | 68.90 | 59.71 | 77.74 |
| Iter 1 | 67.89 | 77.91 | 83.13 | 79.81 | 0.0135 | 45.64 | 56.30 | 78.18 | 73.08 | 86.23 |
| Iter 2 | 72.05 | 82.34 | 83.41 | 82.86 | 0.0202 | 47.45 | 67.36 | 79.72 | 78.05 | 87.69 |
| Iter 3 | 73.19 | 83.04 | 84.79 | 83.80 | 0.0270 | 47.89 | 68.03 | 81.37 | 80.52 | 88.13 |
| Iter 4 | 74.14 | 83.41 | 85.64 | 84.36 | 0.0337 | 49.17 | 69.86 | 81.81 | 81.48 | 88.38 |
| Iter 5 | 74.79 | 84.07 | 85.72 | 84.83 | 0.0404 | 50.27 | 70.72 | 82.36 | 81.89 | 88.72 |
| Iterative Random Selection | ||||||||||
| Iter 0 | 56.87 | 69.67 | 75.17 | 71.31 | 0.0068 | 34.39 | 47.01 | 67.74 | 58.70 | 76.50 |
| Iter 1 | 64.73 | 75.99 | 81.27 | 77.56 | 0.0136 | 41.65 | 53.89 | 74.51 | 71.68 | 81.95 |
| Iter 2 | 66.67 | 77.49 | 82.35 | 79.04 | 0.0205 | 43.23 | 57.50 | 75.31 | 73.76 | 83.55 |
| Iter 3 | 67.63 | 78.16 | 82.69 | 79.75 | 0.0273 | 43.74 | 59.21 | 76.06 | 75.07 | 84.08 |
| Iter 4 | 68.41 | 78.77 | 83.25 | 80.34 | 0.0342 | 44.81 | 60.15 | 76.72 | 75.55 | 84.84 |
| Iter 5 | 69.11 | 79.39 | 83.25 | 80.83 | 0.0410 | 45.02 | 61.73 | 77.08 | 76.33 | 85.36 |
| Dense Supervision | ||||||||||
| Dense | 77.17 | 87.13 | 86.13 | 86.45 | 100.0 | 53.63 | 73.09 | 83.76 | 84.71 | 90.70 |
| Alternative Loss Functions (Final Iteration) | ||||||||||
| Dice | 73.37 | 82.78 | 85.12 | 83.81 | 0.0404 | 48.15 | 68.66 | 81.63 | 80.28 | 88.15 |
| Cross-Entropy | 74.11 | 83.35 | 85.53 | 84.32 | 0.0404 | 48.94 | 69.52 | 82.29 | 81.38 | 88.43 |
| Focal | 74.66 | 83.85 | 85.82 | 84.74 | 0.0404 | 50.26 | 70.24 | 82.07 | 82.04 | 88.69 |
| mIoU | Car | Road | Building | Perm. | BG | |
|---|---|---|---|---|---|---|
| 1 (=standard Dice) | 73.37 | 68.66 | 81.63 | 80.28 | 88.15 | 48.15 |
| 2 | 74.63 | 70.38 | 82.23 | 82.03 | 88.53 | 49.97 |
| 5 (default) | 74.79 | 70.72 | 82.36 | 81.89 | 88.72 | 50.27 |
| 10 | 74.66 | 71.29 | 82.51 | 81.40 | 88.60 | 49.50 |
| 20 | 72.18 | 71.68 | 82.51 | 73.84 | 88.55 | 44.35 |
| Architecture | Encoder | Params | GMACs | iSAGE (%) | Dense | iSAGE/ | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| mIoU | Car | Road | Build. | Perm. | BG | mIoU | Dense | ||||
| U-Net | Eff-B7 | 67.1 | 3.2 | 74.79 | 70.72 | 82.36 | 81.89 | 88.72 | 50.27 | 77.17 | 96.9% |
| U-Net | R101 | 51.5 | 15.6 | 72.32 | 67.95 | 81.52 | 77.12 | 88.01 | 46.99 | 76.63 | 94.4% |
| DLV3+ | R50 | 26.7 | 9.2 | 70.69 | 60.94 | 79.66 | 79.59 | 86.97 | 46.32 | 75.14 | 94.1% |
| SegFormer | MiT-B2 | 24.7 | 5.3 | 71.86 | 61.30 | 80.21 | 81.30 | 88.28 | 48.20 | 75.11 | 95.7% |
| Method | Cost | mIoU | Imperv. | Build. | Tree | Car | Low Veg. |
|---|---|---|---|---|---|---|---|
| Unsupervised Domain Adaptation (no target labels) | |||||||
| No Adaptation | 0% | 30.19 | 35.24 | 42.23 | 43.57 | 1.39 | 28.50 |
| ADVENT [96] | 0% | 40.84 | 53.82 | 56.59 | 49.21 | 24.46 | 20.12 |
| CLAN [97] | 0% | 46.40 | 63.38 | 61.78 | 55.03 | 28.98 | 22.83 |
| Active Domain Adaptation | |||||||
| MADAv2 [98] | 1% | 50.62 | 65.73 | 59.94 | 54.26 | 21.57 | 51.58 |
| RIPU [89] | 0.015% | 69.91 | 80.09 | 86.17 | 65.02 | 53.02 | 65.26 |
| ILM-ASSL [91] | 1% | 70.42 | 81.03 | 87.28 | 64.36 | 57.13 | 62.32 |
| D2ADA [90] | 1% | 71.75 | 80.71 | 87.62 | 66.69 | 57.58 | 66.17 |
| EasySeg [29] | 0.015% | 72.83 | 81.62 | 88.40 | 69.16 | 57.90 | 67.25 |
| Supervised Learning | |||||||
| Fully Supervised | 100% | 76.93 | 85.62 | 89.16 | 71.05 | 68.99 | 69.85 |
| iSAGE | 0.011% | 76.65 | 83.93 | 88.23 | 71.55 | 70.02 | 69.51 |
| Class | Share (%) | Polygons | Vertices ( = 1) | Vertices ( = 2) | Clicks | Ratio ( = 2) |
|---|---|---|---|---|---|---|
| Impervious | 31.4 | 7302 | 209,550 | 123,011 | 5981 | 20.6× |
| Building | 28.2 | 4560 | 72,521 | 37,618 | 5925 | 6.3× |
| Tree | 21.6 | 7131 | 181,470 | 112,534 | 5932 | 19.0× |
| Car | 1.5 | 5074 | 46,396 | 31,074 | 5574 | 5.6× |
| Low vegetation | 17.3 | 8310 | 176,134 | 108,575 | 5640 | 19.3× |
| Total | 100.0 | 32,377 | 686,071 | 412,812 | 29,052 | 14.2× |
| Method | Pixels Labeled (%) | Iter-5 mIoU (%) | Gap to iSAGE (pp) |
|---|---|---|---|
| iSAGE | 0.011 | 76.65 | 0.00 |
| Uncertainty acquisition (oracle entropy budget sweep) [57,69] | |||
| Oracle entropy 1× | 0.011 | 66.38 | −10.27 |
| Oracle entropy 10× | 0.10 | 66.26 | −10.39 |
| Oracle entropy 50× | 0.48 | 67.01 | −9.64 |
| Oracle entropy 100× | 0.95 | 67.85 | −8.80 |
| Self-training pseudo-labels (confidence threshold sweep) [25,26] | |||
| Pseudo-labeling (0.90) | ∼95 | 69.00 | −7.65 |
| Pseudo-labeling (0.95) | 93.0 | 69.35 | −7.30 |
| Pseudo-labeling (0.99) | 84.8 | 69.34 | −7.31 |
| CRF-based label propagation [95] | |||
| DenseCRF (0.95) | 96.7 | 62.25 | −14.40 |
| Control | |||
| Uniform random | 0.011 | 66.60 | −10.05 |
| Method | Annotation | Iter. | HIL | Acq. | Prop. | Pseudo | Consist. | DA | FM |
|---|---|---|---|---|---|---|---|---|---|
| iSAGE | Sparse Clicks | ✓ | ✓ | ||||||
| Weak/sparse supervision (non-iterative) | |||||||||
| Bearman et al. [19] | Sparse points | ||||||||
| ScribbleSup [34] | Scribbles | ✓ | |||||||
| FESTA [20] | Scribbles | ✓ | |||||||
| Liu et al. [36] | One point | ✓ | |||||||
| PAMSNet [54] | Sparse points | ||||||||
| Iterative weak/self-training/semi-supervised | |||||||||
| Tree Energy [47] | Scribbles | ✓ | ✓ | ✓ | |||||
| ScribFormer [99] | Scribbles | ✓ | ✓ | ✓ | |||||
| PyMIC [48] | Points/Scribbles | ✓ | ✓ | ✓ | |||||
| UniMatch V2 [100] | Unlabeled + few | ✓ | ✓ | ✓ | |||||
| PENet [55] | Sparse points | ✓ | ✓ | ✓ | self | ||||
| AESAM [56] | Sparse points | ✓ | ✓ | ✓ | self | ||||
| Active learning/interactive (single-domain) | |||||||||
| Desai and Ghose [51] | Points | ✓ | ✓ | ||||||
| DIAL [27] | Clicks | ✓ | ✓ | ✓ | ✓ | ||||
| HAL-IA [28] | Sparse points | ✓ | ✓ | ✓ | ✓ | ||||
| ViewAL [67] | 3D points | ✓ | ✓ | ||||||
| S4AL [62] | Points | ✓ | ✓ | ✓ | |||||
| ESA [65] | Clicks | ✓ | ✓ | ✓ | ✓ | ||||
| Think Twice [64] | Points | ✓ | ✓ | ||||||
| Cold AL [31] | Points | ✓ | ✓ | ||||||
| Teng and Wang [33] | Points | ✓ | ✓ | ||||||
| Adapt. Superpixel AL [101] | Superpixels | ✓ | ✓ | ✓ | |||||
| BalEntAcq [66] | Sparse pixels | ✓ | ✓ | ||||||
| ALC [80] | Clicks | ✓ | ✓ | ✓ | ✓ | self | |||
| A2LC [81] | Clicks | ✓ | ✓ | ✓ | ✓ | self | |||
| Active domain adaptation | |||||||||
| RIPU [89] | Regions | ✓ | ✓ | ✓ | ✓ | ||||
| D2ADA [90] | Regions | ✓ | ✓ | ✓ | |||||
| ILM-ASSL [91] | Sparse points | ✓ | ✓ | ✓ | ✓ | ||||
| EasySeg [29] | Sparse points | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| Few-shot/zero-shot/foundation-model | |||||||||
| PANet [102] | Support masks | ||||||||
| RePRI [18] | Support masks | ||||||||
| SAM [11] | Prompts/clicks | self | |||||||
| CLIPSeg [12] | Text/image | self | |||||||
| PointSAM [82] | Points/prompts | ✓ | ✓ | self | |||||
| SEEM [103] | Prompts/text | ✓ | self | ||||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Carvalho, O.L.F.d.; de Carvalho Júnior, O.A.; de Albuquerque, A.O.; Silva, D.G.e. iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision. Remote Sens. 2026, 18, 2973. https://doi.org/10.3390/rs18172973
Carvalho OLFd, de Carvalho Júnior OA, de Albuquerque AO, Silva DGe. iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision. Remote Sensing. 2026; 18(17):2973. https://doi.org/10.3390/rs18172973
Chicago/Turabian StyleCarvalho, Osmar Luiz Ferreira de, Osmar Abílio de Carvalho Júnior, Anesmar Olino de Albuquerque, and Daniel Guerreiro e Silva. 2026. "iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision" Remote Sensing 18, no. 17: 2973. https://doi.org/10.3390/rs18172973
APA StyleCarvalho, O. L. F. d., de Carvalho Júnior, O. A., de Albuquerque, A. O., & Silva, D. G. e. (2026). iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision. Remote Sensing, 18(17), 2973. https://doi.org/10.3390/rs18172973

