Next Article in Journal
Stationary State Recognition of a Mobile Platform Based on 6DoF MEMS Inertial Measurement Unit
Next Article in Special Issue
Deep Learning-Based Prediction of Tumor Mutational Burden from Digital Pathology Slides: A Comprehensive Review
Previous Article in Journal
Research on TID Controller Design for Fractional-Order Time-Delay Systems
Previous Article in Special Issue
Machine Learning Methods for Predicting Cancer Complications Using Smartphone Sensor Data: A Prospective Study
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Artificial Intelligence in Medical Diagnostics: Foundations, Clinical Applications, and Future Directions

by
Dorota Bartusik-Aebisher
1,
Daniel Roshan Justin Raj
2 and
David Aebisher
3,*
1
Department of Biochemistry and General Chemistry, Faculty of Medicine, Collegium Medicum, University of Rzeszów, 35-310 Rzeszów, Poland
2
English Division Science Club, Faculty of Medicine, Collegium Medicum, University of Rzeszów, 35-310 Rzeszów, Poland
3
Department of Photomedicine and Physical Chemistry, Faculty of Medicine, Collegium Medicum, University of Rzeszów, 35-310 Rzeszów, Poland
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(2), 728; https://doi.org/10.3390/app16020728
Submission received: 7 December 2025 / Revised: 5 January 2026 / Accepted: 6 January 2026 / Published: 10 January 2026

Abstract

Artificial intelligence (AI) is rapidly transforming medical diagnostics by allowing for early, accurate, and data-driven clinical decision-making. This review provides an overview of how machine learning (ML), deep learning, and emerging multimodal foundation models have been used in diagnostic procedures across imaging, pathology, molecular analysis, physiological monitoring, and electronic health record (EHR)-integrated decision-support systems. We have discussed the basic computational foundations of supervised, unsupervised, and reinforcement learning and have also shown the importance of data curation, validation metrics, interpretability methods, and feature engineering. The use of AI in many different applications has shown that it can find abnormalities and integrate some features from multi-omics and imaging, which has shown improvements in prognostic modeling. However, concerns about data heterogeneity, model drift, bias, and strict regulatory guidelines still remain and are yet to be addressed in this field. Looking forward, future advancements in federated learning, generative AI, and low-resource diagnostics will pave the way for adaptable and globally accessible AI-assisted diagnostics.

Graphical Abstract

1. Introduction

Artificial intelligence (AI) is an umbrella term that refers to the machine-based systems that, when given a certain set of goals set by humans, can make predictions, suggestions, or choices and perform tasks that would normally require the intelligence of humans. AI in healthcare includes systems that can assist with or automate clinical tasks such as diagnosis, prognosis, workflow optimization, recommendation of treatments, and patient monitoring [1]. Machine learning (ML), on the other hand, is a well-defined subfield of AI where algorithms, from data examples, learn patterns and can make predictions or decisions on new data based on the patterns learned. ML involves supervised, unsupervised, and reinforcement learning, while deep learning, which involves neural networks with many layers, is a commonly used ML approach in modern medical imaging and electronic health record (EHR) analysis [2].
The early implementation of AI in medicine relied on expert systems that were using predefined “if-then” rules to simulate clinical reasoning. A good example is the MYCIN program from the 1970s for infectious diseases, showing that computers were able to provide explainable medical recommendations. However, rule-based systems were limited at the time by the difficulty of manual knowledge encoding and poor adaptability [3], which resulted in the MYCIN program never being used in practice. As clinical data increased between the 1980s and 1990s, AI also shifted towards probabilistic reasoning and statistical learning, including the Bayesian network and pattern-recognition methods. These systems were learning from data rather than just fixed rules, which marked a transition towards machine learning in diagnostics [4]. Machine learning-based computer-aided detection (CAD) systems emerged in radiology between the 1990s and 2000s, especially in mammography and chest imaging, where algorithms were analyzing medical images to highlight any abnormalities for radiologists, which showed the first clinical integration of AI tools [5]. Algorithms such as support vector machines, random forests, and logistic regression became the standard for diagnostic prediction and biomarker discovery. These classical ML models from between the 2000s and 2010s relied on engineered features and were being used in radiomics, genomics, and supporting clinical decisions [6]. The deep learning revolution between the 2010s and now, with the rise of deep neural networks (DNNs), has supported AI in medicine by enabling end-to-end learning from raw data. Convolutional neural networks (CNNs) also achieved specialist-level accuracy in diagnostics that were image-based, such as in skin cancer and retinal disease detection [7]. The focus on AI research today is mostly on its clinical validation, generalizability, fairness, and explainability, as all regulatory pathways require precise evaluation before it can be deployed in patient care [8].
AI tools are able to reduce diagnostic errors through detecting subtle findings and minimizing possible oversight, mainly during imaging triage and interpretation. However, issues such as algorithmic bias and over-reliance, such as automation bias, could produce new errors if the AI models are poorly validated or if they were trained on non-representative data [9]. Clinical trials have shown that AI is able to improve diagnostic accuracy and workflow efficiency, but there is limited evidence that shows consistent benefits in hard outcomes such as survival and morbidity, as most studies have reported process improvements rather than its direct impact on patients [10]. As such, AI systems face key barriers like the limited generalizability across populations, lack of explainability, issues with data privacy, and insufficient regulatory frameworks. The integration of AI into clinical workflows in a safe way will require transparent validation and prospective trials [11]. The use of AI in healthcare also raises questions about accountability, transparency, and equity, which further hinder its adoption. Regulatory bodies such as the FDA and EMA also require robust clinical validation, bias assessment, and post-market surveillance for AI diagnostic tools before they allow for their clinical adoption [12].
This review addresses the main research question of how modern-day artificial intelligence and machine learning can be systematically integrated into medical diagnostics across fields such as imaging, molecular and omics data, physiological signals, and clinical records, as well as the factors that currently limit their safe and effective adoption. The motivation behind this work comes from the rapid expansion of AI-driven diagnostic tools and the challenges surrounding generalizability, interpretability, clinical validation, and implementation in the real world. While there are numerous studies that have reviewed AI applications with individual diagnostic domains, there remains a limited number of comprehensive and cross-domain syntheses that combine computational foundations, clinical performance, and translational barriers. This review aims to address this gap by giving a perspective that includes methodological principles with evidence from diverse diagnostic modalities [13]. The contribution of this review is its joint discussion of learning paradigms, model architectures, validation and interpretability frameworks, and regulatory considerations, while also highlighting the emerging trends such as federated learning, generative AI, and multimodal foundation models. The remainder or the review is organized as follows: Section 2 outlines the theoretical and computational foundations of AI in diagnostics; Section 3 reviews the main AI applications across major diagnostic procedures; Section 4 examines predictive analytics and clinical decision-support systems; Section 5 discusses the challenges involved with implementation; Section 6 explores the field’s future directions; and Section 7 concludes the review. To help navigate the diverse literature included in this review, Table 1 summarizes the key representative studies, which are organized by diagnostic domain, data modality, and their clinical contribution.
Figure 1 shows a PRISMA workflow diagram showing the methodology for selecting the articles included in this review.
The methods used for reviewing the literature involved a search of the PubMed and PubMed Central (PMC) databases, which was conducted in accordance with PRISMA guidelines, as seen in Figure 1. Peer-reviewed articles published up to 2025 were searched using keywords related to artificial intelligence, machine learning, and medical diagnostics. After the removal of duplicates and the automated exclusion of non-medical AI and retracted publications, titles and abstracts were screened for relevance to the topic. Full-text articles were assessed using a predefined inclusion criterion that required a clear description of AI methodology, relevance to clinical or diagnostic applications, and the presence of an empirical evaluation or validation. Studies that were not validated or did not have clinical relevance were excluded from the review. A total of 171 studies met the set eligibility criteria and were included in the final review.

2. Theoretical and Computational Foundations

2.1. Overview of Main Learning Paradigms

Artificial intelligence (AI) in medical diagnostics has three main learning paradigms, which are supervised, unsupervised, and reinforcement learning, as seen in Table 2. Supervised learning is based on labeled datasets and supports most diagnostic models in healthcare, such as in CNNs for radiology and pathology and in algorithms used in ECGs or lab-based detections [2,13,15]. Unsupervised learning is able to reveal hidden data structures for the stratification of patients, anomaly detection, and the discovery of biomarkers using clustering or autoencoders [69,70]. Reinforcement learning, on the other hand, uses sequential decision-making, which is then applied to adaptive imaging and diagnostic workflows [71]. On a computational level, all these methods share their foundations in probabilistic interference, gradient optimization, and representation learning [72]. These methods do have their drawbacks, such as the limited interpretability, label bias, and gaps in validations [19]. Up-and-coming hybrid and self-supervised frameworks have integrated large, unlabeled biomedical data to improve the performance of diagnostics and generalization, which form the theoretical foundation of data-driven, clinical AI [73].
Deep learning architectures such as conventional neural networks (CNNs), recurrent neural networks (RNNs), and transformers are very important in medical imaging and signal analysis in AI-assisted diagnostics. CNNs have been shown to be highly effective for spatial feature extractions and are now widely used in radiology for the detection of tumors, organ segmentation, and in the classification of diseases [17]. For example, the U-Net system and its associated variants remain as strong foundations in biomedical image segmentation in computerized tomography (CT), magnetic resonance imaging (MRI) and in histopathology [78], while systems such as CheXNet have demonstrated CNN-based detection of pneumonia in chest X-rays [16]. Frameworks such as the nnU-Net also further standardized medical segmentation processes [79]. RNNs, on the other hand, in particular, long short-term memory (LSTM) networks, are proven to be well suited for physiological time-series signals, such as in electrocardiograms (ECG) and electroencephalograms (EEG). For example, arrhythmia detection with performance comparable to domain experts on specific benchmark datasets could be achieved by deep RNN-based classifiers [31], and other similar architectures have been able to support seizure prediction from EEGs [35]. Transformers have also recently advanced in multimodal and long-range dependency modeling. Some hybrid architectures, such as TransUNet, improved segmentation by the integration of self-attention and CNN features [80], while other transformer-based ECG models allowed for the strong classification of cardiac pathology [32].

2.2. Model Reliability, Interpretability, and Data Foundations in AI Diagnostics

The receiver operating characteristic (ROC) curve, area under the curve (AUC), and precision–recall analyses are algorithmic validation metrics that are crucial in assessing the various AI-assisted diagnostic systems in medicine. The ROC curve is a metric that can evaluate a particular model’s trade-offs between sensitivity and specificity, while AUC can measure its overall ability to discriminate, where showing values that are closer to 1.0 will mean stronger performance. Precision–recall curves give important insight by putting importance on positive predictive value and recalling all possibilities. Figure 2 shows how diagnostic performance in AI systems comes from the interaction between data curation, learning paradigms, validation metrics, and interpretability mechanisms, rather than just from the algorithm alone. By positioning commonly reported metrics such as the ROC curves and AUC downstream of these upstream design decisions, the figure shows why high performance in discrimination does not mean clinical reliability or utility. The figure further emphasizes interpretability and uncertainty estimation as important mediators between model output and clinical decision-making. Works conducted recently have demonstrated the importance of these metrics in many different applications. For example, Li et al. 2025 developed an ML-based model for Staphylococcus aureus bacteremia recurrence, which was also validated with the ROC-AUC metrics to ensure its clinical reliability [81]. Similarly, in a study by Urundai Meeran et al. 2025, AUC and precision–recall curves were used to validate a ResNet-50-based renal malignancy prediction framework [82]. Neuwieser et al. 2025 also highlighted the interpretability and metric-based validation in the classification of CNN-based ulcer images [83]. All of these studies have highlighted the importance of balanced metric evaluation in order to ensure the diagnostic reliability of AI-driven medicine. While the ROC and AUC metrics show a model’s capacity to discriminate, they do not guarantee its clinical utility. Only a few studies have reported model calibration, evaluated clinically meaningful decision thresholds, or assessed the results of false positives and false negatives. A model with good AUC metrics can still produce poorly calibrated probabilities, which hinders its potential in risk stratification or treatment decisions. Without an alignment to clinical procedures and decision thresholds, performance metrics alone offer limited guidance for implementation in the real world.
Explainable artificial intelligence (XAI) is a key part of AI-assisted medical diagnostics because non-transparent, “black-box” models can be an obstruction to the trust of clinicians, informed consent, and regulatory approval, which are factors that slow down AI’s clinical adoption [56]. Intrinsically interpretable models (rules, linear models) and post hoc techniques such as saliency/heat maps, Shapley additive explanation (SHAP), and local interpretable model-agnostic explanations (LIME) are common XAI strategies that assign model outputs to inputs [57]. However, various interpretability issues remain, such as the instability or sensitivity to small input changes in many post hoc explanations, the fact that they can be misleading about causal mechanisms, and how they are unable to answer clinicians’ specific “why” or “what if” questions [84]. Integration into practical use is also challenging, as the explanations given should be tailored to each individual user, be validated clinically, and also be embedded into procedures and interfaces rather than be given as just technical plots [59]. To progress forward, the field requires standardized evaluation protocols, clinician-in-the-loop studies, and regulatory frameworks that would demand the demonstration of explainable fidelity and its clinical benefits before deployment. The combination of built-in interpretability with human-centered design and strict clinical validation will set the foundation for the growth of reliable AI diagnostics [58].
Clinical data have often been heterogeneous, noisy, and incomplete, and thus, data curation, preprocessing, and feature engineering are essential for reliable AI-assisted medical diagnostics. The removal of duplicates, correction of errors, and balancing of imaging and laboratory formats are ensured by proper data curation, which will help reduce possible bias and improve model reproducibility [14]. Broad, well-curated datasets such as the MIMIC-III show how standardized clinical records allow for algorithmic development [42]. Preprocessing techniques such as normalization, the removal of artifacts in imaging, handling of missing values in electronic health records (EHRs), etc., are necessary to avoid distorted patterns and improve model stability [18]. In medical imaging, preprocessing techniques such as intensity correction and segmentation have been shown to improve the diagnostic performance of deep learning tasks [7]. Feature engineering can also translate raw clinical signals into meaningful representations. For example, when applied in radiomics, it brings out quantitative tumor phenotypes that can enhance predictive accuracy in the field of cancer diagnostics [54,55]. When used in EHR, interpretable feature selection supports risk prediction models that are clinically transparent [38]. Effective data curation and engineered features ultimately strengthen generalizability, reduce overfitting, and support the trustworthy deployment in clinical practice [73].

3. AI Application Across Diagnostic Procedures

3.1. Imaging Diagnostics

AI-assisted image analysis in radiology, like the deep learning-convolutional neural networks, has been applied across CT, MRI, X-ray, and ultrasound to improve lesion detection, image quality, and procedure efficiency, which can be seen in Table 3. In a 2021 review, for example, it was found that AI in ultrasound can allow for automatic anatomy localization, segmentation, and lesion detection across the abdomen and organs such as the thyroid and breast [26]. Similarly, another review on radiology in 2023 highlighted the shift in AI from task-specific models to larger “foundation” models that could handle multimodal imaging such as X-ray, CT, MRI, and positive electron tomography (PET) and also support segmentation or classification tasks [20,21]. Another meta-analysis found that the collaboration between humans and AI reduced reading times by around 27% and maintained a sensitivity at around 1.12 times human alone [85]. In pathological fields, digitized whole-slide imaging (WSI) in combination with AI has allowed for the automated classification of tissues, mitotic figure detection, and prognostic feature extraction. A review from 2023 highlighted that FDA-approved whole-slide imaging (WSI) that was used with AI for primary diagnosis signified a major milestone in the integration of computational pathology into clinical procedures [23]. A systematic review and meta-analysis from 2024 that covered over 150,000 WSIs reported a combined mean sensitivity of 96.3% and specificity of 93.3% for AI in digital pathology, even though it was with significant heterogeneity and bias risk [24]. Although many imaging studies and meta-analyses have reported a high diagnostic accuracy, a direct comparison across models and modalities is difficult. The reported performance differences are usually due to variations in study design, including single-center vs. multicenter cohorts, disease prevalence, reference standards, and image acquisition protocols. Models that are trained and evaluated on homogenous retrospective datasets often show inflated performance when compared with models that were validated across institutions or devices. As a result, metrics such as sensitivity and specificity should be interpreted as being due to these methodological differences, rather than as indicators of clinical superiority. In ophthalmology-related fields, retinal imaging such as fundus photography and optical coherence tomography (OCT) has had rapid AI development for the screening and monitoring of diseases such as diabetic retinopathy (DR), age-related macular degeneration (AMD), retinal vein occlusion, and retinopathy of prematurity (ROP). A paper from 2024 talks about the AI models that were applied to these diseases and notes elevated screening outputs and the potential to reduce disparities [28]. Additionally, another review on chorioretinal pathology using AI highlighted the segmentation, classification, and prediction tasks of assessing retinal disease progression [29]. These studies together show that AI is becoming embedded in medical diagnostic methods and, thereby, increasing speed and improving sensitivity and reproducibility. However, a strong focus must be put on its data quality, interpretability, external validation, and its integration into existing procedures.

3.2. Molecular and Omics Diagnostics

AI has been transforming molecular and omics diagnostics by converting high-dimensional biological data into clinically applicable signals for diagnosis, prognosis, and therapy. In genomics, for example, ML and deep learning have accelerated variant calling, pathogenicity prediction, and genotype–phenotype mapping, which has improved the sensitivity of variant detection and helps with the interpretation of variants of uncertain significance. Multi-omics integration, such as in genomics, transcriptomics, proteomics, and metabolomics, uses ML to find coordinated signatures that single-omic analysis may miss. Methods include regularized regression, tree ensembles, graph-based models, and deep generative models like variational autoencoders (VAEs) for dimensionality reduction, imputation, and batch correction, which allows for better patient stratification and biomarker panels [53]. In metabolomic fields, AI handles complex, correlated metabolite profiles to detect disease-related patterns, build predictive diagnostic models, and place importance on candidate biomarkers. Recent reviews have shown that AI procedures improve feature selection, handle any missingness, and aid in pathway-level interpretation in metabolic diseases and cancer applications [50]. Proteomics benefits from ML through mass spectrometry signal interpretation, peptide identification, and the prediction of peptide properties from sequence. Deep models have now been able to enhance peptide-spectrum matching, reduce false identifications, and allow for quantitative biomarker discovery procedures that establish the connection between proteomic patterns and clinical endpoints [48,49]. Across different omics, AI facilitates biomarker discovery by attaining candidate features, building multimarker classifiers, and evaluating the generalizability through cross-validation and external cohorts. However, challenges such as cohort heterogeneity, small sample sizes against feature counts, batch effects, and the need for interpretable models and meticulous external validation remain a hindrance to its clinical adoption [52]. Figure 3 summarizes the role of artificial intelligence across various diagnostic modalities by showing how imaging, molecular omics, physiological signals, and clinical records join through analytical frameworks. Rather than just depicting these domains as standalone pipelines, the figure emphasizes the cross-domain interaction as a key part in improving diagnostic sensitivity and prognostic performance. The figure also supports the argument that the recent advancements in medical AI have arisen from the fusion of complementary data types rather than modality-specific algorithmic optimization alone. Overall, AI in molecular diagnostics improves sensitivity, speed, and integrative understanding, which moves the field towards precision medicine while also highlighting the importance of careful study design, strong validation, interpretability, and data management before clinical translation [47].

3.3. Physiological and Clinical Data

AI plays a huge role in analyzing physiological and clinical data, supporting early diagnosis and the continuous monitoring of patient health. In ECG analysis, deep learning models have been able to detect arrhythmias, ischemia, conduction abnormalities, and cardiomyopathies with high levels of accuracy. Hannun et al. developed a deep neural network that demonstrated arrhythmia classification performance comparable to cardiologists in controlled retrospective evaluations of single-lead ECG recordings [31]. In another similar example, Ribeiro et al. demonstrated the deep learning-based interpretation of 12-lead ECGs in large clinical datasets, which improved the prediction of atrial fibrillation risks and structural cardiac abnormalities [34]. For real-time monitoring, the continuous ECG stream analysis allows for the early detection of cardiac deterioration in intensive care settings [43]. In EEGs, AI has supported seizure detection, classification of sleep stages, and neurological disease diagnosis. For example, Roy et al. conducted a systematic review showing that deep learning models improved EEG-based seizure detection and neurological classification accuracy in comparison to traditional feature-engineered approaches [36]. In another study, Hussein et al. applied CNNs to EEGs for early Alzheimer’s disease detection by capturing the subtle spectral changes [37]. AI has also enhanced the monitoring of vital signs, which allows for the identification of deterioration patterns before the patient’s clinical symptoms become obvious. For example, machine learning in early warning systems (EWSs) has the ability to predict sepsis, respiratory failure, and ICU transfer by analyzing the patient’s heart rate, blood pressure, oxygen saturation, and respiratory rate all together. Mao et al. showed that sepsis prediction improved when using ML-based EHR models [44], and in another study, Kwon et al. reported the improvement of in-hospital cardiac arrest prediction when using deep recurrent networks [45]. The integration of physiological data with EHRs has also shown a great increase in diagnostic power. Natural language processing (NLP) allows for the automated extraction of clinical concepts from physician notes. Devlin et al. [39] made a transformer-based model called BERT that kickstarted medical NLP models and was later adapted for clinical text as ClinicalBERT [40]. NLP models so far have been successfully applied to clinical tasks such as phenotype extraction, diagnosis coding, and outcome prediction. For example, nurse and physician note content has been shown to predict the risk of heart failure rehospitalization and to automate decision-making in large health system datasets [87]. In general, multimodal AI systems that combine ECG, EEG, vital signs, structured EHR fields, and clinical text have been shown to allow for better proactive monitoring, earlier diagnosis, and reduced burden on clinicians.

3.4. AI-Driven Predictive and Prognostic Modeling

The application of artificial intelligence in predictive and prognostic modeling has been quickly advancing, particularly when multimodal inputs such as imaging, omics, text, and clinical data are integrated to predict outcomes. In cardiovascular disease (CVD), for example, a systematic review found that AI survival models such as Random Survival Forests and DeepSurv outperformed traditional models in predefined benchmark tests for time-to-event prediction [88]. Imaging–omics integration, for example, in radiogenomics, allows for the prediction of recurrence, progression-free survival, and treatment response in oncological fields [89]. For instance, AI-powered radiogenomic procedures have combined CT and MRI image features with genomic profiles to categorize risks and guide cancer therapy [90]. In vascular medicine, AI models have been used to predict results in aortic aneurysm, peripheral arterial disease, and lower-limb diseases [91,92]. In non-small cell lung cancer (NSCLC), by combining CT and PET with clinical data, multimodal AI predicted postsurgical recurrence risk [93]. On a broader level, AI models in oncology use digital pathology, molecular omics, and imaging to predict outcomes and also provide personalized treatment suggestions [94]. A key part of prognostic modeling is the alignment of multiple data streams, such as imaging features capturing the anatomical or functional phenotype, omics capturing the molecular drivers, and clinical or EHR data providing context regarding the patient. This integrated method provides a higher predictive accuracy and clinical relevance. However, key challenges remain; for example, in the study by Teshale et al., it was noted that only a few studies included the social determinants of health and gender categorization [88]. Overall, AI-assisted predictive and prognostic models using multimodal data enable a transformative step towards diagnostics and patient management, ultimately allowing for better risk management, early intervention, and personalized therapy planning.

4. Predictive Analytics and Clinical Decision-Support Systems

4.1. Predictive Analysis

Predictive analysis in AI-assisted medical diagnostics helps in detecting health deterioration early and predicts disease trajectories such as sepsis, cardiac arrest, and cancer recurrence by learning patterns from abundant clinical time-series, imaging, and molecular data, as shown in Table 4. So far, several studies have been conducted on this, and they show its feasibility and risks. For example, in sepsis, it has been shown that interpretable ICU models using streaming vitals and labs are able to predict the onset of disease even hours before its clinical recognition [95]. Early work performed in this field showed that even a minimal amount of EHR input allows for the prediction of sepsis [96]. Systematic reviews and meta-analyses have confirmed that ML models do outperform traditional methods in defined diagnostic subtasks, but have also highlighted widespread heterogeneity and reporting bias [97]. Large multisite validations and real-world deployments such as COMPOSER show the possible improvements in outcomes but also give insight into the challenges involved in the generalizability and implementation in the real world. For cardiac arrest, time-series and ECG-derived heart rate variability (HRV) models such as DeepEWS achieved considerably higher sensitivity than conventional early-warning scores for in-hospital events, thereby allowing for a 0.5–24 h prediction window in ICU and ward backgrounds [98,99]. In oncology, radiomics and radiogenomics combine imaging phenotypes with molecular profiles, which lets them predict recurrence and survival levels. Some foundational works established that quantitative imaging features correlated with gene expression and outcomes [100,101]. Reviews and studies in radiogenomics show methods but also raise questions on reproducibility, pushing for standardized procedures and external validation [102,103]. Across various domains, there are some common limitations, such as label noise and variable case definitions, dataset shifts between sites, limited external validation, and insufficient reporting of calibration and clinical utility. Figure 4 demonstrates how predictive analytics models show clinical impact only when they are embedded within EHR-integrated decision-support systems that include validation, governance, and monitoring mechanisms. By situating AI predictions within clinical workflows, feedback loops, and post-deployment surveillance, the figure clarifies why high predictive accuracy alone is not enough for safe and effective real-world usage. The figure also highlights the important role of human oversight, usability, and continuous performance evaluation in preventing automation bias and silent model failure. Transparent, multi-institutional datasets and prospective trials remain essential in order for predictive analysis to move from a retrospective method to a safe clinical practice.

4.2. Personalized Medicine

Personalized medicine through AI models has been using patient-level data like genomics, imaging, proteomics, and clinical history to tailor diagnosis, predict the response to treatment, and choose the best therapy for patients. AI shines at identifying complex, non-linear patterns throughout high-dimensional omics and imaging features that conventional statistics usually fail to do [109]. Radiogenomic procedures combine imaging phenotypes with tumor genomics to understand molecular subtypes and predict recurrence or drug sensitivity, which improves stratification for targeted therapies [110]. When machine learning is applied to pharmacogenomics, drug-sensitivity datasets can predict which patients will respond to specific chemotherapies or targeted agents and which patients will not; this allows for more precise dosing and reduced toxicity [111]. The multi-omics fusion of genomics, transcriptomics, and proteomics with AI supports biomarker discovery for prognostic and predictive signatures, speeding up companion-diagnostic developments [112]. AI also helps in adaptive clinical decision support by combining EHR phenotypes with molecular data to be able to suggest personalized treatment plans and clinical trial matches [113]. There have been key translational successes as well, such as some AI models that can improve the prediction of postsurgical recurrence and therapy response in lung and breast cancers through the combination of imaging, pathology, and molecular profiles [114]. Questions remain regarding its generalizability and explainability, but overall, AI-assisted personalized medicine is becoming a clinical utility by translating multimodal data into practical, patient-specific treatment strategies.

4.3. Decision-Support Systems in EHR Procedures

Decision-support systems (DSSs) and their integration into EHRs have been shown to be effective due to their ability to provide patient-specific recommendations directly within clinical procedures. HER-embedded DSSs have also demonstrated improved diagnostic accuracy in retrospective evaluations and improved standard care when they were properly used in clinician tasks, as shown in the study of EHR-integrated clinical decision-support (CDS) systems by Solomon et al. [60]. The way that it integrates into existing procedures seamlessly is also a reason for its success, with Patterson et al. showing the importance of its usability and clinician engagement [61]. CDS modules that are enhanced with AI increasingly incorporate ML models to improve early detection and risk management. It has been highlighted that advanced CDS can meaningfully improve diagnostic decision-making when used along with high-quality data and standards-based integration [115]. Clinical studies have summarized and highlighted better diagnostic support but also some variability depending on model validation and alert design [116]. Safety considerations and regulations are crucial for its widespread use. Weissman GE describes the FDA’s perspective on predictive CDS, stressing its transparency, need for human oversight, and risk-based classification [62]. Similarly, Karnik K brings attention to regulatory principles for clinical software, emphasizing reliability and traceability [63]. The effective deployment of DSSs depends on ongoing governance. Studies push for CDS stewardship systems to reduce alert fatigue and also to ensure safe performance over time [117]. Reviews on emerging AI-driven systems indicate the importance of explainability to maintain the trust of clinicians [118]. On the whole, EHR-integrated AI-DSSs remain promising for improving diagnostic quality, provided they remain easily usable, validated, transparent, and clinically relevant.

4.4. Validation, Regulation, and Real-World Performance

Clinical validation, regulatory clearance, and post-deployment learning are essential for ensuring that predictive analytics and AI-assisted clinical decision-support systems are safe and clinically beneficial. Clinical validation requires the demonstration of an AI model maintaining its performance across diverse populations, environments, and times [119]. Regulators, particularly the FDA, evaluate AI and ML software as a medical device (SaMD) within already-existing device frameworks while encouraging “good machine learning practice” and pre-specified algorithm change protocols for adaptive models. Evidence packages for regulatory clearance point at clinical utility, patient-centered outcomes, and transparent reporting. However, reviews have shown inconsistent demographic reporting and gaps in safety data across devices that have been approved [120]. Practical deployments so far benefit from staged rollouts, prospective monitoring, and integration with clinical procedures to measure impact on the real world and any harmful side effects [121]. For continuously developing learning systems, regulators and experts have recommended predefined algorithm change protocols, dynamic clinical trial designs, and the constant tracking of performance to allow for safe updates while also maintaining accountability [122,123]. Post-deployment monitoring must be able to detect any dataset shifts and performance changes, such as in label-agnostic shift detection and retraining triggers, with processes for rapid remediation and traceability of data and model versions [124]. Recent case studies and implementation reports have described the best practices, such as cross-site external validation, real-time monitoring dashboards, and clinician-in-the-loop escalation, pathways that are vital to translating the promising potential algorithms into safe and fair clinical care [125]. Real-world deployments have revealed failure modes that were not noticed during retrospective validation. Low performance due to dataset shift, changes in clinical documentation methods, and differences in patient populations have been observed in institutions. In some cases, high alert frequency has led to clinician alert fatigue, and overreliance on algorithmic output has led to automation bias. These failures show that the accuracy of the model alone is insufficient and that continuous monitoring, recalibration, and clinician oversight will be essential to maintain safety and effectiveness after its deployment.

4.5. Case Studies

Case studies from real-world applications show how predictive AI systems can translate into useful diagnostic improvements. A good example is Google DeepMind’s advanced optical coherence tomography (OCT) model for retinal disease, which gave triage recommendations across many different sight-threatening conditions that were comparable to domain experts under controlled evaluation conditions, showing that deep learning groups can generalize to heterogeneous clinical data [30]. The earlier system for diabetic retinopathy screening also delivered high sensitivity and specificity across external datasets, which supports scalable triage where there are limited human graders [126]. Google Health’s multisite breast cancer screening model also reduced both false positives and false negatives when compared with radiologists in retrospective analysis and simulated prospective deployment, which emphasizes the value of well-curated imaging procedures and rigorous cross-site validation [67]. At the Mayo Clinic, the AI-ECG atrial fibrillation predictor showed that a standard sinus-rhythm ECG can find unrevealed physiologic signatures that are detectable only through deep neural networks, thereby allowing for a noninvasive, population-level strategy for early detection [68]. In thoracic imaging fields, deep convolutional systems such as CheXNet-style architectures showed diagnostic accuracy comparable to radiologists on pneumonia detection and a few other chest X-ray procedures, drawing attention to how large datasets and post hoc interpretability can support adoption in high-volume diagnostic procedures [16]. IBM Watson for Oncology, however, showed an instructive contrast, where, although early deployments had shown moderate concordance with multidisciplinary tumor boards, performance varied significantly across institutions. This stresses the importance of sourcing transparent evidence, local customization, and the continuous updating of clinical knowledge bases [127,128]. All of these case studies together show that successful predictive systems rely not only on algorithmic accuracy but also on their integration into procedures and rigorous post-deployment learning frameworks.

5. Innovation and Implementation Challenges

5.1. Data Integrity, Fairness, and Oversight in Medical AI Systems

Substantial innovation and implementation challenges are introduced by AI-assisted medical diagnostics, particularly around data privacy, algorithmic bias, ethical responsibility, and the continuously evolving regulatory frameworks. Clinical AI systems use a large amount of sensitive patient data; however, traditional de-identification is often not enough because high-dimensional medical data, such as in imaging and genomics, has remained re-identifiable even after anonymization processes [65]. There are also concerns over algorithmic bias, as studies have shown that AI models that were trained on skewed datasets may underperform for racial and socio-economic minorities, which can potentially increase healthcare disparities [64]. Bias can stem from issues such as data imbalance, label subjectivity, historical inequalities, or non-representative clinical environments, all of which, call for strict subgroup reporting and continuous monitoring [66]. “Black-box” systems, from an ethical point of view, complicate auditability and informed consent. Explainability failures will diminish trust and make it difficult for clinicians to justify medical decisions to patients [129]. On the regulatory and legal side, many jurisdictions do not yet have tailored frameworks. In the United States, for example, the Health Insurance Portability and Accountability Act (HIPAA) protects health data, but AI systems raise questions of liability about who—clinicians, hospitals, or the developer—would take responsibility if AI happens to make an error [130]. In the European Union, the new AI Act classifies medical AI as “high-risk” and mandates rigorous assessment before market deployment, post-market monitoring, transparency, and human oversight [131]. This must integrate with existing regulations such as the Medical Device Regulation (MDR) and General Data Protection Regulation, which are crucial for privacy and accountability [132]. Healthcare institutions will also have to face operating challenges such as integrating AI into existing procedures, training clinicians well, creating validation methods, and creating AI management committees. Without well-structured oversight, dataset drift and silent failure modes will degrade performance and quality over time [133]. Figure 5 organizes the major challenges in clinical AI into a hierarchical pyramid, showing how issues such as clinician trust, bias amplification, and scalability come from the foundational limitations in data integrity, model reliability, and governance. By structuring these challenges as interdependent layers, the figure shows that ethical and regulatory concerns cannot be addressed without taking technical design and data stewardship into consideration. This representation explains why surface-level interventions, such as an interface redesign or post hoc bias audits, are insufficient without strong foundational controls. The deployment of these AI systems in a responsible way requires a comprehensive strategy that combines technical safeguards, ethical principles, transparent management, and adaptive regulation to ensure safety, fairness, and people’s trust in clinical AI.

5.2. Challenges in System Integration, Model Reliability, and Clinician Trust

Technical challenges surrounding AI in medical diagnostics are mostly related to its interoperability, model drift, and the lack of generalizability, which are further discussed in Table 5. Interoperability is an issue because clinical data are usually fragmented across EHR vendors, imaging modalities, and laboratory systems, and the safe exchange of data and versioning is difficult if there are ad hoc integrations and inconsistent standards, for example, in formats, terminologies, and provenance. Without standard application program interfaces (APIs), clear preprocessing contracts, and proper provenance tracking, models will fail silently when fields are missing or units differ, which will result in increased patient safety risks [1]. Model drift is when input distributions, clinical procedures, or disease prevalence change after the AI model has been deployed, which causes degraded accuracy and calibration. In practice, both sudden shifts, for example, pandemic-driven case mix changes, and gradual covariate drift have been observed. Effective handling of these shifts requires automated drift detection, continuous calibration monitoring, and governance rules for retaining, rollback, or a human review. Several studies have also highlighted the need for practical detection and remediation procedures [86,134]. The lack of generalizability is a result of training on narrow cohorts, single-center, or single-vendor devices. Proactive steps such as stress-testing across external sites, multi-vendor datasets, domain adaptation techniques, and prospective multisite validation will reduce under-specification issues and expose any brittle behavior before clinical use. Even still, many tools have shown variable cross-site performance, so external validation and scrutiny from regulators remain necessary [135]. The collaboration between humans and AI, along with the trust of physicians, depends on usable, transparent interfaces, reliable explanations, and clear decision roles. Clinicians place a high level of trust in systems that provide calibrated probabilities, uncertainty estimates, and audit trails. Clinician training, co-design, and using AI as decision support rather than as an authoritative expert will improve overall acceptance and safety. The governing bodies can take measures such as versioning, logging, and easy override, which will increase accountability [136]. Operationalizing these insights will require standards-based interoperability, for example, Fast Healthcare Interoperability Resources (FHIR); robust monitoring setups for drift detection and calibration; and multidisciplinary governance that will include clinicians, data engineers, and regulators. Recent works show that drift-triggered continual updating combined with clinician-in-the-loop review will preserve performance and help sustain physician trust [137]. Investments must be made in educating clinicians, shared datasets, and transparent post-market surveillance, as it is crucial to have safe, large-scale, and equitable deployment and ongoing multidisciplinary evaluation.
Table 5. Possible failures of clinical AI systems, with their sources, manifestations, severity, detectability, and mitigation. Clinical Severity shows the potential impact of an AI system failure on the safety of patients and clinical outcomes. Severity levels were given based on the likelihood of patient harm, the clinical importance of the affected decision, and the potential for delayed diagnosis. Detectability levels, on the other hand, show how readily a failure can be identified during routine clinical use and were determined by considering the presence of observable performance inconsistencies, the availability of monitoring or alerting systems, and the level of clinician oversight in the workflow.
Table 5. Possible failures of clinical AI systems, with their sources, manifestations, severity, detectability, and mitigation. Clinical Severity shows the potential impact of an AI system failure on the safety of patients and clinical outcomes. Severity levels were given based on the likelihood of patient harm, the clinical importance of the affected decision, and the potential for delayed diagnosis. Detectability levels, on the other hand, show how readily a failure can be identified during routine clinical use and were determined by considering the presence of observable performance inconsistencies, the availability of monitoring or alerting systems, and the level of clinician oversight in the workflow.
Type of FailureSource of IssueObservable ManifestationClinical SeverityDetectabilityExample ScenarioMitigation Strategy
Model drift [138,139]Shift in population, new devices, new diseasesDecreased level of accuracy, uncalibrated outputsHighModerateCOVID-19 disrupting pneumonia modelsContinuous monitoring, periodic retraining
Data source errors [140]Incorrect preprocessing, missing values, HER mapping errorsSilent model failure, unintelligible predictionsCriticalDifficultLab units mismatching (mg/dL vs. mmol/L)Standardized preprocessing, unit normalization
Algorithmic bias [141]Imbalanced datasets, biased annotationWorse outcomes for subgroupsHighModerateLower melanoma detection in darker skin typesSubgroup performance auditing, re-sampling
Adversarial vulnerability [142]Input perturbations, image compressionMisclassification under subtle changesMediumDifficultSlight noise altering CT diagnosisRigorous training, adversarial defenses
Explanation failures (XAI) [143]Unstable saliency mapsInconsistent heatmaps result in the loss of clinician trustMediumEasy Two nearly identical images that produce different saliencyEnsemble explanation, smoothening techniques
Interoperability Failures [144]Non-FHIR EHR formats, missing API supportClinically unable to deployCriticalEasyCDSS not compatible with existing hospital information systems (HIS)FHIR adoption, Health Level Seven (HL7) compliance
Human-AI misalignment [145]Over-trust or Under-trustAutomation bias or alert fatigueHighModerateClinicians overriding or deferring without evaluatingAdequate clinician training, UI co-design, uncertainty estimates

5.3. Academia, Industry, and Healthcare Involvement

Startups have been pushing the speedy translation of AI research into clinically usable diagnostic tools by putting their focus on narrow, clinically tractable use cases, such as in imaging, pathology, and triage, which lower regulatory and adaptation hurdles. These startup ventures often partner up with health systems for real-world validation [146]. Academia–industry partnerships have accelerated development by combining methodological rigor and datasets from universities with engineering and deployment experience in the industry. Structured models for licensing, spinouts, and co-development, along with the early involvement of clinicians, will improve clinical relevance and trust [147]. Healthcare innovation trends today have shown a shift to multimodal and generative models that can combine images, signals, and clinical texts; also, they are focused on achieving prospective, real-world clinical validation and regulatory alignment. Recent trends have also shown business models that blend Software as a Service (SaaS), device approvals, and hospital partnerships to scale [148]. However, obstacles remain. The availability of high-quality, large-scale healthcare data is fundamental to training reliable ML models, but real-world datasets are usually incomplete, heterogeneous, and noisy [149]. The interoperability between systems is restricted by fragmented EHR architectures, inconsistent data standards, and poor semantic alignment across institutions, all of which are hindering the sharing of data and deployment of AI. Moreover, to have a trustworthy deployment of AI, it is necessary to have strong governance. Privacy, patient consent, and data ownership are issues that call for socio-technical frameworks for the proper management of collected information, especially when data is being transferred from an academic health center to the private industry [150]. Without strong governance, these partnerships are at risk of having ethical drawbacks, a lack of transparency, and unfair outcomes.

6. Perspectives, Limitations, and Future Outlook

6.1. Trends in Secure and Multimodal Medical AI

Federated learning (FL) and allied privacy-preserving techniques such as secure aggregation, differential privacy, and homomorphic encryption are quickly shifting from research proofs-of-concept into multicenter clinical studies. This is due to the fact that they allow for the training of models without centralizing patient data, which directly addresses regulatory and institutional data-sovereignty challenges. Recent reviews have shown FL’s potential for collaborative disease prediction, as seen in Table 6, imaging, and clinical note models, but have also highlighted the constant issues surrounding it, such as non-independent and identically distributed (IID) data across sites, communication or computation overhead, susceptibility to model or inference attacks, and the governance needs for consent, auditing, and liability [151,152]. Simultaneously, foundation models and biomedical large language models (LLMs), such as domain-adapted models like Med-PaLM/Med-PaLM2 and BioGPT, are transforming what automated diagnostic support can look like, going from narrow, task-specific classifiers to generalist systems that are able to reason over text, images, and structured data [153]. These models are showing strong gains in medical quality assurance, knowledge recall, and multimodal tasks; however, they remain restricted by training data sources, calibration, hallucination risks, and requirements for clinical validation [154]. Privacy-preserving foundation models are an upcoming trend where researchers are testing FL and split-model approaches to fine-tune large medical models on local hospital data so that institutions keep raw records while contributing to a shared model and are also exploring certifiable privacy such as data protection (DP), trusted execution environments (TEEs), and strong evaluation procedures to measure fairness and safety. Governance, auditability, and health-equity assessments still remain the top priority, so even when there is a case of biases being amplified by foundation models on erroneous training sets, there must be federated protocols that include bias detection, consistent external validation, and clear data-use governance [155]. For the future, we can expect operational advances such as multi-hospital FL, along with locally fine-tuned foundation models, to grow over the next few years, provided that there are strong technical developments in privacy, standardized external evaluation, and clear governance frameworks supporting the growth of AI use.

6.2. Generative AI, Low-Resource Diagnostics and Robotic Integration

Emerging trend generative models such as generative adversarial networks (GANs), variational autoencoders (VAEs), diffusion models, and large language models are being used more and more to produce realistic synthetic images, physiological time series, EHR text, and tabular records, as they help to augment small datasets, protect privacy, and stress-test diagnostic systems. For example, GANs have been widely applied in healthcare for image and signal synthesis, which allows for anonymization and data augmentation [162]. Figure 6 uses a plant-based schematic that illustrates how future clinical capabilities in diagnostic AI comes from layered technical and institutional foundations. The roots represent the enabling infrastructures such as privacy-preserving computation, federated learning, data standardization, and governance, all of which, are necessary to support the trunk of developing technologies such as self-supervised learning, multimodal foundation models, and continuous model monitoring. From the trunk, branches depict active innovation pathways such as generative synthetic data ecosystems, wearables and edge AI diagnostics, and privacy-safe training of models across hospitals, which give rise to leaves and flowers which represent the emerging trends and clinical outcomes. By connecting future applications with their underlying dependencies, the figure shows that progress in diagnostic AI is limited by foundational readiness rather than algorithmic ambitions alone. A recent scoping review found that adversarial networks, probabilistic models, and large language models were the dominant architectures for the generation of synthetic longitudinal data, time series and medical text, with an emphasis on the goal of privacy preservation and addressing class imbalance [163]. Synthetic data based on GAN has been proposed not just for augmentation but also for creating “synthetic control arms” in clinical trials or for medical education, which reduces the reliance on real patient data [164]. However, challenges remain surrounding mode collapse, the lack of standard evaluation metrics, and potential overfitting to synthetic artifacts rather than the real biology. Some studies on structured EHR data generation have compared probabilistic models, imputation-based models, and GANs, which showed trade-offs in realism, fidelity, and privacy risk [165]. An emerging trend in low- and middle-income countries (LMICs) is that AI is increasingly being seen as a way to strengthen health systems, particularly in improving diagnostics, triage, treatment planning, and upgrading community health procedures [166]. AI is very promising in resource-poor settings due to the growing smartphone usage, cloud computing, and the digitalization of health data; however, it has been cautioned that AI tools must be integrated into existing systems rather than built in isolation [167]. To sustain AI integration in such settings, a framework has to address the available infrastructure, skilled personnel, and local relevance [168]. There are also significant equity concerns as many AI systems come from high-income countries and may not be able to generalize to low-resource countries, which further raises questions on bias, data sovereignty, and even the digital divide [169]. It was found in a review of real-world AI applications in LMICs that although studies report a potential for use in triage and clinical decision support, adoption is hindered by the level of usability, lack of trust from users, and infrastructure limitations such as unstable internet in some places and the poor integration with existing health records [166]. Future perspective wearables with features such as ECG, photoplethysmography (PPG), and inertial sensors can produce continuous multivariate time series that are ideal for lightweight, on-device AI models. Generative models can help synthesize physiologic sequences to pretrain or augment such edge models, which improve anomaly detection or personalization when real data is not widely available. GANs have also indeed been used to produce synthetic Internet-of-Medical-Things (IoMT) sensor datasets that were validated against real sensor data [170]. In robotic platforms such as automated microscopy, robotic ultrasound, or sample-handling robots, synthetic data is being used to simulate sensor results and patient interactions, which allows for safe training and closed-loop testing before its deployment in real clinical settings. A promising research program in the future would include the combination of generative models with some physical simulations such as digital twins, synthetic and real regulatory validation procedures, and human-in-the-loop oversight to ensure safety, ethical considerations, and practicality.

6.3. Perspectives and Drawbacks of Current Study

Although the results in this study sound promising, there are some limitations that must be addressed before considering its findings fully applicable in the real-world. Firstly, the analysis relies on retrospective datasets, which can have issues like selection bias, incomplete annotations, and variability in data acquisition methods. Such limitations can inflate its estimated performance and reduce the generalizability of results when it is translated to real-world clinical environments. Along with this, the heterogeneity of the datasets that were included, while they show real clinical complexity, they can also introduce confounding factors that are often difficult to fully control and may even hide the true contribution of the approach. Secondly, from a methodological point of view, the proposed framework depends hugely on the quality of available data. Any differences in the handling of missing data, feature normalization, or labeling strategies can significantly affect the model’s performance and reproducibility. Moreover, the high computational requirements that are associated with model training and optimization can limit scalability, particularly in environments that are resource-constrained or in institutions lacking advanced infrastructure. The interpretability of the resulting models also remains a challenge as although some performance metrics can indicate a strong predictive capability, translating this predictive capability into clinically meaningful and usable insights require additional validation and user-centered designs. Lastly, practical implementation raises regulatory and operational concerns. The lack of standardized validation procedures and external evaluations does not allow for it to be directly compared with other methodologies. Furthermore, its integration into existing clinical procedures can face challenges regarding transparency, accountability, and automation bias. Addressing these challenges will need multicenter studies, standardized reporting practices, and a closer collaboration between method developers, clinicians, and regulatory bodies so we can ensure that the proposed methodology can be safely deployed in real-world settings.

7. Conclusions

Artificial intelligence is evolving from an early, rule-based system into a diverse system of machine learning, deep learning, and multimodal foundation models, which together, are transforming the field of medical diagnostics. Across its many different applications, such as imaging, pathology, molecular profiling, and clinical decision support, AI showcases its ability to detect small abnormalities, integrate complex multimodal signals, and support an earlier and more accurate diagnosis. However, the prolonged success and responsible deployment of diagnostic AI heavily depend on rooting ethical, interpretable, and equitable principles into every stage of development and implementation. Modern AI systems these days are not only technological advancements but also socio-technical systems that can directly affect patient care. Rigorous oversight will be necessary to address issues such as bias in training data, black-box model decision-making, constant risks of dataset drifts, and the limited generalizability across populations and clinical environments. Ethically, AI in diagnostics requires transparent modeling, calibrated outputs, explainability frameworks that clinicians can trust, and bias evaluation across all demographic groups. It is also very important to ensure equitable access, making sure that AI will improve and not worsen global healthcare disparities. This means the design of systems must be such that it is reliable in diverse, low-resource, real-world environments rather than just in high-income, data-rich settings. In the future, the most useful progress in AI diagnostics will come from the crucial collaborations between clinicians, engineers, and policymakers. Clinicians will bring their domain expertise and contextual judgment, engineers can provide algorithmic innovation and system design, policymakers and regulatory authorities will implement safety, accountability, and equity frameworks, and healthcare institutions will provide the real-world infrastructure and validation that AI will need to be deployed successfully. Multidisciplinary governance is also essential for drift monitoring, post-deployment learning, continuous updates, and maintaining clinical trust. AI has already shown its transformative potential in medical diagnostics, and if it is further developed responsibly with transparency, fairness, and clinical relevance, AI will not only improve diagnostic accuracy but also will help form a more personalized, efficient, and equitable future for global healthcare.

Author Contributions

Conceptualization, D.B.-A., D.R.J.R. and D.A.; methodology, D.B.-A., D.R.J.R. and D.A.; software, D.B.-A., D.R.J.R. and D.A.; validation, D.B.-A., D.R.J.R. and D.A.; formal analysis, D.B.-A., D.R.J.R. and D.A.; investigation, D.B.-A., D.R.J.R. and D.A.; resources, D.B.-A., D.R.J.R. and D.A.; data curation, D.B.-A. and D.A.; writing—original draft preparation, D.B.-A. and D.A.; writing—review and editing, D.B.-A., D.R.J.R. and D.A.; visualization, D.B.-A., D.R.J.R. and D.A.; supervision, D.B.-A., D.R.J.R. and D.A.; project administration, D.A.; funding acquisition, D.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Bajwa, J.; Munir, U.; Nori, A.; Williams, B. Artificial intelligence in healthcare: Transforming the practice of medicine. Future Healthc. J. 2021, 8, e188–e194. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  2. Habehh, H.; Gohel, S. Machine Learning in Healthcare. Curr. Genom. 2021, 22, 291–300. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  3. Shortliffe, E.H. Mycin: A Knowledge-Based Computer Program Applied to Infectious Diseases. In Proceedings of the Annual Symposium on Computer Application in Medical Care, Las Vegas, NV, USA, 5 October 1977; pp. 66–69. [Google Scholar] [PubMed Central]
  4. Perry, C.A. Knowledge bases in medicine: A review. Bull. Med. Libr. Assoc. 1990, 78, 271–282. [Google Scholar] [PubMed] [PubMed Central]
  5. Castellino, R.A. Computer aided detection (CAD): An overview. Cancer Imaging 2005, 5, 17–19. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  6. Orrù, G.; Pettersson-Yeo, W.; Marquand, A.F.; Sartori, G.; Mechelli, A. Using Support Vector Machine to identify imaging biomarkers of neurological and psychiatric disease: A critical review. Neurosci. Biobehav. Rev. 2012, 36, 1140–1152. [Google Scholar] [CrossRef] [PubMed]
  7. Esteva, A.; Kuprel, B.; Novoa, R.A.; Ko, J.; Swetter, S.M.; Blau, H.M.; Thrun, S. Dermatologist-level classification of skin cancer with deep neural networks. Nature 2017, 542, 115–118, Erratum in Nature 2017, 546, 686. https://doi.org/10.1038/nature22985. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  8. Zhou, S.K.; Greenspan, H.; Davatzikos, C.; Duncan, J.S.; van Ginneken, B.; Madabhushi, A.; Prince, J.L.; Rueckert, D.; Summers, R.M. A review of deep learning in medical imaging: Imaging traits, technology trends, case studies with progress highlights, and future promises. Proc. IEEE Inst. Electr. Electron. Eng. 2021, 109, 820–838. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  9. Cross, J.L.; Choma, M.A.; Onofrey, J.A. Bias in medical AI: Implications for clinical decision-making. PLoS Digit. Health 2024, 3, e0000651. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  10. Han, R.; Acosta, J.N.; Shakeri, Z.; Ioannidis, J.P.A.; Topol, E.J.; Rajpurkar, P. Randomised controlled trials evaluating artificial intelligence in clinical practice: A scoping review. Lancet Digit. Health 2024, 6, e367–e373. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  11. Ahmed, M.I.; Spooner, B.; Isherwood, J.; Lane, M.; Orrock, E.; Dennison, A. A Systematic Review of the Barriers to the Implementation of Artificial Intelligence in Healthcare. Cureus 2023, 15, e46454. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  12. Park, S.H.; Choi, J.I.; Fournier, L.; Vasey, B. Randomized Clinical Trials of Artificial Intelligence in Medicine: Why, When, and How? Korean J. Radiol. 2022, 23, 1119–1125. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  13. Aggarwal, R.; Sounderajah, V.; Martin, G.; Ting, D.S.W.; Karthikesalingam, A.; King, D.; Ashrafian, H.; Darzi, A. Diagnostic accuracy of deep learning in medical imaging: A systematic review and meta-analysis. NPJ Digit. Med. 2021, 4, 65. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  14. Beam, A.L.; Kohane, I.S. Big Data and Machine Learning in Health Care. JAMA 2018, 319, 1317–1318. [Google Scholar] [CrossRef] [PubMed]
  15. Roy, S.; Meena, T.; Lim, S.J. Demystifying Supervised Learning in Healthcare 4.0: A New Reality of Transforming Diagnostic Medicine. Diagnostics 2022, 12, 2549. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  16. Rajpurkar, P.; Irvin, J.; Ball, R.L.; Zhu, K.; Yang, B.; Mehta, H.; Duan, T.; Ding, D.; Bagul, A.; Langlotz, C.P.; et al. Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practicing radiologists. PLoS Med. 2018, 15, e1002686. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  17. Litjens, G.; Kooi, T.; Bejnordi, B.E.; Setio, A.A.A.; Ciompi, F.; Ghafoorian, M.; van der Laak, J.A.W.M.; van Ginneken, B.; Sánchez, C.I. A survey on deep learning in medical image analysis. Med. Image Anal. 2017, 42, 60–88. [Google Scholar] [CrossRef] [PubMed]
  18. Lundervold, A.S.; Lundervold, A. An overview of deep learning in medical imaging focusing on MRI. Z. Med. Phys. 2019, 29, 102–127. [Google Scholar] [CrossRef] [PubMed]
  19. Kelly, C.J.; Karthikesalingam, A.; Suleyman, M.; Corrado, G.; King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019, 17, 195. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  20. Bian, Y.; Li, J.; Ye, C.; Jia, X.; Yang, Q. Artificial intelligence in medical imaging: From task-specific models to large-scale foundation models. Chin. Med. J. 2025, 138, 651–663. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  21. Najjar, R. Redefining Radiology: A Review of Artificial Intelligence Integration in Medical Imaging. Diagnostics 2023, 13, 2760. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  22. Campanella, G.; Hanna, M.G.; Geneslaw, L.; Miraflor, A.; Werneck Krauss Silva, V.; Busam, K.J.; Brogi, E.; Reuter, V.E.; Klimstra, D.S.; Fuchs, T.J. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat. Med. 2019, 25, 1301–1309. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  23. Shafi, S.; Parwani, A.V. Artificial intelligence in diagnostic pathology. Diagn. Pathol. 2023, 18, 109. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  24. McGenity, C.; Clarke, E.L.; Jennings, C.; Matthews, G.; Cartlidge, C.; Freduah-Agyemang, H.; Stocken, D.D.; Treanor, D. Artificial intelligence in digital pathology: A systematic review and meta-analysis of diagnostic test accuracy. NPJ Digit. Med. 2024, 7, 114. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  25. Dolezal, J.M.; Kochanny, S.; Dyer, E.; Ramesh, S.; Srisuwananukorn, A.; Sacco, M.; Howard, F.M.; Li, A.; Mohan, P.; Pearson, A.T. Slideflow: Deep learning for digital histopathology with real-time whole-slide visualization. BMC Bioinform. 2024, 25, 134. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  26. Shen, Y.T.; Chen, L.; Yue, W.W.; Xu, H.X. Artificial intelligence in ultrasound. Eur. J. Radiol. 2021, 139, 109717. [Google Scholar] [CrossRef] [PubMed]
  27. Meng, Q.; Sinclair, M.; Zimmer, V.; Hou, B.; Rajchl, M.; Toussaint, N.; Oktay, O.; Schlemper, J.; Gomez, A.; Housden, J.; et al. Weakly Supervised Estimation of Shadow Confidence Maps in Fetal Ultrasound Imaging. IEEE Trans. Med. Imaging 2019, 38, 2755–2767. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  28. Parmar, U.P.S.; Surico, P.L.; Singh, R.B.; Romano, F.; Salati, C.; Spadea, L.; Musa, M.; Gagliano, C.; Mori, T.; Zeppieri, M. Artificial Intelligence (AI) for Early Diagnosis of Retinal Diseases. Medicina 2024, 60, 527. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  29. Driban, M.; Yan, A.; Selvam, A.; Ong, J.; Vupparaboina, K.K.; Chhablani, J. Artificial intelligence in chorioretinal pathology through fundoscopy: A comprehensive review. Int. J. Retin. Vitr. 2024, 10, 36. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  30. De Fauw, J.; Ledsam, J.R.; Romera-Paredes, B.; Nikolov, S.; Tomasev, N.; Blackwell, S.; Askham, H.; Glorot, X.; O’Donoghue, B.; Visentin, D.; et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat. Med. 2018, 24, 1342–1350. [Google Scholar] [CrossRef] [PubMed]
  31. Hannun, A.Y.; Rajpurkar, P.; Haghpanahi, M.; Tison, G.H.; Bourn, C.; Turakhia, M.P.; Ng, A.Y. Cardiologist-level arrhythmia detection and classification in ambulatory electrocardiograms using a deep neural network. Nat. Med. 2019, 25, 65–69, Erratum in Nat. Med. 2019, 25, 530. https://doi.org/10.1038/s41591-019-0359-9. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  32. Meng, L.; Tan, W.; Ma, J.; Wang, R.; Yin, X.; Zhang, Y. Enhancing dynamic ECG heartbeat classification with lightweight transformer model. Artif. Intell. Med. 2022, 124, 102236. [Google Scholar] [CrossRef] [PubMed]
  33. Jaya Prakash, A.; Nasreddine Belkacem, A.; Elfadel, I.M.; Jelinek, H.F.; Atef, M. Advances in machine and deep learning for ECG beat classification: A systematic review. Front. Digit. Health 2025, 7, 1649923. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  34. Ribeiro, A.H.; Ribeiro, M.H.; Paixão, G.M.M.; Oliveira, D.M.; Gomes, P.R.; Canazart, J.A.; Ferreira, M.P.S.; Andersson, C.R.; Macfarlane, P.W.; Meira, W., Jr.; et al. Automatic diagnosis of the 12-lead ECG using a deep neural network. Nat. Commun. 2020, 11, 1760, Erratum in Nat. Commun. 2020, 11, 2227. https://doi.org/10.1038/s41467-020-16172-1. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  35. Acharya, U.R.; Hagiwara, Y.; Adeli, H. Automated seizure prediction. Epilepsy Behav. 2018, 88, 251–261. [Google Scholar] [CrossRef] [PubMed]
  36. Roy, Y.; Banville, H.; Albuquerque, I.; Gramfort, A.; Falk, T.H.; Faubert, J. Deep learning-based electroencephalography analysis: A systematic review. J. Neural Eng. 2019, 16, 051001. [Google Scholar] [CrossRef] [PubMed]
  37. Hussein, S.; Cao, K.; Song, Q.; Bagci, U. Risk stratification of lung nodules using 3D CNN-based multi-task learning. In Information Processing in Medical Imaging. IPMI 2017, Boone, NC, USA, 25–30 June 2017; Springer: New York, NY, USA, 2017. [Google Scholar]
  38. Choi, E.; Bahadori, M.T.; Schuetz, A.; Stewart, W.F.; Sun, J. Doctor AI: Predicting Clinical Events via Recurrent Neural Networks. JMLR Workshop Conf. Proc. 2016, 56, 301–318. [Google Scholar] [PubMed] [PubMed Central]
  39. Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), Minneapolis, MN, USA, 2–7 June 2019. [Google Scholar] [CrossRef]
  40. Alsentzer, E.; Murphy, J.R.; Boag, W.; Weng, W.-H.; Jin, D.; Naumann, T.; McDermott, M. Publicly Available Clinical BERT Embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop (ClinicalNLP), NAACL-HLT, Minneapolis, MN, USA, 7 June 2019. [Google Scholar] [CrossRef]
  41. Acharya, A.; Shrestha, S.; Chen, A.; Conte, J.; Avramovic, S.; Sikdar, S.; Anastasopoulos, A.; Das, S. Clinical risk prediction using language models: Benefits and considerations. J. Am. Med. Inform. Assoc. 2024, 31, 1856–1864. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  42. Johnson, A.E.; Pollard, T.J.; Shen, L.; Lehman, L.W.; Feng, M.; Ghassemi, M.; Moody, B.; Szolovits, P.; Celi, L.A.; Mark, R.G. MIMIC-III, a freely accessible critical care database. Sci. Data 2016, 3, 160035. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  43. Moody, G.B.; Mark, R.G.; Goldberger, A.L. PhysioNet: Physiologic signals, time series and related open source software for basic, clinical, and applied research. Annu. Int. Conf. IEEE Eng. Med. Biol. Soc. 2011, 2011, 8327–8330. [Google Scholar] [CrossRef] [PubMed]
  44. Mao, Q.; Jay, M.; Hoffman, J.L.; Calvert, J.; Barton, C.; Shimabukuro, D.; Shieh, L.; Chettipally, U.; Fletcher, G.; Kerem, Y.; et al. Multicentre validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU. BMJ Open 2018, 8, e017833. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  45. Kwon, J.M.; Lee, Y.; Lee, Y.; Lee, S.; Park, J. An Algorithm Based on Deep Learning for Predicting In-Hospital Cardiac Arrest. J. Am. Heart Assoc. 2018, 7, e008678. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  46. Athanasopoulou, K.; Michalopoulou, V.I.; Scorilas, A.; Adamopoulos, P.G. Integrating Artificial Intelligence in Next-Generation Sequencing: Advances, Challenges, and Future Directions. Curr. Issues Mol. Biol. 2025, 47, 470. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  47. Baião, A.R.; Cai, Z.; Poulos, R.C.; Robinson, P.J.; Reddel, R.R.; Zhong, Q.; Vinga, S.; Gonçalves, E. A technical review of multi-omics data integration methods: From classical statistical to deep generative approaches. Brief. Bioinform. 2025, 26, bbaf355. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  48. Mann, M.; Kumar, C.; Zeng, W.F.; Strauss, M.T. Artificial intelligence for proteomics and biomarker discovery. Cell Syst. 2021, 12, 759–770. [Google Scholar] [CrossRef] [PubMed]
  49. Kitaoka, Y.; Uchihashi, T.; Kawata, S.; Nishiura, A.; Yamamoto, T.; Hiraoka, S.-i.; Yokota, Y.; Isomura, E.T.; Kogo, M.; Tanaka, S.; et al. Role and Potential of Artificial Intelligence in Biomarker Discovery and Development of Treatment Strategies for Amyotrophic Lateral Sclerosis. Int. J. Mol. Sci. 2025, 26, 4346. [Google Scholar] [CrossRef] [PubMed]
  50. Chi, J.; Shu, J.; Li, M.; Mudappathi, R.; Jin, Y.; Lewis, F.; Boon, A.; Qin, X.; Liu, L.; Gu, H. Artificial Intelligence in Metabolomics: A Current Review. Trends Anal. Chem. 2024, 178, 117852. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  51. Gloaguen, Y.; Kirwan, J.A.; Beule, D. Deep Learning-Assisted Peak Curation for Large-Scale LC-MS Metabolomics. Anal. Chem. 2022, 94, 4930–4937. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  52. Acharya, D.; Mukhopadhyay, A. A comprehensive review of machine learning techniques for multi-omics data integration: Challenges and applications in precision oncology. Brief. Funct. Genom. 2024, 23, 549–560. [Google Scholar] [CrossRef] [PubMed]
  53. Lin, M.; Guo, J.; Gu, Z.; Tang, W.; Tao, H.; You, S.; Jia, D.; Sun, Y.; Jia, P. Machine learning and multi-omics integration: Advancing cardiovascular translational research and clinical practice. J. Transl. Med. 2025, 23, 388. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  54. Parmar, C.; Grossmann, P.; Rietveld, D.; Rietbergen, M.M.; Lambin, P.; Aerts, H.J. Radiomic Machine-Learning Classifiers for Prognostic Biomarkers of Head and Neck Cancer. Front. Oncol. 2015, 5, 272. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  55. Wu, W.; Parmar, C.; Grossmann, P.; Quackenbush, J.; Lambin, P.; Bussink, J.; Mak, R.; Aerts, H.J. Exploratory Study to Identify Radiomics Classifiers for Lung Cancer Histology. Front. Oncol. 2016, 6, 71. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  56. Alkhanbouli, R.; Matar Abdulla Almadhaani, H.; Alhosani, F.; Simsekler, M.C.E. The role of explainable artificial intelligence in disease prediction: A systematic literature review and future research directions. BMC Med. Inform. Decis. Mak. 2025, 25, 110. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  57. Fuhrman, J.D.; Gorre, N.; Hu, Q.; Li, H.; El Naqa, I.; Giger, M.L. A review of explainable and interpretable AI with applications in COVID-19 imaging. Med. Phys. 2022, 49, 1–14. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  58. Saarela, M.; Podgorelec, V. Recent Applications of Explainable AI (XAI): A Systematic Literature Review. Appl. Sci. 2024, 14, 8884. [Google Scholar] [CrossRef]
  59. Salimparsa, M.; Sedig, K.; Lizotte, D.J.; Abdullah, S.S.; Chalabianloo, N.; Muanda, F.T. Explainable AI for Clinical Decision Support Systems: Literature Review, Key Gaps, and Research Synthesis. Informatics 2025, 12, 119. [Google Scholar] [CrossRef]
  60. Solomon, J.; Dauber-Decker, K.; Richardson, S.; Levy, S.; Khan, S.; Coleman, B.; Persaud, R.; Chelico, J.; King, D.; Spyropoulos, A.; et al. Integrating Clinical Decision Support Into Electronic Health Record Systems Using a Novel Platform (EvidencePoint): Developmental Study. JMIR Form. Res. 2023, 7, e44065. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  61. Patterson, B.W.; Pulia, M.S.; Ravi, S.; Hoonakker, P.L.T.; Schoofs Hundt, A.; Wiegmann, D.; Wirkus, E.J.; Johnson, S.; Carayon, P. Scope and Influence of Electronic Health Record-Integrated Clinical Decision Support in the Emergency Department: A Systematic Review. Ann. Emerg. Med. 2019, 74, 285–296. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  62. Weissman, G.E. FDA Regulation of Predictive Clinical Decision-Support Tools: What Does It Mean for Hospitals? J. Hosp. Med. 2021, 16, 244–246. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  63. Karnik, K. FDA regulation of clinical decision support software. J. Law Biosci. 2014, 1, 202–208. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  64. Ueda, D.; Kakinuma, T.; Fujita, S.; Kamagata, K.; Fushimi, Y.; Ito, R.; Matsui, Y.; Nozaki, T.; Nakaura, T.; Fujima, N.; et al. Fairness of artificial intelligence in healthcare: Review and recommendations. Jpn. J. Radiol. 2024, 42, 3–15. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  65. Nasir, M.; Siddiqui, K.; Ahmed, S. Ethical-legal implications of AI-powered healthcare in critical perspective. Front. Artif. Intell. 2025, 8, 1619463. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  66. Park, M.K.; Ashwood, N.; Capes, N. Ethics of Artificial Intelligence in Medicine. Cureus 2025, 17, e83567. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  67. McKinney, S.M.; Sieniek, M.; Godbole, V.; Godwin, J.; Antropova, N.; Ashrafian, H.; Back, T.; Chesus, M.; Corrado, G.S.; Darzi, A.; et al. International evaluation of an AI system for breast cancer screening. Nature 2020, 577, 89–94, Erratum in Nature 2020, 586, E19. https://doi.org/10.1038/s41586-020-2679-9. [Google Scholar] [CrossRef] [PubMed]
  68. Attia, Z.I.; Noseworthy, P.A.; Lopez-Jimenez, F.; Asirvatham, S.J.; Deshmukh, A.J.; Gersh, B.J.; Carter, R.E.; Yao, X.; Rabinstein, A.A.; Erickson, B.J.; et al. An artificial intelligence-enabled ECG algorithm for the identification of patients with atrial fibrillation during sinus rhythm: A retrospective analysis of outcome prediction. Lancet 2019, 394, 861–867. [Google Scholar] [CrossRef] [PubMed]
  69. Neijzen, D.; Lunter, G. Unsupervised learning for medical data: A review of probabilistic factorization methods. Stat. Med. 2023, 42, 5541–5554. [Google Scholar] [CrossRef] [PubMed]
  70. Raza, K.; Singh, N.K. A Tour of Unsupervised Deep Learning for Medical Image Analysis. Curr. Med. Imaging Rev. 2021, 17, 1059–1077. [Google Scholar] [CrossRef] [PubMed]
  71. Hu, M.; Zhang, J.; Matkovic, L.; Liu, T.; Yang, X. Reinforcement learning in medical image analysis: Concepts, applications, challenges, and future directions. J. Appl. Clin. Med. Phys. 2023, 24, e13898. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  72. Ching, T.; Himmelstein, D.S.; Beaulieu-Jones, B.K.; Kalinin, A.A.; Do, B.T.; Way, G.P.; Ferrero, E.; Agapow, P.M.; Zietz, M.; Hoffman, M.M.; et al. Opportunities and obstacles for deep learning in biology and medicine. J. R. Soc. Interface 2018, 15, 20170387. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  73. Topol, E.J. High-performance medicine: The convergence of human and artificial intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef] [PubMed]
  74. Delord, M.; Sun, X.; Learoyd, A.; Curcin, V.; Wolfe, C.; Ashworth, M.; Douiri, A. Patient-oriented unsupervised learning to uncover the patterns of multimorbidity associated with stroke using primary care electronic health records. BMC Prim. Care 2024, 25, 419. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  75. Mehari, T.; Strodthoff, N. Self-supervised representation learning from 12-lead ECG data. Comput. Biol. Med. 2022, 141, 105114. [Google Scholar] [CrossRef] [PubMed]
  76. Fenwick, D.; NaderiAlizadeh, N.; Tarokh, V.; Clark, D.; Rajagopal, J.; Kapadia, A.; Felice, N.; Samei, E.; Abadi, E. Black-box Optimization of CT Acquisition and Reconstruction Parameters: A Reinforcement Learning Approach. Proc. SPIE Int. Soc. Opt. Eng. 2025, 13405, 1340528. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  77. Hirte, A.U.; Platscher, M.; Joyce, T.; Heit, J.J.; Tranvinh, E.; Federau, C. Realistic generation of diffusion-weighted magnetic resonance brain images with deep generative models. Magn. Reson. Imaging 2021, 81, 60–66. [Google Scholar] [CrossRef] [PubMed]
  78. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015; Navab, N., Hornegger, J., Wells, W., Frangi, A., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2015; Volume 9351. [Google Scholar] [CrossRef]
  79. Isensee, F.; Jaeger, P.F.; Kohl, S.A.A.; Petersen, J.; Maier-Hein, K.H. nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods 2021, 18, 203–211. [Google Scholar] [CrossRef] [PubMed]
  80. Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef]
  81. Li, Y.; Song, S.; Zhu, L.; Zhang, X.; Mou, Y.; Lei, M.; Wang, W.; Tao, Z. Machine learning-based prediction model for patients with recurrent Staphylococcus aureus bacteremia. BMC Med. Inform. Decis. Mak. 2025, 25, 99. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  82. Urundai Meeran, S.B. Walrus Optimization-Enhanced ResNet-50 for AI-Driven Renal Malignancy Prediction with Occlusion Sensitivity-Based Interpretation. Asian Pac. J. Cancer Prev. 2025, 26, 2995–3004. [Google Scholar] [CrossRef] [PubMed]
  83. Neuwieser, H.; Jami, N.V.S.J.; Meier, R.J.; Liebsch, G.; Felthaus, O.; Klein, S.; Schreml, S.; Berneburg, M.; Prantl, L.; Leutheuser, H.; et al. Interpreting Venous and Arterial Ulcer Images Through the Grad-CAM Lens: Insights and Implications in CNN-Based Wound Image Classification. Diagnostics 2025, 15, 2184. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  84. Ennab, M.; Mcheick, H. Enhancing interpretability and accuracy of AI models in healthcare: A comprehensive review on challenges and future directions. Front. Robot. AI 2024, 11, 1444763. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  85. Chen, M.; Wang, Y.; Wang, Q.; Shi, J.; Wang, H.; Ye, Z.; Xue, P.; Qiao, Y. Impact of human and artificial intelligence collaboration on workload reduction in medical image interpretation. NPJ Digit. Med. 2024, 7, 349. [Google Scholar] [CrossRef]
  86. Sahiner, B.; Chen, W.; Samala, R.K.; Petrick, N. Data drift in medical machine learning: Implications and potential remedies. Br. J. Radiol. 2023, 96, 20220878. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  87. Cunningham, J.W.; Singh, P.; Reeder, C.; Claggett, B.; Marti-Castellote, P.M.; Lau, E.S.; Khurshid, S.; Batra, P.; Lubitz, S.A.; Maddah, M.; et al. Natural Language Processing for Adjudication of Heart Failure Hospitalizations in a Multi-Center Clinical Trial. medRxiv 2023, 23, 23294234. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  88. Teshale, A.B.; Htun, H.L.; Vered, M.; Owen, A.J.; Freak-Poli, R. A Systematic Review of Artificial Intelligence Models for Time-to-Event Outcome Applied in Cardiovascular Disease Risk Prediction. J. Med. Syst. 2024, 48, 68. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  89. Tran, K.A.; Kondrashova, O.; Bradley, A.; Williams, E.D.; Pearson, J.V.; Waddell, N. Deep learning in cancer diagnosis, prognosis and treatment selection. Genome Med. 2021, 13, 152. [Google Scholar] [CrossRef] [PubMed]
  90. Saxena, S.; Jena, B.; Gupta, N.; Das, S.; Sarmah, D.; Bhattacharya, P.; Nath, T.; Paul, S.; Fouda, M.M.; Kalra, M.; et al. Role of Artificial Intelligence in Radiogenomics for Cancers in the Era of Precision Medicine. Cancers 2022, 14, 2860. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  91. Lareyre, F.; Chaudhuri, A.; Behrendt, C.A.; Pouhin, A.; Teraa, M.; Boyle, J.R.; Tulamo, R.; Raffort, J. Artificial intelligence-based predictive models in vascular diseases. Semin. Vasc. Surg. 2023, 36, 440–447. [Google Scholar] [CrossRef] [PubMed]
  92. Goffart, S.; Delingette, H.; Chierici, A.; Guzzi, L.; Nasr, B.; Lareyre, F.; Raffort, J. Artificial Intelligence Techniques for Prognostic and Diagnostic Assessments in Peripheral Artery Disease: A Scoping Review. Angiology 2025. Epub ahead of printing. [Google Scholar] [CrossRef] [PubMed]
  93. Mehri-Kakavand, G.; Mdletshe, S.; Wang, A. A Comprehensive Review on the Application of Artificial Intelligence for Predicting Postsurgical Recurrence Risk in Early-Stage Non-Small Cell Lung Cancer Using Computed Tomography, Positron Emission Tomography, and Clinical Data. J. Med. Radiat. Sci. 2025, 72, 280–296. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  94. Azadi Moghadam, P.; Bashashati, A.; Goldenberg, S.L. Artificial Intelligence and Pathomics: Prostate Cancer. Urol. Clin. N. Am. 2024, 51, 15–26. [Google Scholar] [CrossRef] [PubMed]
  95. Nemati, S.; Holder, A.; Razmi, F.; Stanley, M.D.; Clifford, G.D.; Buchman, T.G. An Interpretable Machine Learning Model for Accurate Prediction of Sepsis in the ICU. Crit. Care Med. 2018, 46, 547–553. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  96. Desautels, T.; Calvert, J.; Hoffman, J.; Jay, M.; Kerem, Y.; Shieh, L.; Shimabukuro, D.; Chettipally, U.; Feldman, M.D.; Barton, C.; et al. Prediction of Sepsis in the Intensive Care Unit With Minimal Electronic Health Record Data: A Machine Learning Approach. JMIR Med. Inform. 2016, 4, e28. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  97. Fleuren, L.M.; Klausch, T.L.T.; Zwager, C.L.; Schoonmade, L.J.; Guo, T.; Roggeveen, L.F.; Swart, E.L.; Girbes, A.R.J.; Thoral, P.; Ercole, A.; et al. Machine learning for the prediction of sepsis: A systematic review and meta-analysis of diagnostic test accuracy. Intensive Care Med. 2020, 46, 383–400. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  98. Lee, Y.; Kwon, J.M.; Lee, Y.; Park, H.; Cho, H.; Park, J. Deep Learning in the Medical Domain: Predicting Cardiac Arrest Using Deep Learning. Acute Crit. Care 2018, 33, 117–120. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  99. Lee, H.; Yang, H.L.; Ryu, H.G.; Jung, C.W.; Cho, Y.J.; Yoon, S.B.; Yoon, H.K.; Lee, H.C. Real-time machine learning model to predict in-hospital cardiac arrest using heart rate variability in ICU. NPJ Digit. Med. 2023, 6, 215. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  100. Aerts, H.J.; Velazquez, E.R.; Leijenaar, R.T.; Parmar, C.; Grossmann, P.; Carvalho, S.; Bussink, J.; Monshouwer, R.; Haibe-Kains, B.; Rietveld, D.; et al. Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach. Nat. Commun. 2014, 5, 4006, Erratum in Nat. Commun. 2014, 5, 4644. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  101. Lambin, P.; Leijenaar, R.T.H.; Deist, T.M.; Peerlings, J.; de Jong, E.E.C.; van Timmeren, J.; Sanduleanu, S.; Larue, R.T.H.M.; Even, A.J.G.; Jochems, A.; et al. Radiomics: The bridge between medical imaging and personalized medicine. Nat. Rev. Clin. Oncol. 2017, 14, 749–762. [Google Scholar] [CrossRef] [PubMed]
  102. van Timmeren, J.E.; Cester, D.; Tanadini-Lang, S.; Alkadhi, H.; Baessler, B. Radiomics in medical imaging-“how-to” guide and critical reflection. Insights Imaging 2020, 11, 91. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  103. Amrollahi, F.; Shashikumar, S.P.; Razmi, F.; Nemati, S. Contextual Embeddings from Clinical Notes Improves Prediction of Sepsis. AMIA Annu. Symp. Proc. 2021, 2020, 197–202. [Google Scholar] [PubMed] [PubMed Central]
  104. Rosnati, M.; Fortuin, V. MGP-AttTCN: An interpretable machine learning model for the prediction of sepsis. PLoS ONE 2021, 16, e0251248. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  105. Ueno, R.; Xu, L.; Uegami, W.; Matsui, H.; Okui, J.; Hayashi, H.; Miyajima, T.; Hayashi, Y.; Pilcher, D.; Jones, D. Value of laboratory results in addition to vital signs in a machine learning algorithm to predict in-hospital cardiac arrest: A single-center retrospective cohort study. PLoS ONE 2020, 15, e0235835. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  106. Ai, Y.; Liu, J.; Li, Y.; Wang, F.; Du, X.; Jain, R.K.; Lin, L.; Chen, Y.W. SAMA: A Self-and-Mutual Attention Network for Accurate Recurrence Prediction of Non-Small Cell Lung Cancer Using Genetic and CT Data. IEEE J. Biomed. Health Inform. 2025, 29, 3220–3233. [Google Scholar] [CrossRef] [PubMed]
  107. Liu, Y.; Yu, Y.; Ouyang, J.; Jiang, B.; Yang, G.; Ostmeier, S.; Wintermark, M.; Michel, P.; Liebeskind, D.S.; Lansberg, M.G.; et al. Functional Outcome Prediction in Acute Ischemic Stroke Using a Fused Imaging and Clinical Deep Learning Model. Stroke 2023, 54, 2316–2327. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  108. Deng, Y.; Liu, S.; Wang, Z.; Wang, Y.; Jiang, Y.; Liu, B. Explainable time-series deep learning models for the prediction of mortality, prolonged length of stay and 30-day readmission in intensive care patients. Front. Med. 2022, 9, 933037. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  109. Johnson, K.B.; Wei, W.Q.; Weeraratne, D.; Frisse, M.E.; Misulis, K.; Rhee, K.; Zhao, J.; Snowdon, J.L. Precision Medicine, AI, and the Future of Personalized Health Care. Clin. Transl. Sci. 2021, 14, 86–93. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  110. Trivizakis, E.; Papadakis, G.Z.; Souglakos, I.; Papanikolaou, N.; Koumakis, L.; Spandidos, D.A.; Tsatsakis, A.; Karantanas, A.H.; Marias, K. Artificial intelligence radiogenomics for advancing precision and effectiveness in oncologic care (Review). Int. J. Oncol. 2020, 57, 43–53. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  111. Sotudian, S.; Paschalidis, I.C. Machine Learning for Pharmacogenomics and Personalized Medicine: A Ranking Model for Drug Sensitivity Prediction. EEE/ACM Trans. Comput. Biol. Bioinform. 2022, 19, 2324–2333. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  112. Kirienko, M.; Sollini, M.; Corbetta, M.; Voulaz, E.; Gozzi, N.; Interlenghi, M.; Gallivanone, F.; Castiglioni, I.; Asselta, R.; Duga, S.; et al. Radiomics and gene expression profile to characterise the disease and predict outcome in patients with lung cancer. Eur. J. Nucl. Med. Mol. Imaging 2021, 48, 3643–3655. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  113. Binczyk, F.; Prazuch, W.; Bozek, P.; Polanska, J. Radiomics and artificial intelligence in lung cancer screening. Transl. Lung Cancer Res. 2021, 10, 1186–1199. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  114. He, W.; Huang, W.; Zhang, L.; Wu, X.; Zhang, S.; Zhang, B. Radiogenomics: Bridging the gap between imaging and genomics for precision oncology. MedComm 2024, 5, e722. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  115. Chen, Z.; Liang, N.; Zhang, H.; Li, H.; Yang, Y.; Zong, X.; Chen, Y.; Wang, Y.; Shi, N. Harnessing the power of clinical decision support systems: Challenges and opportunities. Open Heart 2023, 10, e002432. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  116. Susanto, A.P.; Lyell, D.; Widyantoro, B.; Berkovsky, S.; Magrabi, F. Effects of machine learning-based clinical decision support systems on decision-making, care delivery, and patient outcomes: A scoping review. J. Am. Med. Inform. Assoc. 2023, 30, 2050–2063. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  117. Chaparro, J.D.; Beus, J.M.; Dziorny, A.C.; Hagedorn, P.A.; Hernandez, S.; Kandaswamy, S.; Kirkendall, E.S.; McCoy, A.B.; Muthu, N.; Orenstein, E.W. Clinical Decision Support Stewardship: Best Practices and Techniques to Monitor and Improve Interruptive Alerts. Appl. Clin. Inform. 2022, 13, 560–568. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  118. Elhaddad, M.; Hamam, S. AI-Driven Clinical Decision Support Systems: An Ongoing Pursuit of Potential. Cureus 2024, 16, e57728. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  119. Benjamens, S.; Dhunnoo, P.; Meskó, B. The state of artificial intelligence-based FDA-approved medical devices and algorithms: An online database. npj Digit. Med. 2020, 3, 118. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  120. Muralidharan, V.; Adewale, B.A.; Huang, C.J.; Nta, M.T.; Ademiju, P.O.; Pathmarajah, P.; Hang, M.K.; Adesanya, O.; Abdullateef, R.O.; Babatunde, A.O.; et al. A scoping review of reporting gaps in FDA-approved AI medical devices. npj Digit. Med. 2024, 7, 273. [Google Scholar] [CrossRef]
  121. Lee, J.T.; Moffett, A.T.; Maliha, G.; Faraji, Z.; Kanter, G.P.; Weissman, G.E. Analysis of Devices Authorized by the FDA for Clinical Decision Support in Critical Care. JAMA Intern. Med. 2023, 183, 1399–1401. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  122. Gilbert, S.; Fenech, M.; Hirsch, M.; Upadhyay, S.; Biasiucci, A.; Starlinger, J. Algorithm Change Protocols in the Regulation of Adaptive Machine Learning-Based Medical Devices. J. Med. Internet Res. 2021, 23, e30545. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  123. Rosenthal, J.T.; Beecy, A.; Sabuncu, M.R. Rethinking clinical trials for medical AI with dynamic deployments of adaptive systems. npj Digit. Med. 2025, 8, 252. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  124. Subasri, V.; Krishnan, A.; Kore, A.; Dhalla, A.; Pandya, D.; Wang, B.; Malkin, D.; Razak, F.; Verma, A.A.; Goldenberg, A.; et al. Detecting and Remediating Harmful Data Shifts for the Responsible Deployment of Clinical AI Models. JAMA Netw. Open 2025, 8, e2513685. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  125. Lam, J.Y.; Lu, X.; Shashikumar, S.P.; Lee, Y.S.; Miller, M.; Pour, H.; Boussina, A.E.; Pearce, A.K.; Malhotra, A.; Nemati, S. Development, deployment, and continuous monitoring of a machine learning model to predict respiratory failure in critically ill patients. JAMIA Open 2024, 7, ooae141. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  126. Gulshan, V.; Peng, L.; Coram, M.; Stumpe, M.C.; Wu, D.; Narayanaswamy, A.; Venugopalan, S.; Widner, K.; Madams, T.; Cuadros, J.; et al. Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs. JAMA 2016, 316, 2402–2410. [Google Scholar] [CrossRef] [PubMed]
  127. Zhou, N.; Zhang, C.T.; Lv, H.Y.; Hao, C.X.; Li, T.J.; Zhu, J.J.; Zhu, H.; Jiang, M.; Liu, K.W.; Hou, H.L.; et al. Concordance Study Between IBM Watson for Oncology and Clinical Practice for Patients with Cancer in China. Oncologist 2019, 24, 812–819. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  128. Somashekhar, S.P.; Sepúlveda, M.J.; Puglielli, S.; Norden, A.D.; Shortliffe, E.H.; Rohit Kumar, C.; Rauthan, A.; Arun Kumar, N.; Patil, P.; Rhee, K.; et al. Watson for Oncology and breast cancer treatment recommendations: Agreement with an expert multidisciplinary tumor board. Ann. Oncol. 2018, 29, 418–423. [Google Scholar] [CrossRef] [PubMed]
  129. Montalbano, P. Ethical and Legal Considerations of Medical Artificial Intelligence. Prim. Care 2025, 52, 769–779. [Google Scholar] [CrossRef] [PubMed]
  130. Pham, T. Ethical and legal considerations in healthcare AI: Innovation and policy for safe and fair use. R. Soc. Open Sci. 2025, 12, 241873. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  131. Vardas, E.P.; Marketou, M.; Vardas, P.E. Medicine, healthcare and the AI act: Gaps, challenges and future implications. Eur. Heart J. Digit. Health 2025, 6, 833–839. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  132. Ardic, N.; Dinc, R. Artificial Intelligence in Healthcare: Current Regulatory Landscape and Future Directions. Br. J. Hosp. Med. 2025, 86, 1–21. [Google Scholar] [CrossRef] [PubMed]
  133. Bignami, E.; Darhour, L.J.; Franco, G.; Guarnieri, M.; Bellini, V. AI policy in healthcare: A checklist-based methodology for structured implementation. J. Anesth. Analg. Crit. Care 2025, 5, 56. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  134. Andersen, E.S.; Birk-Korch, J.B.; Hansen, R.S.; Fly, L.H.; Röttger, R.; Arcani, D.M.C.; Brasen, C.L.; Brandslund, I.; Madsen, J.S. Monitoring performance of clinical artificial intelligence in health care: A scoping review. JBI Evid. Synth. 2024, 22, 2423–2446. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  135. Eche, T.; Schwartz, L.H.; Mokrane, F.Z.; Dercle, L. Toward Generalizability in the Deployment of Artificial Intelligence in Radiology: Role of Computation Stress Testing to Overcome Underspecification. Radiol. Artif. Intell. 2021, 3, e210097. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  136. Asan, O.; Bayrak, A.E.; Choudhury, A. Artificial Intelligence and Human Trust in Healthcare: Focus on Clinicians. J. Med. Internet Res. 2020, 22, e15154. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  137. Davis, S.E.; Greevy, R.A., Jr.; Lasko, T.A.; Walsh, C.G.; Matheny, M.E. Detection of calibration drift in clinical prediction models to inform model updating. J. Biomed. Inform. 2020, 112, 103611. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  138. Parikh, R.B.; Zhang, Y.; Kolla, L.; Chivers, C.; Courtright, K.R.; Zhu, J.; Navathe, A.S.; Chen, J. Performance drift in a mortality prediction algorithm among patients with cancer during the SARS-CoV-2 pandemic. J. Am. Med. Inform. Assoc. 2023, 30, 348–354. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  139. Rahmani, K.; Thapa, R.; Tsou, P.; Casie Chetty, S.; Barnes, G.; Lam, C.; Foon Tso, C. Assessing the effects of data drift on the performance of machine learning models used in clinical sepsis prediction. Int. J. Med. Inform. 2023, 173, 104930. [Google Scholar] [CrossRef] [PubMed]
  140. Bradwell, K.R.; Wooldridge, J.T.; Amor, B.; Bennett, T.D.; Anand, A.; Bremer, C.; Yoo, Y.J.; Qian, Z.; Johnson, S.G.; Pfaff, E.R.; et al. Harmonizing units and values of quantitative data elements in a very large nationally pooled electronic health record (EHR) dataset. J. Am. Med. Inform. Assoc. 2022, 29, 1172–1182. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  141. Daneshjou, R.; Vodrahalli, K.; Novoa, R.A.; Jenkins, M.; Liang, W.; Rotemberg, V.; Ko, J.; Swetter, S.M.; Bailey, E.E.; Gevaert, O.; et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci. Adv. 2022, 8, eabq6147. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  142. Wang, W.; Wildgruber, M.; Wang, Y. Attacking medical images with minimal noise: Exploiting vulnerabilities in medical deep-learning systems. Quant. Imaging Med. Surg. 2024, 14, 9374–9384. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  143. Venkatesh, K.; Mutasa, S.; Moore, F.; Sulam, J.; Yi, P.H. Gradient-Based Saliency Maps Are Not Trustworthy Visual Explanations of Automated AI Musculoskeletal Diagnoses. J. Imaging Inform. Med. 2024, 37, 2490–2499. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  144. Taber, P.; Radloff, C.; Del Fiol, G.; Staes, C.; Kawamoto, K. New Standards for Clinical Decision Support: A Survey of The State of Implementation. Yearb. Med. Inform. 2021, 30, 159–171. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  145. Goddard, K.; Roudsari, A.; Wyatt, J.C. Automation bias: A systematic review of frequency, effect mediators, and mitigators. J. Am. Med.Inform. Assoc. 2012, 19, 121–127. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  146. Alhejaily, A.G. Artificial intelligence in healthcare (Review). Biomed. Rep. 2024, 22, 11. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  147. Pantanowitz, L.; Bui, M.M.; Chauhan, C.; ElGabry, E.; Hassell, L.; Li, Z.; Parwani, A.V.; Salama, M.E.; Sebastian, M.M.; Tulman, D.; et al. Rules of engagement: Promoting academic-industry partnership in the era of digital pathology and artificial intelligence. Acad. Pathol. 2022, 9, 100026. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  148. Samah, T.; Samar, M. Investigating the Key Trends in Applying Artificial Intelligence to Health Technologies: A Scoping Review. PLoS ONE 2025, 20, e0322197. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  149. Stanfill, M.H.; Marc, D.T. Health Information Management: Implications of Artificial Intelligence on Healthcare Data and Information Management. Yearb. Med. Inform. 2019, 28, 56–64. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  150. Ambalavanan, R.; Snead, R.S.; Marczika, J.; Towett, G.; Malioukis, A.; Mbogori-Kairichi, M. Challenges and strategies in building a foundational digital health data integration ecosystem: A systematic review and thematic synthesis. Front. Health Serv. 2025, 5, 1600689. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  151. Khalid, N.; Qayyum, A.; Bilal, M.; Al-Fuqaha, A.; Qadir, J. Privacy-preserving artificial intelligence in healthcare: Techniques and applications. Comput. Biol. Med. 2023, 158, 106848. [Google Scholar] [CrossRef] [PubMed]
  152. Teo, Z.L.; Jin, L.; Li, S.; Miao, D.; Zhang, X.; Ng, W.Y.; Tan, T.F.; Lee, D.M.; Chua, K.J.; Heng, J.; et al. Federated machine learning in healthcare: A systematic review on clinical applications and technical architecture. Cell Rep. Med. 2024, 5, 101419, Erratum in Cell Rep. Med. 2024, 5, 101481. https://doi.org/10.1016/j.xcrm.2024.101481. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  153. Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Amin, M.; Hou, L.; Clark, K.; Pfohl, S.R.; Cole-Lewis, H.; et al. Toward expert-level medical question answering with large language models. Nat. Med. 2025, 31, 943–950. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  154. Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef] [PubMed]
  155. Eden, R.; Chukwudi, I.; Bain, C.; Barbieri, S.; Callaway, L.; de Jersey, S.; George, Y.; Gorse, A.D.; Lawley, M.; Marendy, P.; et al. A scoping review of the governance of federated learning in healthcare. npj Digit. Med. 2025, 8, 427. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  156. Guan, H.; Yap, P.T.; Bozoki, A.; Liu, M. Federated learning for medical image analysis: A survey. Pattern Recognit. 2024, 151, 110424. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  157. Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. Large language models encode clinical knowledge. Nature 2023, 620, 172–180, Erratum in Nature 2023, 620, E19. https://doi.org/10.1038/s41586-023-06455-0. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  158. Wang, S.; Zhou, X.; Li, C.; Wang, S.; Li, Y.; Tan, T.; Zheng, H. Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation. Research 2025, 8, 1029. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  159. Sabry, F.; Eltaras, T.; Labda, W.; Alzoubi, K.; Malluhi, Q. Machine Learning for Healthcare Wearable Devices: The Big Picture. J. Healthc. Eng. 2022, 2022, 4653923. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  160. Ning, G.; Zhang, X.; Liao, H. Autonomic Robotic Ultrasound Imaging System Based on Reinforcement Learning. IEEE Trans. Biomed. Eng. 2021, 68, 2787–2797. [Google Scholar] [CrossRef] [PubMed]
  161. Mulder, S.T.; Omidvari, A.H.; Rueten-Budde, A.J.; Huang, P.H.; Kim, K.H.; Bais, B.; Rousian, M.; Hai, R.; Akgun, C.; van Lennep, J.R.; et al. Dynamic Digital Twin: Diagnosis, Treatment, Prediction, and Prevention of Disease During the Life Course. J. Med. Internet Res. 2022, 24, e35675. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  162. Akpinar, M.H.; Sengur, A.; Salvi, M.; Seoni, S.; Faust, O.; Mir, H.; Molinari, F.; Acharya, U.R. Synthetic Data Generation via Generative Adversarial Networks in Healthcare: A Systematic Review of Image- and Signal-Based Studies. IEEE Open J. Eng. Med. Biol. 2024, 6, 183–192. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  163. Loni, M.; Poursalim, F.; Asadi, M.; Gharehbaghi, A. A review on generative AI models for synthetic medical text, time series, and longitudinal data. npj Digit. Med. 2025, 8, 281. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  164. Arora, A.; Arora, A. Generative adversarial networks and synthetic patient data: Current challenges and future perspectives. Future Healthc. J. 2022, 9, 190–193. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  165. Goncalves, A.; Ray, P.; Soper, B.; Stevens, J.; Coyle, L.; Sales, A.P. Generation and evaluation of synthetic patient data. BMC Med. Res. Methodol. 2020, 20, 108. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  166. Ciecierski-Holmes, T.; Singh, R.; Axt, M.; Brenner, S.; Barteit, S. Artificial intelligence for strengthening healthcare systems in low- and middle-income countries: A systematic scoping review. npj Digit. Med. 2022, 5, 162. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  167. Wahl, B.; Cossy-Gantner, A.; Germann, S.; Schwalbe, N.R. Artificial intelligence (AI) and global health: How can AI contribute to health in resource-poor settings? BMJ Glob. Health 2018, 3, e000798. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  168. Hadley, T.D.; Pettit, R.W.; Malik, T.; Khoei, A.A.; Salihu, H.M. Artificial Intelligence in Global Health -A Framework and Strategy for Adoption and Sustainability. Int. J. MCH AIDS 2020, 9, 121–127. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  169. Alami, H.; Rivard, L.; Lehoux, P.; Hoffman, S.J.; Cadeddu, S.B.M.; Savoldelli, M.; Samri, M.A.; Ag Ahmed, M.A.; Fleet, R.; Fortin, J.P. Artificial intelligence in health care: Laying the Foundation for Responsible, sustainable, and inclusive innovation in low- and middle-income countries. Glob. Health 2020, 16, 52. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  170. Vaccari, I.; Orani, V.; Paglialonga, A.; Cambiaso, E.; Mongelli, M. A Generative Adversarial Network (GAN) Technique for Internet of Medical Things Data. Sensors 2021, 21, 3726. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
Figure 1. Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) flow diagram showing the identification, screening, eligibility assessment, and inclusion of studies for this review. Records were identified through PubMed and PubMed Central (PMC) databases, with duplicates and non-eligible records removed prior to screening. Titles and abstracts of perspective articles were screened for relevance to artificial intelligence-assisted medical diagnostics, followed by assessment based on inclusion and exclusion criteria. A total of 171 studies were included in the final review.
Figure 1. Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) flow diagram showing the identification, screening, eligibility assessment, and inclusion of studies for this review. Records were identified through PubMed and PubMed Central (PMC) databases, with duplicates and non-eligible records removed prior to screening. Titles and abstracts of perspective articles were screened for relevance to artificial intelligence-assisted medical diagnostics, followed by assessment based on inclusion and exclusion criteria. A total of 171 studies were included in the final review.
Applsci 16 00728 g001
Figure 2. Summarizes the theoretical and computational foundations of AI-assisted diagnostics, showing how diverse clinical data types feed into different machine learning paradigms, computational mechanisms, and model architectures to produce interpretable and clinically actionable outputs. Created in BioRender. Roshan, D. (2025) https://BioRender.com/h132ht7 (accessed on 5 January 2026).
Figure 2. Summarizes the theoretical and computational foundations of AI-assisted diagnostics, showing how diverse clinical data types feed into different machine learning paradigms, computational mechanisms, and model architectures to produce interpretable and clinically actionable outputs. Created in BioRender. Roshan, D. (2025) https://BioRender.com/h132ht7 (accessed on 5 January 2026).
Applsci 16 00728 g002
Figure 3. Radial schematic showing the role of AI across diagnostic modalities and the cross-domain interactions that support integrated clinical decision-making. Created in BioRender. Roshan, D. (2025) https://BioRender.com/ehmctxs (accessed on 5 January 2026).
Figure 3. Radial schematic showing the role of AI across diagnostic modalities and the cross-domain interactions that support integrated clinical decision-making. Created in BioRender. Roshan, D. (2025) https://BioRender.com/ehmctxs (accessed on 5 January 2026).
Applsci 16 00728 g003
Figure 4. Overview of predictive analytics, integration of multimodal patient data, AI model outputs, and EHR-embedded decision support with validation and monitoring steps. Created in BioRender. Roshan, D. (2025) https://BioRender.com/aycchk9 (accessed on 5 January 2026).
Figure 4. Overview of predictive analytics, integration of multimodal patient data, AI model outputs, and EHR-embedded decision support with validation and monitoring steps. Created in BioRender. Roshan, D. (2025) https://BioRender.com/aycchk9 (accessed on 5 January 2026).
Applsci 16 00728 g004
Figure 5. A pyramid showing the innovation and implementation challenges for clinical AI. The lower layers, data integrity and platform reliability, form the technical foundation, while governance, clinician trust, and sustainable scale rest above. Side panels summarize the common failure modes from Table 5 and the recommended mitigations. Created in BioRender. Roshan, D. (2025) https://BioRender.com/zseg1w7 (accessed on 5 January 2026).
Figure 5. A pyramid showing the innovation and implementation challenges for clinical AI. The lower layers, data integrity and platform reliability, form the technical foundation, while governance, clinician trust, and sustainable scale rest above. Side panels summarize the common failure modes from Table 5 and the recommended mitigations. Created in BioRender. Roshan, D. (2025) https://BioRender.com/zseg1w7 (accessed on 5 January 2026).
Applsci 16 00728 g005
Figure 6. Plant schematic showing how foundational privacy, governance, and data infrastructures (roots) support emerging diagnostic AI technologies (branches) and future clinical capabilities (leaves). Created in BioRender. Roshan, D. (2025) https://BioRender.com/f3xqt41 (accessed on 5 January 2026).
Figure 6. Plant schematic showing how foundational privacy, governance, and data infrastructures (roots) support emerging diagnostic AI technologies (branches) and future clinical capabilities (leaves). Created in BioRender. Roshan, D. (2025) https://BioRender.com/f3xqt41 (accessed on 5 January 2026).
Applsci 16 00728 g006
Table 1. Summary of key representative studies included in this review, grouped together by diagnostic domain, data modality, and clinical contribution.
Table 1. Summary of key representative studies included in this review, grouped together by diagnostic domain, data modality, and clinical contribution.
Diagnostic DomainData ModalityAI Methods UsedRepresentative Included Key StudiesMain Clinical Contribution
Foundations and Early AIRule-based clinical logicExpert systemsBajwa et al. (2021) [1], Shortliffe (MYCIN, 1977) [3], Perry (1990) [4], Beam and Kohane (2018) [14]Showed the feasibility and limits of computer-assisted diagnosis
Classical ML in DiagnosticsImaging, tabular clinical dataSVM, RF, Logistic RegressionHabehh and Gohel (2021) [2], Castellino (2005) [5], Orrù et al. (2012) [6], Roy et al. (2022) [15]Feature-engineered diagnostic prediction and early CAD
Deep Learning in Medical ImagingX-ray, CT, MRICNNsEsteva et al. (2017) [7], Rajpurkar et al. (2018) [16], Aggarwal et al. (2021) [13], Litjens et al. (2017) [17], Lundervold and Lundervold (2019) [18]Demonstrated image classification performance comparable to specialists in controlled retrospective evaluations
Radiology and Its Foundation ModelsMultimodal imagingCNNs, TransformersZhou et al. (2021) [8], Kelly et al. (2019) [19], Bian et al. (2025) [20], Najjar (2023) [21],Shift from task-specific to usable multimodal models
Digital PathologyWhole-slide images (WSI)MIL, Vision TransformersCampanella et al. (2019) [22], Shafi and Parwani (2023) [23], McGenity et al. (2024) [24], Dolezal et al. (2024) [25]Automated cancer detection, grading, and prognostication
Ultrasound DiagnosticsReal-time imagingCNN–RNN hybridsShen et al. (2021) [26], Meng et al. (2019) [27]Reduced operator dependence in clinical settings
OphthalmologyFundus photography, OCTCNNsParmar et al. (2024) [28], Driban et al. (2024) [29], De Fauw et al. (2018) [30] Adaptable retinal disease screening
ECG Physiological SignalsTime-series waveformsCNNs, LSTM, TransformersHannun et al. (2019) [31], Meng et al. (2022) [32], Jaya Prakash et al. (2025) [33], Ribeiro et al. (2020) [34]These systems can produce performance that is comparable to cardiologists in arrhythmia detection
EEG Physiological SignalsTime-series signalsCNNs, DL classifiersAcharya et al. (2018) [35], Roy et al. (2019) [36], Hussein et al. (2017) [37]Seizure and neurodegenerative disease detection
EHR and Clinical TextStructured and unstructured dataRNNs, NLP TransformersChoi et al. (Doctor AI, 2016) [38], Devlin et al. (BERT, 2019) [39], Altsentzer et al. (ClinicalBERT, 2019) [40], Acharya et al. (2024) [41]Risk prediction and phenotyping
Predictive Analytics (ICU)Vitals, labs, waveformsRNNs, GRU, Ensemble MLJohnson et al. (MIMIC-III, 2016) [42], Moody et al. (PhysioNet, 2011) [43], Mao et al. (2018) [44], Kwon et al. (2018) [45]Early detection of deterioration
GenomicsNGS dataDL, TransformersAthanasopoulou et al. (2025) [46], Baião et al. (2025) [47]Variant interpretation and pathogenicity prediction
ProteomicsMass spectrometryML, DLMann et al. (2021) [48], Kitaoka et al. (2025) [49]Biomarker discovery
MetabolomicsLC–MS profilesML, VAEsChi et al. (2024) [50], Gloaguen et al. (2022) [51]Disease signature detection
Multi-Omics IntegrationGenomics, proteomics, imagingHybrid ML, generative modelsAcharya and Mukhopadhyay (2024) [52], Lin et al. (2025) [53]Precision medicine stratification
RadiogenomicsImaging and genomicsCNN + ML fusionParmar et al. (2015) [54], Wu et al. (2016) [55]Outcome and recurrence prediction
Explainable AI (XAI)All modalitiesSHAP, Grad-CAM, LIMEAlkhanbouli et al. (2025) [56], Fuhrman et al. (2022) [57], Saarela and Podgorelec (2024) [58] Model transparency and trust
Clinical Decision Support (CDS)EHR-integrated systemsML + rulesSalimparsa et al. (2025) [59], Solomon et al. (2023) [60], Patterson et al. (2019) [61]Workflow-integrated diagnostics
Validation and RegulationMultisite clinical dataSaMD frameworksHan et al. (2024) [10], Park et al. (2022) [12], Weissman (FDA, 2021) [62] Karnik (2014) [63]Clinical safety and approval pathways
Bias, Ethics and FairnessDemographically stratified dataAudit frameworksCross et al. (2024) [9], Ueda et al. (2023) [64], Nasir et al. (2025) [65], Park et al. (2025) [66]Responsible deployment of clinical AI systems
Real-World Case StudiesImaging, ECG, pathologyDeployed DL systemsDe Fauw et al. (2018) [30], Mckinney et al. (2020) [67], Attia et al. (2019) [68]Transition from lab to clinic
Table 2. Comparison of supervised, unsupervised, reinforcement, and self-supervised learning. Each ML paradigm is given a Clinical Suitability Score (1–5) after the comparison, which shows the estimated readiness of a given paradigm’s clinical diagnostic use in the real world. These scores were given based on five criteria: maturity of the paradigm in clinical literature; availability of clinically annotated data; robustness to noise and distribution shift; how it handles interpretability issues and uncertainty; and feasibility of regulatory approval and validation. A score of 5 shows a strong alignment across these criteria, while lower scores indicate increasing barriers to clinical adoption.
Table 2. Comparison of supervised, unsupervised, reinforcement, and self-supervised learning. Each ML paradigm is given a Clinical Suitability Score (1–5) after the comparison, which shows the estimated readiness of a given paradigm’s clinical diagnostic use in the real world. These scores were given based on five criteria: maturity of the paradigm in clinical literature; availability of clinically annotated data; robustness to noise and distribution shift; how it handles interpretability issues and uncertainty; and feasibility of regulatory approval and validation. A score of 5 shows a strong alignment across these criteria, while lower scores indicate increasing barriers to clinical adoption.
ML ParadigmsLabel DependencyData RequirementsComputational ComplexityUncertainty HandlingExample Diagnostic TasksLimitation in DiagnosticsClinical Suitability (Score out of 5)
Supervised Learning [16]HighLarge, labeled datasetsModerate to highSoftmax probability, calibrationPneumonia detection, ECG arrhythmia classificationRequires high-quality labels, annotation costs5/5
Weakly Supervised Learning [22]Partial labels or noisy labelsMedium to large datasetsModerateMIL, label noise correctionWhole-slide pathology classification for slide-level labels onlyLabel noise causes instability4/5
Unsupervised Learning [74]NoneMedium to large unlabeled datasetsLow to moderateClustering uncertainty metricsPatient stratification, anomaly detectionPoor clinical interpretability3/5
Self-Supervised Learning
[75]
None (pretext tasks)Very large unlabelled datasetsHighContrastive uncertainty modelingImaging pretraining, ECG representation learningRequires massive compute and data5/5
Reinforcement Learning
[76]
Reward labels onlySequence dataVery highQ-value uncertaintyImaging acquisition optimizationHard to validate clinically2/3
Generative ML [77]Low but depends on the modelVery large datasetsVery highLatent space variance estimatesSynthetic MRI, data augmentationRisk of hallucinated features3/3
Table 3. Inputs, algorithms, tasks, and clinical benefits of AI in different modalities, along with its integration difficulty score, which shows the estimated technical, infrastructural, and organizational effort required to deploy and maintain AI-based systems with existing clinical workflows. Scores were given based on five criteria: complexity of data acquisition and standardization; preprocessing and quality control requirements; computational and hardware demands; compatibility with existing clinical information systems; and validation, monitoring, and maintenance burden. A “low” integration difficulty indicates little to no disruption to existing procedures, whereas moderate to high scores indicate increasing requirements for its integration.
Table 3. Inputs, algorithms, tasks, and clinical benefits of AI in different modalities, along with its integration difficulty score, which shows the estimated technical, infrastructural, and organizational effort required to deploy and maintain AI-based systems with existing clinical workflows. Scores were given based on five criteria: complexity of data acquisition and standardization; preprocessing and quality control requirements; computational and hardware demands; compatibility with existing clinical information systems; and validation, monitoring, and maintenance burden. A “low” integration difficulty indicates little to no disruption to existing procedures, whereas moderate to high scores indicate increasing requirements for its integration.
Diagnostic ModalitiesData CharacteristicsRequired PreprocessingBest Model ClassesCommon AI Failure ModesExample Clinical ApplicationsIntegration Difficulty
Radiology (CT/MRI/X-ray) [21,86]High-resolution, continuous, spatially structuredDICOM parsing, normalization, segmentationCNNs, TransformersScanner variability drift, overfittingTumor detection, triage Moderate
Digital Pathology (WSI) [25]Ultra-high-resolution gigapixel images, multi-scaleTile extraction, stain normalizationVision Transformers, MIL networksTile selection bias, stain variabilityCancer subtype predictionHigh
Ultrasound [27]Noisy, operator-dependent, temporal–spatialSpeckle noise reduction, temporal smoothingCNN-RNN hybridsOperator variability, shadow artifactsFetal assessment, cardiac imagingHigh
ECG/EEG/Physiological signals [33]Time-series, waveform-basedFiltering, beat segmentation, artifact removal1D CNNs, LSTMsBaseline wander errors, electrode misplacementArrhythmia and seizure detectionLow
Genomics (NGS) [46]Categorical sequence data, high dimensionalityQC filtering, alignment, variant callingTransformers, ensemblesPipeline variability, batch effectsVariant interpretationHigh
Metabolomics and Proteomics [51]Mass spectra, high dimensionalityPeak picking, normalization, noise removalSVMs, VAEs, DL modelsBatch effects, ion suppression artifactsBiomarker discoveryModerate
EHR and Clinical Notes [41]Mixed structured and unstructuredTokenization, feature extraction, imputationNLP transformers, multimodal modelsMissingness, coding differencesRisk prediction, CDSSModerate
Table 4. Input fusion strategies, benchmark datasets, and real-world performance gaps in multimodal predictive models in diagnostics. The reported performance metrics show retrospective discrimination (AUROC or -index) and do not indicate calibration quality or potential clinical utilization.
Table 4. Input fusion strategies, benchmark datasets, and real-world performance gaps in multimodal predictive models in diagnostics. The reported performance metrics show retrospective discrimination (AUROC or -index) and do not indicate calibration quality or potential clinical utilization.
Clinical Prediction TaskMultimodal Input UsedFusion Strategy
(Early, Late, or Hybrid)
Benchmark Datasets UsedPerformance RangeReal-World Performance GapsMain Clinical Barrier
Sepsis onset prediction [104,105]Vitals, labs, nursing, notes, medicationsEarly, attention-based fusionMIMIC-III/IV, eICUArea under the receiver operating characteristic curve (AUROC) 0.80–0.92Drops to around 0.65–0.75 in external hospitals Shift due to different documentation patterns
Cardiac arrest [105]ECG, waveform morphology, blood gasesHybrid, CNN, and RNN fusionPhysioNet ECGDB, Telemetric datasetsAUROC 0.85–0.95Wearable vs. in-hospital data mismatchSampling frequency inconsistency
Cancer recurrence [106]MRI/CT, genomics, clinical historyLate fusion The Cancer Genome Atlas (TCGA) and institutional datasetsC-index 0.70–0.82Limited genomic completenessIntegration of omics into EHR
Stroke outcome [107]CT perfusion, National Institutes of Health Stroke Scale (NIHSS), and comorbiditiesHybrid fusionISLES, local stroke centersAUROC 0.78–0.89Scanner heterogeneityRadiology protocol variations
Mortality or deterioration in ICU [108]Vitals, ventilator data, and labsEarly fusion and gated recurrent unit (GRU)/LSTMMIMIC-IVAUROC 0.85–0.92Operational driftVariability in monitoring frequency
Table 6. The technological architectures, dependencies, risks, and clinical readiness of the upcoming diagnostic AI systems. The clinical readiness score shows the estimated maturity of a given AI technology for its deployment in real-world diagnostics. Scores were given based on four criteria: the extent of clinical validation and real-world evidence; regulatory progress or approval status; robustness, reliability, and calibration under clinical conditions; and clinician trust, usability, and governance requirements. A score of 5 is given for technologies that have already demonstrated clinical deployment and regulatory alignment, while lower scores are given due to validation, integration, or safety issues.
Table 6. The technological architectures, dependencies, risks, and clinical readiness of the upcoming diagnostic AI systems. The clinical readiness score shows the estimated maturity of a given AI technology for its deployment in real-world diagnostics. Scores were given based on four criteria: the extent of clinical validation and real-world evidence; regulatory progress or approval status; robustness, reliability, and calibration under clinical conditions; and clinician trust, usability, and governance requirements. A score of 5 is given for technologies that have already demonstrated clinical deployment and regulatory alignment, while lower scores are given due to validation, integration, or safety issues.
TechnologyMain Architecture UsedKey DependenciesDiagnostic Use CasesCurrent RisksClinical Readiness (Points out of 5)10-Year Outlook
Federated learning [156]FedAvg, FedProx, secure aggregationMulti-hospital networks, stable connectivity, DP protocolsGlobal training of imaging omics modelsNon-IID data, privacy attacks3/5The standard for multicenter research
Foundation models like Med-PaLM, LLaVA-Med [153,157]Transformer and multimodal encodersTrillion-scale tokens, aligned medical corporaGeneralist clinical reasoning, radiology QAHallucination, lack of calibration2/5Highly trusted diagnostic copilots
Generative AI (Diffusion, GANs) [158]Latent diffusion modelsHigh-quality annotated datasetsSynthetic medical imaging, rare disease augmentationFabricated features, legal/IP issues2/5Partially could replace real datasets for low-prevalence diseases
Wearable and edge AI [159]TinyML, optimized CNNs, on-chip inferenceLow-power chips, continuous data captureReal-time arrhythmia, seizure, hypoxia alerts Battery limits, sensor drifts4/5Ubiquitous real-time diagnostic
Robotics-AI integration [160]Vision transformers and reinforcement learningHigh-precision actuators, real-time sensing Robotic ultrasound, automated microscopySafety, real-time latency2/5Routine semi-autonomous diagnostic procedures
Digital twins [161]Physics-informed and ML hybridLongitudinal patient-specific dataVirtual patient simulations, treatment predictionsData sparsity, validation difficulty1/5Widely used for RCT simulation and personalized monitoring
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bartusik-Aebisher, D.; Justin Raj, D.R.; Aebisher, D. Artificial Intelligence in Medical Diagnostics: Foundations, Clinical Applications, and Future Directions. Appl. Sci. 2026, 16, 728. https://doi.org/10.3390/app16020728

AMA Style

Bartusik-Aebisher D, Justin Raj DR, Aebisher D. Artificial Intelligence in Medical Diagnostics: Foundations, Clinical Applications, and Future Directions. Applied Sciences. 2026; 16(2):728. https://doi.org/10.3390/app16020728

Chicago/Turabian Style

Bartusik-Aebisher, Dorota, Daniel Roshan Justin Raj, and David Aebisher. 2026. "Artificial Intelligence in Medical Diagnostics: Foundations, Clinical Applications, and Future Directions" Applied Sciences 16, no. 2: 728. https://doi.org/10.3390/app16020728

APA Style

Bartusik-Aebisher, D., Justin Raj, D. R., & Aebisher, D. (2026). Artificial Intelligence in Medical Diagnostics: Foundations, Clinical Applications, and Future Directions. Applied Sciences, 16(2), 728. https://doi.org/10.3390/app16020728

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop