A Comparative Review of Artificial Intelligence Applications in Small Molecule Versus Peptide Drug Discovery
Abstract
1. Introduction
2. Foundational Divergence: Molecular Language and Data Ecosystems
| Representation | Modality | Description | Advantage | Disadvantage | Architectures |
|---|---|---|---|---|---|
| SMILES String | Small Molecule | Encodes molecular structure as a 1D ASCII string | Compact, easy to store, compatible with NLP models | Non-unique representation, fragile grammar, loses 3D information | RNN [44], Long Short-Term Memory (LSTM) [45], Transformer |
| Molecular Fingerprint | Small Molecule | Binary vector representing substructure presence | Fast computation, suitable for large-scale similarity search | Discrete, information loss (bit collision), weak similarity encoding ability | Classical ML (RF [46], SVM [47]) |
| Molecular Graph (2D/3D) | Small Molecule | Graph structure with atoms as nodes and bonds as edges | Most natural representation, rich information, preserves topological/spatial structure | Complex generation, high computational cost for 3D conformation | GNN, CNN [48] (3D) |
| Amino Acid Sequence | Peptide | 1D sequence composed of amino acid residues | Simple, intuitive, large data volume, easy to obtain | Ignores 3D structure and chemical modifications; similar sequences may have vastly different structures/functions | RNN, LSTM, Transformer |
| Sequence Embeddings (PLMs) | Peptide | Dense vectors extracted from large-scale protein language models | Captures context and evolutionary information, strong representation ability | Depends on the quality and coverage of pre-trained models | Transformer, MLP [49] |
| 3D Structure | Peptide | 3D coordinates of atoms or residues | Directly encodes function-related spatial information, high accuracy | Conformational flexibility leads to difficult representation, extremely high computational cost, data scarcity | GNN, SE(3)-Equivariant Networks [50] |
| Dataset | Modality | Description | Primary Use | Scale/Scope |
|---|---|---|---|---|
| DUDE-Z [32] | Small Molecule | Enhanced Useful Decoy Set, provides high-quality actives and property-matched decoys | Benchmark for structured virtual screening | Contains 102 targets, each with actives and 50× decoys |
| MUV [51] | Small Molecule | Maximum Unbiased Validation dataset, aims to minimize analog bias and artificial enrichment | Benchmark for ligand-based and structure-based virtual screening | Contains 17 targets, embedding actives within the chemical space of decoys |
| ChEMBL [33] | Small Molecule | Large public database containing drugs, targets, and their bioactivity data | Model training, QSAR, target discovery | Over 2 million compounds and 15 million activity data points |
| BindingDB [34] | Small Molecule | Focuses on binding affinity data for drug targets and small molecule interactions | Model training (affinity prediction) | Contains various affinity data such as Kd, Ki, and IC50 |
| PepBDB [37,38] | Peptide | Curated database of peptide–protein complex structures from PDB | Structure-based peptide design and docking studies | Provides “clean” complex structures for computational research |
| SATPdb [38] | Peptide | Structurally annotated therapeutic peptide database, integrates 22 other databases | Browsing and searching for peptides with specific structures or functions | Over 19,000 unique therapeutic peptide sequences |
| THPdb [52] | Peptide | FDA-approved therapeutic peptide and protein database | Research on marketed drug properties | Contains sequence, mechanism of action, pharmacokinetics, etc. |
| DBAASP [39] | Peptide | Database of Antimicrobial Activity and Selectivity of Peptides | Research and design of antimicrobial peptides (AMPs) | Contains a large number of AMP sequences and their activity against different microorganisms |
3. AI in Virtual Screening: A Tale of Speed Versus Flexibility
4. The Role of AI in Optimizing Structure–Activity Relationship (SAR) Performance During Lead Optimization Tasks
5. De Novo Design: Contrasting Philosophies of Property-First vs. Structure-First
6. From Bits to Beakers: AI’s Role in Synthesis Planning
7. Predicting Biological Function: The Distinct ADMET Challenges
| ADMET Prediction | Small Molecules | Peptides |
|---|---|---|
| Absorption/ Permeability | Description: Predict oral absorption-related properties, like Caco-2 permeability [110], blood–brain barrier permeability. Datasets: Caco-2 [110], BBBP [111], etc. AI Methods: QSAR models (GNN, RF, SVM). | Description: Generally poor permeability; predicting passive diffusion and active transport is a major challenge. Datasets: Data is scarce, mostly internal or from small literature sets. AI Methods: Sequence- and structure-based predictive models, still in exploratory stages. |
| Metabolism/ Stability | Description: Predict whether metabolized or inhibited by CYP450 enzymes. Datasets: Large public datasets on CYP inhibition/substrates [51,102]. AI Methods: Classification/regression models for CYP subtype specificity. | Description: The main challenge is rapid degradation by proteases, leading to a short half-life. Datasets: Protease cleavage site databases, but data is limited. AI Methods: Predict protease cleavage sites to guide sequence modification for enhanced stability. |
| Distribution | Description: Predict plasma protein binding [112] (PPB) and volume of distribution (Vd). Datasets: Public PPB datasets [113,114]. AI Methods: QSAR regression models. | Description: Distribution is limited by size and stability; targeted delivery is key. Datasets: Extremely scarce. AI Methods: No mature, general predictive models yet. |
| Toxicity (Specific) | Description: Predict hERG potassium channel inhibition (cardiotoxicity), a major cause of drug withdrawal. Datasets: Large public hERG activity datasets [115,116]. AI Methods: High-precision classification/regression QSAR models. | Description: Generally not a primary toxicity concern for peptides. |
| Toxicity (Systemic) | Description: Not a major consideration for most small molecules (except haptens). | Description: The most critical and complex safety risk is immunogenicity; prediction of T-cell/B-cell epitopes. Datasets: IEDB [117] and other immune epitope databases, but data is complex and HLA genotype-dependent. AI Methods: Prediction of peptide-MHC binding, a frontier in immunoinformatics. |
8. Conclusions and Future Perspective
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| AL | Active Learning |
| ML | Machine Learning |
| DL | Deep Learning |
| ADMET | Absorption, Distribution, Metabolism, Excretion, and Toxicity |
| PPI | Protein–Protein Interaction |
| SMILES | Simplified Molecular Input Line Entry System |
| GNN | Graph Neural Network |
| PLM | Protein Language Model |
| SA | Synthetic Accessibility |
| MOO | Multi-Objective Optimization |
| SPPS | Solid-Phase Peptide Synthesis |
| HTS | High-Throughput Screening |
| VS | Virtual Screening |
| SAR | Structure–Activity Relationship |
| QSAR | Quantitative Structure–Activity Relationship |
| RL | Reinforcement Learning |
| LLM | Large Language Model |
| MD | Molecular Dynamics |
| FEP | Free Energy Perturbation |
| XAI | Explainable Artificial Intelligence |
References
- Niazi, S.K.; Mariam, Z. Artificial intelligence in drug development: Reshaping the therapeutic landscape. Ther. Adv. Drug Saf. 2025, 16, 20420986251321704. [Google Scholar] [CrossRef] [PubMed]
- Blau, B.M.; Griffith, T.G.; Whitby, R.J. Pharmaceutical innovation and access to financial markets. PLoS ONE 2022, 17, e0278875. [Google Scholar] [CrossRef] [PubMed]
- Kim, H.; Kim, E.; Lee, I.; Bae, B.; Park, M.; Nam, H. Artificial Intelligence in Drug Discovery: A Comprehensive Review of Data-driven and Machine Learning Approaches. Biotechnol. Bioprocess Eng. 2020, 25, 895–930. [Google Scholar] [CrossRef]
- Shan, Y.; Mysore, V.P.; Leffler, A.E.; Kim, E.T.; Sagawa, S.; Shaw, D.E. How does a small molecule bind at a cryptic binding site? PLoS Comput. Biol. 2022, 18, e1009817. [Google Scholar] [CrossRef] [PubMed]
- Hopkins, A.L.; Groom, C.R. The druggable genome. Nat. Rev. Drug Discov. 2002, 1, 727–730. [Google Scholar] [CrossRef]
- Fosgerau, K.; Hoffmann, T. Peptide therapeutics: Current status and future directions. Drug Discov. Today 2015, 20, 122–128. [Google Scholar] [CrossRef]
- Muttenthaler, M.; King, G.F.; Adams, D.J.; Alewood, P.F. Trends in peptide drug discovery. Nat. Rev. Drug Discov. 2021, 20, 309–325. [Google Scholar] [CrossRef]
- Veselinovic, A.M.; Veselinovic, J.B.; Zivkovic, J.V.; Nikolic, G.M. Application of SMILES Notation Based Optimal Descriptors in Drug Discovery and Design. Curr. Top. Med. Chem. 2015, 15, 1768–1779. [Google Scholar] [CrossRef]
- Withers, C.A.; Rufai, A.M.; Venkatesan, A.; Tirunagari, S.; Lobentanzer, S.; Harrison, M.; Zdrazil, B. Natural language processing in drug discovery: Bridging the gap between text and therapeutics with artificial intelligence. Expert. Opin. Drug Discov. 2025, 20, 765–783. [Google Scholar] [CrossRef]
- Luong, K.D.; Singh, A. Application of Transformers in Cheminformatics. J. Chem. Inf. Model. 2024, 64, 4392–4409. [Google Scholar] [CrossRef]
- Ross, J.; Belgodere, B.; Chenthamarakshan, V.; Padhi, I.; Mroueh, Y.; Das, P. Large-scale chemical language representations capture molecular structure and properties. Nat. Mach. Intell. 2022, 4, 1256–1264. [Google Scholar] [CrossRef]
- Chithrananda, S.; Grand, G.; Ramsundar, B. ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction. arXiv 2020, arXiv:2010.09885. [Google Scholar]
- Weininger, D. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 1988, 28, 31–36. [Google Scholar] [CrossRef]
- Gomez-Bombarelli, R.; Wei, J.N.; Duvenaud, D.; Hernandez-Lobato, J.M.; Sanchez-Lengeling, B.; Sheberla, D.; Aguilera-Iparraguirre, J.; Hirzel, T.D.; Adams, R.P.; Aspuru-Guzik, A. Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules. ACS Cent. Sci. 2018, 4, 268–276. [Google Scholar] [CrossRef]
- Lo, A.; Pollice, R.; Nigam, A.; White, A.D.; Krenn, M.; Aspuru-Guzik, A. Recent advances in the self-referencing embedded strings (SELFIES) library. Digit. Discov. 2023, 2, 897–908. [Google Scholar] [CrossRef]
- Rogers, D.; Hahn, M. Extended-connectivity fingerprints. J. Chem. Inf. Model. 2010, 50, 742–754. [Google Scholar] [CrossRef] [PubMed]
- Zhang, O.; Lin, H.; Zhang, X.; Wang, X.; Wu, Z.; Ye, Q.; Zhao, W.; Wang, J.; Ying, K.; Kang, Y.; et al. Graph Neural Networks in Modern AI-Aided Drug Discovery. Chem. Rev. 2025, 125, 10001–10103. [Google Scholar] [CrossRef] [PubMed]
- Carracedo-Reboredo, P.; Linares-Blanco, J.; Rodriguez-Fernandez, N.; Cedron, F.; Novoa, F.J.; Carballal, A.; Maojo, V.; Pazos, A.; Fernandez-Lozano, C. A review on machine learning approaches and trends in drug discovery. Comput. Struct. Biotechnol. J. 2021, 19, 4538–4558. [Google Scholar] [CrossRef]
- Zhao, T.; Hu, Y.; Valsdottir, L.R.; Zang, T.; Peng, J. Identifying drug-target interactions based on graph convolutional network and deep neural network. Brief. Bioinform. 2021, 22, 2141–2150. [Google Scholar]
- Li, J.; Sun, Q.; Zhang, F.; Yang, B. Meta-structure-based graph attention networks. Neural Netw. 2024, 171, 362–373. [Google Scholar] [CrossRef]
- Kearnes, S.; McCloskey, K.; Berndl, M.; Pande, V.; Riley, P. Molecular graph convolutions: Moving beyond fingerprints. J. Comput.-Aided Mol. Des. 2016, 30, 595–608. [Google Scholar] [CrossRef]
- Xiong, Z.; Wang, D.; Liu, X.; Zhong, F.; Wan, X.; Li, X.; Li, Z.; Luo, X.; Chen, K.; Jiang, H.; et al. Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism. J. Med. Chem. 2020, 63, 8749–8760. [Google Scholar] [CrossRef]
- Jiang, L.; Sun, N.; Zhang, Y.; Yu, X.; Liu, X. Bioactive Peptide Recognition Based on NLP Pre-Train Algorithm. IEEE/ACM Trans. Comput. Biol. Bioinform. 2023, 20, 3809–3819. [Google Scholar] [CrossRef] [PubMed]
- Madani, A.; Krause, B.; Greene, E.R.; Subramanian, S.; Mohr, B.P.; Holton, J.M.; Olmos, J.L., Jr.; Xiong, C.; Sun, Z.Z.; Socher, R.; et al. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 2023, 41, 1099–1106. [Google Scholar] [CrossRef] [PubMed]
- Lin, Z.; Akin, H.; Rao, R.; Hie, B.; Zhu, Z.; Lu, W.; Smetanin, N.; Verkuil, R.; Kabeli, O.; Shmueli, Y.; et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023, 379, 1123–1130. [Google Scholar] [CrossRef] [PubMed]
- Elnaggar, A.; Essam, H.; Salah-Eldin, W.; Moustafa, W.; Elkerdawy, M.; Rochereau, C.; Rost, B. Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling. arXiv 2023, arXiv:2301.06568. [Google Scholar] [CrossRef]
- Elnaggar, A.; Heinzinger, M.; Dallago, C.; Rihawi, G.; Wang, Y.; Jones, L.; Gibbs, T.; Feher, T.; Angerer, C.; Steinegger, M.; et al. ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing. arXiv 2020, arXiv:2007.06225. [Google Scholar]
- Agoni, C.; Fernandez-Diaz, R.; Timmons, P.B.; Adelfio, A.; Gomez, H.; Shields, D.C. Molecular Modelling in Bioactive Peptide Discovery and Characterisation. Biomolecules 2025, 15, 524. [Google Scholar] [CrossRef]
- Oikawa, Y.; Uzawa, T.; Berenger, F.; Minagawa, N.; Yumoto, A.; Takaku, H.; Tamura, R.; Ito, Y.; Tsuda, K. GPepT: A Foundation Language Model for Peptidomimetics Incorporating Noncanonical Amino Acids. ACS Med. Chem. Lett. 2025, 16, 1670–1675. [Google Scholar] [CrossRef]
- Garzon Otero, D.; Akbari, O.; Bilodeau, C. PepMNet: A hybrid deep learning model for predicting peptide properties using hierarchical graph representations. Mol. Syst. Des. Eng. 2025, 10, 205–218. [Google Scholar] [CrossRef]
- Feller, A.L.; Secor, M.; Swanson, S.; Wilke, C.O.; Deibler, K. PeptideMTR: Scaling SMILES-Based Language Models for Therapeutic Peptide Engineering. bioRxiv 2026. [Google Scholar] [CrossRef]
- Stein, R.M.; Yang, Y.; Balius, T.E.; O’Meara, M.J.; Lyu, J.; Young, J.; Tang, K.; Shoichet, B.K.; Irwin, J.J. Property-Unmatched Decoys in Docking Benchmarks. J. Chem. Inf. Model. 2021, 61, 699–714. [Google Scholar] [CrossRef]
- Mutowo, P.; Bento, A.P.; Dedman, N.; Gaulton, A.; Hersey, A.; Lomax, J.; Overington, J.P. A drug target slim: Using gene ontology and gene ontology annotations to navigate protein-ligand target space in ChEMBL. J. Biomed. Semant. 2016, 7, 59. [Google Scholar] [CrossRef]
- Gilson, M.K.; Liu, T.; Baitaluk, M.; Nicola, G.; Hwang, L.; Chong, J. BindingDB in 2015: A public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic Acids Res. 2016, 44, D1045–D1053. [Google Scholar] [CrossRef] [PubMed]
- Kim, M.T.; Wang, W.; Sedykh, A.; Zhu, H. Curating and Preparing High-Throughput Screening Data for Quantitative Structure-Activity Relationship Modeling. Methods Mol. Biol. 2016, 1473, 161–172. [Google Scholar]
- Zhang, S.; Wu, C.; Wang, S.; Vogel, H.; Li, Y.; Yuan, S. Discovering Potent and Diverse Agonists for the beta(2)-Adrenergic Receptor via Machine Learning. ACS Pharmacol. Transl. Sci. 2025, 8, 4297–4311. [Google Scholar] [CrossRef] [PubMed]
- Wen, Z.; He, J.; Tao, H.; Huang, S.-Y. PepBDB: A comprehensive structural database of biological peptide–protein interactions. Bioinformatics 2018, 35, 175–177. [Google Scholar] [CrossRef] [PubMed]
- Singh, S.; Chaudhary, K.; Dhanda, S.K.; Bhalla, S.; Usmani, S.S.; Gautam, A.; Tuknait, A.; Agrawal, P.; Mathur, D.; Raghava, G.P. SATPdb: A database of structurally annotated therapeutic peptides. Nucleic Acids Res. 2016, 44, D1119–D1126. [Google Scholar] [CrossRef]
- Pirtskhalava, M.; Amstrong, A.A.; Grigolava, M.; Chubinidze, M.; Alimbarashvili, E.; Vishnepolsky, B.; Gabrielian, A.; Rosenthal, A.; Hurt, D.E.; Tartakovsky, M. DBAASP v3: Database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics. Nucleic Acids Res. 2021, 49, D288–D297. [Google Scholar] [CrossRef]
- Visan, A.I.; Negut, I. Integrating Artificial Intelligence for Drug Discovery in the Context of Revolutionizing Drug Delivery. Life 2024, 14, 233. [Google Scholar] [CrossRef]
- Stokes, J.M.; Yang, K.; Swanson, K.; Jin, W.; Cubillos-Ruiz, A.; Donghia, N.M.; MacNair, C.R.; French, S.; Carfrae, L.A.; Bloom-Ackermann, Z.; et al. A Deep Learning Approach to Antibiotic Discovery. Cell 2020, 181, 475–483. [Google Scholar] [CrossRef]
- Wang, X.; Gao, C.; Han, P.; Li, X.; Chen, W.; Rodriguez Paton, A.; Wang, S.; Zheng, P. PETrans: De Novo Drug Design with Protein-Specific Encoding Based on Transfer Learning. Int. J. Mol. Sci. 2023, 24, 1146. [Google Scholar] [CrossRef]
- Bae, D.; Kim, M.; Seo, J.; Nam, H. AI-guided discovery and optimization of antimicrobial peptides through species-aware language model. Brief. Bioinform. 2025, 26, bbaf343. [Google Scholar] [CrossRef] [PubMed]
- Lipton, Z.C.; Berkowitz, J.; Elkan, C. A Critical Review of Recurrent Neural Networks for Sequence Learning. arXiv 2015, arXiv:1506.00019. [Google Scholar] [CrossRef]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [PubMed]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
- Hearst, M.A.; Dumais, S.T.; Osuna, E.; Platt, J.; Scholkopf, B. Support vector machines. IEEE Intell. Syst. Their Appl. 1998, 13, 18–28. [Google Scholar] [CrossRef]
- Li, Z.; Liu, F.; Yang, W.; Peng, S.; Zhou, J. A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects. IEEE Trans. Neural Netw. Learn. Syst. 2022, 33, 6999–7019. [Google Scholar] [CrossRef]
- Taud, H.; Mas, J.F. Multilayer Perceptron (MLP). In Geomatic Approaches for Modeling Land Change Scenarios; Camacho Olmedo, M.T., Paegelow, M., Mas, J.-F., Escobar, F., Eds.; Springer International Publishing: Cham, Switzerland, 2018; pp. 451–455. [Google Scholar]
- Hu, J.; Shen, L.; Albanie, S.; Sun, G.; Wu, E. Squeeze-and-Excitation Networks. arXiv 2017, arXiv:1709.01507. [Google Scholar]
- Rohrer, S.G.; Baumann, K. Maximum Unbiased Validation (MUV) Data Sets for Virtual Screening Based on PubChem Bioactivity Data. J. Chem. Inf. Model. 2009, 49, 169–184. [Google Scholar] [CrossRef]
- Usmani, S.S.; Bedi, G.; Samuel, J.S.; Singh, S.; Kalra, S.; Kumar, P.; Ahuja, A.A.; Sharma, M.; Gautam, A.; Raghava, G.P.S. THPdb: Database of FDA-approved peptide and protein therapeutics. PLoS ONE 2017, 12, e0181748. [Google Scholar] [CrossRef]
- Li, H.; Sun, X.; Cui, W.; Xu, M.; Dong, J.; Ekundayo, B.E.; Ni, D.; Rao, Z.; Guo, L.; Stahlberg, H.; et al. Computational drug development for membrane protein targets. Nat. Biotechnol. 2024, 42, 229–242. [Google Scholar] [CrossRef]
- Eberhardt, J.; Santos-Martins, D.; Tillack, A.F.; Forli, S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings. J. Chem. Inf. Model. 2021, 61, 3891–3898. [Google Scholar] [CrossRef]
- Yang, Y.; Yao, K.; Repasky, M.P.; Leswing, K.; Abel, R.; Shoichet, B.K.; Jerome, S.V. Efficient Exploration of Chemical Space with Docking and Deep Learning. J. Chem. Theory Comput. 2021, 17, 7106–7119. [Google Scholar] [CrossRef]
- Sha, C.M.; Wang, J.; Dokholyan, N.V. NeuralDock: Rapid and Conformation-Agnostic Docking of Small Molecules. Front. Mol. Biosci. 2022, 9, 867241. [Google Scholar] [CrossRef]
- Zhang, X.; Zhang, O.; Shen, C.; Qu, W.; Chen, S.; Cao, H.; Kang, Y.; Wang, Z.; Wang, E.; Zhang, J.; et al. Efficient and accurate large library ligand docking with KarmaDock. Nat. Comput. Sci. 2023, 3, 789–804. [Google Scholar] [CrossRef]
- Abramson, J.; Adler, J.; Dunger, J.; Evans, R.; Green, T.; Pritzel, A.; Ronneberger, O.; Willmore, L.; Ballard, A.J.; Bambrick, J.; et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024, 630, 493–500. [Google Scholar] [CrossRef] [PubMed]
- Passaro, S.; Corso, G.; Wohlwend, J.; Reveiz, M.; Thaler, S.; Somnath, V.R.; Getz, N.; Portnoi, T.; Roy, J.; Stark, H.; et al. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. bioRxiv 2025. [Google Scholar] [CrossRef]
- Zhang, Y.; Sanner, M.F. AutoDock CrankPep: Combining folding and docking to predict protein-peptide complexes. Bioinformatics 2019, 35, 5121–5127. [Google Scholar] [CrossRef] [PubMed]
- Omidi, A.; Moller, M.H.; Malhis, N.; Bui, J.M.; Gsponer, J. AlphaFold-Multimer accurately captures interactions and dynamics of intrinsically disordered protein regions. Proc. Natl. Acad. Sci. USA 2024, 121, e2406407121. [Google Scholar] [CrossRef] [PubMed]
- Zalewski, M.; Wallner, B.; Kmiecik, S. Protein-Peptide Docking with ESMFold Language Model. J. Chem. Theory Comput. 2025, 21, 2817–2821. [Google Scholar]
- Medina, E.; Latham, D.R.; Sanabria, H. Unraveling protein’s structural dynamics: From configurational dynamics to ensemble switching guides functional mesoscale assemblies. Curr. Opin. Struct. Biol. 2021, 66, 129–138. [Google Scholar] [CrossRef]
- Wang, C.; Greene, D.; Xiao, L.; Qi, R.; Luo, R. Recent Developments and Applications of the MMPBSA Method. Front. Mol. Biosci. 2017, 4, 87. [Google Scholar] [CrossRef]
- Abraham, M.J.; Murtola, T.; Schulz, R.; Páll, S.; Smith, J.C.; Hess, B.; Lindahl, E. GROMACS: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX 2015, 1–2, 19–25. [Google Scholar] [CrossRef]
- Case, D.A.; Aktulga, H.M.; Belfon, K.; Cerutti, D.S.; Cisneros, G.A.; Cruzeiro, V.W.D.; Forouzesh, N.; Giese, T.J.; Götz, A.W.; Gohlke, H.; et al. AmberTools. J. Chem. Inf. Model. 2023, 63, 6183–6191. [Google Scholar] [CrossRef]
- Case, D.A.; Cerutti, D.S.; Cruzeiro, V.W.D.; Darden, T.A.; Duke, R.E.; Ghazimirsaeed, M.; Giambaşu, G.M.; Giese, T.J.; Götz, A.W.; Harris, J.A.; et al. Recent Developments in Amber Biomolecular Simulations. J. Chem. Inf. Model. 2025, 65, 7835–7843. [Google Scholar] [CrossRef]
- Loeffler, H.H.; He, J.; Tibo, A.; Janet, J.P.; Voronov, A.; Mervin, L.H.; Engkvist, O. Reinvent 4: Modern AI-driven generative molecule design. J. Cheminform. 2024, 16, 20. [Google Scholar] [CrossRef]
- Murphy, K. Reinforcement Learning: An Overview. arXiv 2024, arXiv:2412.05265. [Google Scholar]
- Ross, G.A.; Lu, C.; Scarabelli, G.; Albanese, S.K.; Houang, E.; Abel, R.; Harder, E.D.; Wang, L. The maximal and current accuracy of rigorous protein-ligand binding free energy calculations. Commun. Chem. 2023, 6, 222. [Google Scholar] [CrossRef] [PubMed]
- de Ruiter, A.; Oostenbrink, C. Advances in the calculation of binding free energies. Curr. Opin. Struct. Biol. 2020, 61, 207–212. [Google Scholar] [CrossRef] [PubMed]
- Wang, L.; Wu, Y.; Deng, Y.; Kim, B.; Pierce, L.; Krilov, G.; Lupyan, D.; Robinson, S.; Dahlgren, M.K.; Greenwood, J.; et al. Accurate and reliable prediction of relative ligand binding potency in prospective drug discovery by way of a modern free-energy calculation protocol and force field. J. Am. Chem. Soc. 2015, 137, 2695–2703. [Google Scholar] [CrossRef]
- van Pinxteren, D.J.M.; Jespers, W. Integrating Machine Learning into Free Energy Perturbation Workflows. J. Chem. Inf. Model. 2025, 65, 9856–9864. [Google Scholar] [CrossRef]
- Tosca, E.M.; Aiello, L.; De Carlo, A.; Magni, P. Pharmacometrics in the Age of Large Language Models: A Vision of the Future. Pharmaceutics 2025, 17, 1274. [Google Scholar] [CrossRef] [PubMed]
- Ovchinnikov, S.; Huang, P.S. Structure-based protein design with deep learning. Curr. Opin. Chem. Biol. 2021, 65, 136–144. [Google Scholar] [CrossRef]
- Wan, F.; Wong, F.; Collins, J.J.; de la Fuente-Nunez, C. Machine learning for antimicrobial peptide identification and design. Nat. Rev. Bioeng. 2024, 2, 392–407. [Google Scholar] [CrossRef]
- Lu, S.; Gao, Z.; He, D.; Zhang, L.; Ke, G. Data-driven quantum chemical property prediction leveraging 3D conformations with Uni-Mol. Nat. Commun. 2024, 15, 7104, Erratum in Nat. Commun. 2024, 15, 8548. [Google Scholar] [CrossRef]
- Shin, B.; Park, S.; Bak, J.; Ho, J.C. Controlled Molecule Generator for Optimizing Multiple Chemical Properties. In Proceedings of the Conference on Health, Inference, and Learning, New York, NY, USA, 8–10 April 2021; Volume 2021, pp. 146–153. [Google Scholar]
- Yang, K.; Xie, Z.; Li, Z.; Qian, X.; Sun, N.; He, T.; Xu, Z.; Jiang, J.; Mei, Q.; Wang, J.; et al. MolProphet: A One-Stop, General Purpose, and AI-Based Platform for the Early Stages of Drug Discovery. J. Chem. Inf. Model. 2024, 64, 2941–2947. [Google Scholar] [CrossRef]
- Nicolaou, C.A.; Kannas, C.; Loizidou, E. Multi-objective optimization methods in de novo drug design. Mini Rev. Med. Chem. 2012, 12, 979–987. [Google Scholar] [CrossRef] [PubMed]
- Chen, S.; Jung, Y. Estimating the synthetic accessibility of molecules with building block and reaction-aware SAScore. J. Cheminform 2024, 16, 83. [Google Scholar] [CrossRef]
- Gao, W.; Luo, S.; Coley, C.W. Generative AI for navigating synthesizable chemical space. Proc. Natl. Acad. Sci. USA 2025, 122, e2415665122. [Google Scholar] [CrossRef] [PubMed]
- Watson, J.L.; Juergens, D.; Bennett, N.R.; Trippe, B.L.; Yim, J.; Eisenach, H.E.; Ahern, W.; Borst, A.J.; Ragotte, R.J.; Milles, L.F.; et al. De novo design of protein structure and function with RFdiffusion. Nature 2023, 620, 1089–1100. [Google Scholar] [CrossRef] [PubMed]
- Dauparas, J.; Anishchenko, I.; Bennett, N.; Bai, H.; Ragotte, R.J.; Milles, L.F.; Wicky, B.I.M.; Courbet, A.; de Haas, R.J.; Bethel, N.; et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science 2022, 378, 49–56. [Google Scholar] [CrossRef]
- Rettie, S.A.; Juergens, D.; Adebomi, V.; Bueso, Y.F.; Zhao, Q.; Leveille, A.N.; Liu, A.; Bera, A.K.; Wilms, J.A.; Üffing, A.; et al. Accurate de novo design of high-affinity protein-binding macrocycles using deep learning. Nat. Chem. Biol. 2025, 21, 1948–1956. [Google Scholar] [CrossRef]
- Lipinski, C.A. Lead- and drug-like compounds: The rule-of-five revolution. Drug Discov. Today Technol. 2004, 1, 337–341. [Google Scholar] [CrossRef] [PubMed]
- Lau, J.L.; Dunn, M.K. Therapeutic peptides: Historical perspectives, current development trends, and future directions. Bioorg. Med. Chem. 2018, 26, 2700–2707. [Google Scholar] [CrossRef]
- Jumper, J.; Evans, R.; Pritzel, A.; Green, T.; Figurnov, M.; Ronneberger, O.; Tunyasuvunakool, K.; Bates, R.; Zidek, A.; Potapenko, A.; et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021, 596, 583–589. [Google Scholar] [CrossRef]
- Baek, M.; DiMaio, F.; Anishchenko, I.; Dauparas, J.; Ovchinnikov, S.; Lee, G.R.; Wang, J.; Cong, Q.; Kinch, L.N.; Schaeffer, R.D.; et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science 2021, 373, 871–876. [Google Scholar] [CrossRef]
- Tan, X.; Liu, Q.; Zhou, M.; Fang, Y.; Ouyang, D.; Zeng, W.; Dong, J. pepADMET: A Novel Computational Platform For Systematic ADMET Evaluation of Peptides. J. Chem. Inf. Model. 2026, 66, 936–946. [Google Scholar] [CrossRef]
- Chenthamarakshan, V.; Hoffman, S.C.; Owen, C.D.; Lukacik, P.; Strain-Damerell, C.; Fearon, D.; Malla, T.R.; Tumber, A.; Schofield, C.J.; Duyvesteyn, H.M.E.; et al. Accelerating drug target inhibitor discovery with a deep generative foundation model. Sci. Adv. 2023, 9, eadg7865. [Google Scholar] [CrossRef]
- Nicolaou, K.C. Organic synthesis: The art and science of replicating the molecules of living nature and creating others like them in the laboratory. Proc. Math. Phys. Eng. Sci. 2014, 470, 20130690. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Z. When doctors meet with AlphaGo: Potential application of machine learning to clinical medicine. Ann. Transl. Med. 2016, 4, 125. [Google Scholar] [CrossRef]
- Andronov, M.; Voinarovska, V.; Andronova, N.; Wand, M.; Clevert, D.A.; Schmidhuber, J. Reagent prediction with a molecular transformer improves reaction data quality. Chem. Sci. 2023, 14, 3235–3246. [Google Scholar] [CrossRef]
- Akhondi, S.A.; Rey, H.; Schworer, M.; Maier, M.; Toomey, J.; Nau, H.; Ilchmann, G.; Sheehan, M.; Irmer, M.; Bobach, C.; et al. Automatic identification of relevant chemical compounds from patents. Database 2019, 2019, baz001. [Google Scholar] [CrossRef]
- Tu, Z.; Choure, S.J.; Fong, M.H.; Roh, J.; Levin, I.; Yu, K.; Joung, J.F.; Morgan, N.; Li, S.-C.; Sun, X.; et al. ASKCOS: Open-Source, Data-Driven Synthesis Planning. Acc. Chem. Res. 2025, 58, 1764–1775. [Google Scholar] [CrossRef]
- Schwaller, P.; Probst, D.; Vaucher, A.C.; Nair, V.H.; Kreutter, D.; Laino, T.; Reymond, J.-L. Mapping the space of chemical reactions using attention-based neural networks. Nat. Mach. Intell. 2021, 3, 144–152. [Google Scholar] [CrossRef]
- Mottola, S.; Bene, A.D.; Mazzarella, V.; Cutolo, R.; Boccino, I.; Merlino, F.; Cosconati, S.; Maro, S.D.; Messere, A. Sustainable Ultrasound-Assisted Solid-Phase peptide synthesis (SUS-SPPS): Less Waste, more efficiency. Ultrason. Sonochem 2025, 114, 107257. [Google Scholar] [CrossRef] [PubMed]
- Szaniszlo, S.; Csampai, A.; Horvath, D.; Tomecz, R.; Farkas, V.; Perczel, A. Unveiling the Oxazolidine Character of Pseudoproline Derivatives by Automated Flow Peptide Chemistry. Int. J. Mol. Sci. 2024, 25, 4150. [Google Scholar] [CrossRef] [PubMed]
- Shields, B.J.; Stevens, J.; Li, J.; Parasram, M.; Damani, F.; Alvarado, J.I.M.; Janey, J.M.; Adams, R.P.; Doyle, A.G. Bayesian reaction optimization as a tool for chemical synthesis. Nature 2021, 590, 89–96. [Google Scholar] [CrossRef]
- Jeong, J.; Kim, D.; Choi, J. Application of ToxCast/Tox21 data for toxicity mechanism-based evaluation and prioritization of environmental chemicals: Perspective and limitations. Toxicol. In Vitro 2022, 84, 105451. [Google Scholar] [CrossRef]
- Wu, Z.; Ramsundar, B.; Feinberg, E.N.; Gomes, J.; Geniesse, C.; Pappu, A.S.; Leswing, K.; Pande, V. MoleculeNet: A benchmark for molecular machine learning. Chem. Sci. 2018, 9, 513–530. [Google Scholar] [CrossRef]
- Campillos, M.; Kuhn, M.; Gavin, A.C.; Jensen, L.J.; Bork, P. Drug target identification using side-effect similarity. Science 2008, 321, 263–266. [Google Scholar] [CrossRef]
- Ferguson, C.S.; Tyndale, R.F. Cytochrome P450 enzymes in the brain: Emerging evidence of biological significance. Trends Pharmacol. Sci. 2011, 32, 708–714. [Google Scholar] [CrossRef]
- Roux, G.L.; Jarray, R.; Guyot, A.C.; Pavoni, S.; Costa, N.; Theodoro, F.; Nassor, F.; Pruvost, A.; Tournier, N.; Kiyan, Y.; et al. Proof-of-Concept Study of Drug Brain Permeability Between in Vivo Human Brain and an in Vitro iPSCs-Human Blood-Brain Barrier Model. Sci. Rep. 2019, 9, 16310. [Google Scholar] [CrossRef]
- Zitron, E.; Karle, C.A.; Wendt-Nordahl, G.; Kathofer, S.; Zhang, W.; Thomas, D.; Weretka, S.; Kiehn, J. Bertosamil blocks HERG potassium channels in their open and inactivated states. Br. J. Pharmacol. 2002, 137, 221–228. [Google Scholar] [CrossRef] [PubMed]
- Manikandan, P.; Nagini, S. Cytochrome P450 Structure, Function and Clinical Significance: A Review. Curr. Drug Targets 2018, 19, 38–54. [Google Scholar] [CrossRef]
- Jepson, A.; Banya, W.; Sisay-Joof, F.; Hassan-King, M.; Nunes, C.; Bennett, S.; Whittle, H. Quantification of the relative contribution of major histocompatibility complex (MHC) and non-MHC genes to human immune responses to foreign antigens. Infect. Immun. 1997, 65, 872–876. [Google Scholar] [CrossRef] [PubMed]
- Purushotham, U. Enhanced Discussion on Reassessing Lipinski’s Rule of Five in the Era of AI-Driven Drug Discovery. Adv. Pharm. Bull. 2025, 15, 474–476. [Google Scholar] [CrossRef]
- O’Hagan, S.; Kell, D.B. The apparent permeabilities of Caco-2 cells to marketed drugs: Magnitude, and independence from both biophysical properties and endogenite similarities. PeerJ 2015, 3, e1405. [Google Scholar] [CrossRef]
- Martins, I.F.; Teixeira, A.L.; Pinheiro, L.; Falcao, A.O. A Bayesian approach to in silico blood-brain barrier penetration modeling. J. Chem. Inf. Model. 2012, 52, 1686–1697. [Google Scholar] [CrossRef] [PubMed]
- Dehghan Niestanak, V.; Unsworth, L.D. Detailing Protein-Bound Uremic Toxin Interaction Mechanisms with Human Serum Albumin in the Pursuit of Designing Competitive Binders. Int. J. Mol. Sci. 2023, 24, 7452. [Google Scholar] [CrossRef]
- Mendez, D.; Gaulton, A.; Bento, A.P.; Chambers, J.; De Veij, M.; Felix, E.; Magarinos, M.P.; Mosquera, J.F.; Mutowo, P.; Nowotka, M.; et al. ChEMBL: Towards direct deposition of bioassay data. Nucleic Acids Res. 2019, 47, D930–D940. [Google Scholar] [CrossRef]
- Wishart, D.S.; Feunang, Y.D.; Guo, A.C.; Lo, E.J.; Marcu, A.; Grant, J.R.; Sajed, T.; Johnson, D.; Li, C.; Sayeeda, Z.; et al. DrugBank 5.0: A major update to the DrugBank database for 2018. Nucleic Acids Res. 2018, 46, D1074–D1082. [Google Scholar]
- Sato, T.; Yuki, H.; Ogura, K.; Honma, T. Construction of an integrated database for hERG blocking small molecules. PLoS ONE 2018, 13, e0199348. [Google Scholar] [CrossRef]
- Karim, A.; Lee, M.; Balle, T.; Sattar, A. CardioTox net: A robust predictor for hERG channel blockade based on deep learning meta-feature ensembles. J. Cheminform. 2021, 13, 60. [Google Scholar] [PubMed]
- Mendes, M.; Mahita, J.; Blazeska, N.; Greenbaum, J.; Ha, B.; Wheeler, K.; Wang, J.; Shackelford, D.; Sette, A.; Peters, B. IEDB-3D 2.0: Structural data analysis within the Immune Epitope Database. Protein Sci. 2023, 32, e4605. [Google Scholar]
- Mumuni, F.; Mumuni, A. Explainable artificial intelligence (XAI): From inherent explainability to large language models. arXiv 2025, arXiv:2501.09967. [Google Scholar] [CrossRef]
- Bender, A.; Cortés-Ciriano, I. Artificial intelligence in drug discovery: What is realistic, what are illusions? Part 1: Ways to make an impact, and why we are not there yet. Drug Discov. Today 2021, 26, 511–524. [Google Scholar] [CrossRef] [PubMed]





| Modality | Representative Tool/Method | Core Problem Solved | AI Role |
|---|---|---|---|
| Small Molecule | AI Surrogate Model (Lean Docking), AlphaFold-Ligand | Throughput: Accelerating screening of large databases | Fast filter/ end-to-end predictor |
| Peptide | AlphaFold-Multimer, ESMFold | Flexibility: Handling the vast conformational space of peptide chains | Structure predictor (transforms docking into a folding problem) |
| Modality | Representative Tool/Method | Core Design Philosophy | AI Roles |
|---|---|---|---|
| Small Molecule | MOO frameworks (e.g., MolProphet), SynFormer | Property-First: Balance multiple abstract objectives (activity, ADMET, SA) | Multi-property optimizer/ Synthesis pathway planner |
| Peptide | RFdiffusion, RFpeptides | Structure-First: Design a functional 3D backbone that binds a target | Structural architect/Geometric problem solver |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lin, H.; Vogel, H.; Zhang, H. A Comparative Review of Artificial Intelligence Applications in Small Molecule Versus Peptide Drug Discovery. Int. J. Mol. Sci. 2026, 27, 3142. https://doi.org/10.3390/ijms27073142
Lin H, Vogel H, Zhang H. A Comparative Review of Artificial Intelligence Applications in Small Molecule Versus Peptide Drug Discovery. International Journal of Molecular Sciences. 2026; 27(7):3142. https://doi.org/10.3390/ijms27073142
Chicago/Turabian StyleLin, Han, Horst Vogel, and Huawei Zhang. 2026. "A Comparative Review of Artificial Intelligence Applications in Small Molecule Versus Peptide Drug Discovery" International Journal of Molecular Sciences 27, no. 7: 3142. https://doi.org/10.3390/ijms27073142
APA StyleLin, H., Vogel, H., & Zhang, H. (2026). A Comparative Review of Artificial Intelligence Applications in Small Molecule Versus Peptide Drug Discovery. International Journal of Molecular Sciences, 27(7), 3142. https://doi.org/10.3390/ijms27073142

