A Survey of Diffusion Models: Methods and Applications
Abstract
1. Introduction
1.1. Relation to Existing Surveys and Contributions
- We provide a specialized taxonomy of acceleration techniques, categorizing them into algorithmic solvers, architectural compression, and system-level lightweight paradigms.
- We bridge the gap between high-end foundation models and edge-side applications, highlighting strategies for deploying diffusion models on mobile and embedded devices.
- We offer an updated perspective on emerging architectures (e.g., Mamba/SSMs) and the convergence with LLMs, moving beyond the traditional U-Net-centric view.
1.2. Organization of the Paper
2. Basic Principles of Diffusion Models
2.1. Forward Process
2.2. Reverse Process
2.3. Training Objectives and Loss Functions
3. Methodologies
3.1. Fundamental Frameworks and Architectures
3.1.1. DDPMs (Denoising Diffusion Probabilistic Models)
3.1.2. From Discrete Markov Chains to Continuous SDEs and ODEs
3.1.3. Main Branch and Conditional Injection
3.2. Sampling Acceleration and Efficiency
3.2.1. Breakthroughs in Speed and Efficiency
3.2.2. Sampling and Acceleration
3.2.3. Training Objectives, Sampling Intervals and Weights
3.3. Controllable Generation Mechanisms
3.3.1. Conditional Generation and Guidance
3.3.2. Form and Objectives
3.3.3. Control Mechanism
3.3.4. Evaluation and Benchmarks
4. Applications
4.1. Image Restoration
4.2. Two-Dimensional Image Generation
4.3. Three-Dimensional Model/Content Generation
4.4. Video Generation and Editing
4.5. Audio Generation: From Speech to Music
5. Efficient and Lightweight Diffusion Models
- Sampling Acceleration: Focusing on advanced ODE solvers and scheduling strategies to reduce the number of inference steps from hundreds to mere dozens or single digits.
- Architectural Compression: Employing techniques such as network pruning, quantization, and structural search to minimize parameter count and memory usage.
- Knowledge Distillation: Leveraging teacher–student frameworks to condense the multi-step diffusion trajectory into fewer steps, enabling rapid inference.
5.1. Structural Efficiency and Backbone Optimization
5.2. Quantization and Frequency Domain Learning
5.3. The Generation-Augmented Lightweight Paradigm
6. Challenges and Limitations
6.1. Computational Cost and Environmental Sustainability
6.2. Intellectual Property, Copyright, and Data Provenance
6.3. Bias, Safety, and Misuse
7. Future Research Directions
7.1. Scalable Architectures: From Transformers to State Space Models
7.2. Convergence of Reasoning and Generation (LLM + Diffusion)
7.3. Towards World Simulators and Emergent Capabilities
8. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial networks. Adv. Neural Inf. Process. Syst. 2014, 27, 2672–2680. [Google Scholar] [CrossRef]
- Kingma, D.P.; Welling, M. Auto-encoding variational Bayes. arXiv 2022, arXiv:1312.6114. [Google Scholar] [PubMed]
- Dinh, L.; Sohl-Dickstein, J.; Bengio, S. Density estimation using Real NVP. arXiv 2017, arXiv:1605.08803. [Google Scholar] [CrossRef]
- Weng, L. What Are Diffusion Models? lilianweng.github.io 2021. Available online: https://lilianweng.github.io/posts/2021-07-11-diffusion-models/ (accessed on 10 December 2025).
- Sohl-Dickstein, J.; Weiss, E.A.; Maheswaranathan, N.; Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. arXiv 2015, arXiv:1503.03585. [Google Scholar] [CrossRef]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar]
- Song, J.; Meng, C.; Ermon, S. Denoising diffusion implicit models. arXiv 2022, arXiv:2010.02502. [Google Scholar] [CrossRef]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 10684–10695. [Google Scholar]
- Dhariwal, P.; Nichol, A. Diffusion models beat GANs on image synthesis. Adv. Neural Inf. Process. Syst. 2021, 34, 11237–11252. [Google Scholar]
- Ho, J.; Salimans, T. Classifier-free diffusion guidance. arXiv 2022, arXiv:2207.12598. [Google Scholar] [CrossRef]
- Zhang, L.; Rao, A.; Agrawala, M. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–3 October 2023; pp. 1743–1754. [Google Scholar]
- Hertz, A.; Mokady, R.; Tenenbaum, J.; Aberman, K.; Pritch, Y.; Cohen-Or, D. Prompt-to-prompt image editing with cross attention control. arXiv 2022, arXiv:2208.01626. [Google Scholar]
- Yang, L.; Zhang, Z.; Song, Y.; Hong, S.; Xu, R.; Zhao, Y.; Zhang, W.; Cui, B.; Yang, M.-H. Diffusion models: A comprehensive survey of methods and applications. ACM Comput. Surv. 2023, 56, 1–39. [Google Scholar] [CrossRef]
- Croitoru, F.-A.; Hondru, V.; Ionescu, R.T.; Shah, M. Diffusion models in vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10850–10869. [Google Scholar] [CrossRef]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015; Springer: Munich, Germany, 2015; pp. 234–241. [Google Scholar]
- Peebles, W.; Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Vancouver, BC, Canada, 2023; pp. 15357–15367. [Google Scholar]
- Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.; Ghasemipour, S.K.S.; Ayan, B.K.; Mahdavi, S.S.; Lopes, R.G.; et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv 2022, arXiv:2205.11487. [Google Scholar]
- Karras, T.; Aittala, M.; Aila, T.; Laine, S. Elucidating the design space of diffusion-based generative models. Adv. Neural Inf. Process. Syst. 2022, 35, 27371–27386. [Google Scholar]
- Hang, T.; Gu, S.; Li, C.; Bao, J.; Chen, D.; Hu, H.; Geng, X.; Guo, B. Efficient diffusion training via min-SNR weighting strategy. arXiv 2024, arXiv:2303.09556. [Google Scholar] [CrossRef]
- Song, Y.; Sohl-Dickstein, J.; Kingma, D.P.; Kumar, A.; Ermon, S.; Poole, B. Score-based generative modeling through stochastic differential equations. arXiv 2021, arXiv:2011.13456. [Google Scholar] [CrossRef]
- Anderson, B.D.O. Reverse-time diffusion equation models. Stoch. Processes Their Appl. 1982, 12, 313–326. [Google Scholar] [CrossRef]
- Hyvärinen, A. Estimation of non-normalized statistical models by score matching. J. Mach. Learn. Res. 2005, 6, 695–709. [Google Scholar]
- Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; Zhu, J. DPM-Solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. In Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS); Curran Associates, Inc.: New Orleans, LA, USA, 2022; pp. 5775–5787. [Google Scholar]
- Bao, F.; Nie, S.; Xue, K.; Cao, Y.; Li, C.; Su, H.; Zhu, J. All are worth words: A ViT backbone for diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 15525–15535. [Google Scholar]
- Gao, S.; Zhou, P.; Cheng, M.-M.; Yan, S. MDTv2: Masked diffusion transformer is a strong image synthesizer. arXiv 2024, arXiv:2303.14389. [Google Scholar] [CrossRef]
- Kim, D.; Lai, C.-H.; Liao, W.-H.; Murata, N.; Takida, Y.; Uesaka, T.; He, Y.; Mitsufuji, Y.; Ermon, S. Consistency trajectory models: Learning probability flow ODE trajectory of diffusion. arXiv 2024, arXiv:2310.02279. [Google Scholar] [CrossRef]
- Lipman, Y.; Chen, R.T.Q.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow matching for generative modeling. arXiv 2023, arXiv:2210.02747. [Google Scholar] [CrossRef]
- Wang, Z.; Zheng, H.; He, P.; Chen, W.; Zhou, M. Diffusion-GAN: Training GANs with diffusion. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Liu, L.; Ren, Y.; Lin, Z.; Zhao, Z. Pseudo numerical methods for diffusion models on manifolds. arXiv 2022, arXiv:2202.09778. [Google Scholar] [CrossRef]
- Zhang, M.; Cai, Z.; Pan, L.; Hong, F.; Guo, X.; Yang, L.; Liu, Z. MotionDiffuse: Text-driven human motion generation with diffusion model. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Tel Aviv, Israel, 2022; pp. 616–633. [Google Scholar]
- Zhang, H.; Cao, L.; Ma, J. Text-DiFuse: An interactive multi-modal image fusion framework based on text-modulated diffusion model. arXiv 2024, arXiv:2410.23905. [Google Scholar]
- Hsiao, Y.-T.; Khodadadeh, S.; Duarte, K.; Lin, W.-A.; Qu, H.; Kwon, M.; Kalarot, R. Plug-and-play diffusion distillation. arXiv 2024, arXiv:2406.01954. [Google Scholar]
- Mou, C.; Wang, X.; Xie, L.; Wu, Y.; Zhang, J.; Qi, Z.; Shan, Y. T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. Proc. AAAI Conf. Artif. Intell. 2024, 38, 4296–4304. [Google Scholar] [CrossRef]
- Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; Aberman, K. DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv 2023, arXiv:2208.12242. [Google Scholar]
- Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A.H.; Chechik, G.; Cohen-Or, D. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv 2022, arXiv:2208.01618. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-rank adaptation of large language models. arXiv 2021, arXiv:2106.09685. [Google Scholar]
- Vaccaro, M.; Friday, M.; Zaghi, A. Multi-agentic LLMs for personalizing STEM texts. Appl. Sci. 2025, 15, 7579. [Google Scholar] [CrossRef]
- Li, H.; Lin, Y.; He, W.; Han, W.; Xu, X.; Xu, C.; Gao, E.; Zhao, H.; Gao, X. SANTO: A coarse-to-fine alignment and stitching method for spatial omics. Nat. Commun. 2024, 15, 6048. [Google Scholar] [CrossRef]
- Han, L.; Li, Y.; Zhang, H.; Milanfar, P.; Metaxas, D.; Yang, F. SVDiff: Compact parameter space for diffusion fine-tuning. arXiv 2023, arXiv:2303.11305. [Google Scholar] [CrossRef]
- Tewel, Y.; Gal, R.; Chechik, G.; Atzmon, Y. Key-locked rank one editing for text-to-image personalization. arXiv 2024, arXiv:2305.01644. [Google Scholar]
- Huang, Y.; Huang, J.; Liu, Y.; Yan, M.; Lv, J.; Liu, J.; Xiong, W.; Zhang, H.; Cao, L.; Chen, S. Diffusion model-based image editing: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 4409–4437. [Google Scholar] [CrossRef]
- Chen, R.; Chen, Y.; Jiao, N.; Jia, K. Fantasia3D: Disentangling geometry and appearance for high-quality text-to-3D content creation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–3 October 2023; pp. 22189–22199. [Google Scholar]
- Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. Adv. Neural Inf. Process. Syst. 2018, 31, 265–277. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. Proc. Mach. Learn. Res. 2021, 139, 8748–8761. [Google Scholar]
- Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.-Y.; Ermon, S. SDEdit: Guided image synthesis and editing with stochastic differential equations. arXiv 2022, arXiv:2108.01073. [Google Scholar] [CrossRef]
- Unterthiner, T.; van Steenkiste, S.; Kurach, K.; Marinier, R.; Michalski, M.; Gelly, S. Towards accurate generative models of video: A new metric and challenges. Adv. Neural Inf. Process. Syst. 2019, 32, 11967–11977. [Google Scholar]
- Downs, L.; Francis, A.; Koenig, N.; Kinman, B.; Hickman, R.; Reymann, K.; McHugh, T.B.; Vanhoucke, V. Google scanned objects: A high-quality dataset of 3D scanned household items. arXiv 2022, arXiv:2204.11918. [Google Scholar] [CrossRef]
- Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Zhang, Q.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; et al. eDiff-I: Text-to-image diffusion models with an ensemble of expert denoisers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 18–22 June 2023; pp. 12079–12088. [Google Scholar]
- Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; Chen, M. Hierarchical text-conditional image generation with CLIP latents. arXiv 2022, arXiv:2204.06125. [Google Scholar] [CrossRef]
- Poole, B.; Jain, A.; Barron, J.T.; Mildenhall, B. DreamFusion: Text-to-3D using 2D diffusion. arXiv 2022, arXiv:2209.14988. [Google Scholar]
- Lin, C.-H.; Gao, J.; Tang, L.; Takikawa, T.; Zeng, X.; Huang, X.; Kreis, K.; Fidler, S.; Liu, M.-Y.; Lin, T.-Y. Magic3D: High-resolution text-to-3D content creation. arXiv 2023, arXiv:2211.10440. [Google Scholar]
- Agostinelli, A.; Denk, T.I.; Borsos, Z.; Engel, J.; Verzetti, M.; Caillon, A.; Huang, Q.; Jansen, A.; Roberts, A.; Tagliasacchi, M.; et al. MusicLM: Generating music from text. arXiv 2023, arXiv:2301.11325. [Google Scholar] [CrossRef]
- Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; Plumbley, M.D. AudioLDM: Text-to-audio generation with latent diffusion models. In Proceedings of the International Conference on Machine Learning (ICML), Honolulu, HI, USA, 23–29 July 2023; pp. 21450–21474. [Google Scholar]
- Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; Catanzaro, B. DiffWave: A versatile diffusion model for audio synthesis. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021. [Google Scholar]
- Fei, B.; Lyu, Z.; Pan, L.; Zhang, J.; Yang, W.; Luo, T.; Zhang, B.; Dai, B. Generative diffusion prior for unified image restoration and enhancement. arXiv 2023, arXiv:2304.01247. [Google Scholar]
- Wang, W.; Bao, J.; Zhou, W.; Chen, D.; Chen, D.; Yuan, L.; Li, H. SinDiffusion: Learning a diffusion model from a single natural image. arXiv 2022, arXiv:2211.12445. [Google Scholar] [CrossRef] [PubMed]
- Choi, J.; Kim, S.; Jeong, Y.; Gwon, Y.; Yoon, S. ILVR: Conditioning method for denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 14347–14356. [Google Scholar]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 18–22 June 2018; pp. 586–595. [Google Scholar]
- Gu, S.; Chen, D.; Bao, J.; Wen, F.; Zhang, B.; Chen, D.; Yuan, L.; Guo, B. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 21–24 June 2022; pp. 10686–10696. [Google Scholar]
- Yu, J.; Xu, Y.; Koh, J.Y.; Luong, T.; Baid, G.; Wang, Z.; Vasudevan, V.; Ku, A.; Yang, Y.; Ayan, B.K.; et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv 2022, arXiv:2206.10789. [Google Scholar]
- Chang, H.; Zhang, H.; Barber, J.; Maschinot, A.J.; Lezama, J.; Jiang, L.; Yang, M.H.; Murphy, K.; Freeman, W.T.; Rubinstein, M.; et al. Muse: Text-to-image generation via masked generative transformers. arXiv 2023, arXiv:2301.00704. [Google Scholar]
- Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; Lee, Y.J. GLIGEN: Open-set grounded text-to-image generation. arXiv 2023, arXiv:2301.07093. [Google Scholar]
- Hinz, T.; Heinrich, S.; Wermter, S. Semantic object accuracy for generative text-to-image synthesis. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 1552–1565. [Google Scholar] [CrossRef]
- Kim, S.; Jung, S.; Kim, B.; Choi, M.; Shin, J.; Lee, J. Towards safe self-distillation of internet-scale text-to-image diffusion models. arXiv 2023, arXiv:2307.05977. [Google Scholar]
- Liu, Y.; Lin, C.; Zeng, Z.; Long, X.; Liu, L.; Komura, T.; Wang, W. SyncDreamer: Generating multiview-consistent images from a single-view image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; p. TBD. [Google Scholar]
- Zheng, X.-Y.; Pan, H.; Wang, P.-S.; Tong, X.; Liu, Y.; Shum, H.-Y. Locally attentional SDF diffusion for controllable 3D shape generation. ACM Trans. Graph. 2023, 42, 1–13. [Google Scholar] [CrossRef]
- Voleti, V.; Jolicoeur-Martineau, A.; Pal, C. MCVD: Masked conditional video diffusion for prediction, generation, and interpolation. arXiv 2022, arXiv:2205.09853. [Google Scholar] [CrossRef]
- Harvey, W.; Naderiparizi, S.; Masrani, V.; Weilbach, C.; Wood, F. Flexible diffusion modeling of long videos. arXiv 2022, arXiv:2205.11495. [Google Scholar] [CrossRef]
- Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; et al. Make-A-Video: Text-to-video generation without text-video data. arXiv 2022, arXiv:2209.14792. [Google Scholar]
- Molad, E.; Horwitz, E.; Valevski, D.; Acha, A.R.; Matias, Y.; Pritch, Y.; Leviathan, Y.; Hoshen, Y. Dreamix: Video diffusion models are general video editors. arXiv 2023, arXiv:2302.01329. [Google Scholar] [CrossRef]
- Blattmann, A.; Rombach, R.; Ling, H.; Dockhorn, T.; Kim, S.W.; Fidler, S.; Kreis, K. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 22563–22575. [Google Scholar]
- Zhang, Y.; Wei, Y.; Jiang, D.; Zhang, X.; Zuo, W.; Tian, Q. ControlVideo: Training-free controllable text-to-video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 1734–1743. [Google Scholar]
- Chen, N.; Zhang, Y.; Zen, H.; Weiss, R.J.; Norouzi, M.; Chan, W. WaveGrad: Estimating gradients for waveform generation. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021. [Google Scholar]
- Liu, Y.; Sun, W. A lightweight diffusion model for image generation based on improved MobileViT. In Proceedings of the International Conference on Image Processing, Machine Learning and Pattern Recognition; ACM: Guangzhou, China, 2024; pp. 67–73. [Google Scholar]
- Cai, W.; Gao, W.; Wang, X.; Di, X. Lightweight diffusion model for camouflaged object detection. In Proceedings of the 3rd International Conference on Artificial Intelligence, Internet of Things and Cloud Computing Technology (AIoTC), Wuhan, China, 13–15 September 2024; pp. 91–96. [Google Scholar]
- Grassucci, E.; Pignata, G.; Cicchetti, G.; Comminiello, D. Lightweight diffusion models for resource-constrained semantic communication. IEEE Wirel. Commun. Lett. 2025, 14, 2743–2747. [Google Scholar] [CrossRef]
- Tao, Z.; Gao, Y.; Lin, S. Double branch lightweight finger vein recognition based on diffusion model. Int. J. Adv. Comput. Sci. Appl. 2024, 15. [Google Scholar] [CrossRef]
- An, T.; Xue, B.; Huo, C.; Xiang, S.; Pan, C. Efficient remote sensing image super-resolution via lightweight diffusion models. IEEE Geosci. Remote Sens. Lett. 2024, 21, 1–5. [Google Scholar] [CrossRef]
- Li, F.; Wu, H.; Zhang, J. Lightweight diffusion model for synthesizing malicious network traffic. In Proceedings of the IEEE National Aerospace and Electronics Conference (NAECON), Dayton, OH, USA, 15–18 July 2024; pp. 409–413. [Google Scholar]
- Gao, R.; Kang, J.; Lai, B.; Xu, M.; Sun, G.; Zhang, T.; Zhang, W.; Yang, D. High-quality trajectory generation for autonomous driving: A lightweight federated learning-based diffusion model. In Proceedings of the IEEE Global Communications Conference (GLOBECOM), Cape Town, South Africa, 8–12 December 2024; pp. 1641–1646. [Google Scholar]
- Wilms, M.; Ahsan, A.O.; Ohara, E.Y.; Dagasso, G.; Macavoy, E.; Stanley, E.A.M.; Vigneshwaran, V.; Forkert, N.D. A lightweight 3D conditional diffusion model for self-explainable brain age prediction in adults and children. In Machine Learning in Clinical Neuroimaging—7th International Workshop, MLCN 2024; Bathula, D.R., Nirmala, A.B., Dvornek, N.C., Govindarajan, S.T., Habes, M., Kumar, V., Nebli, A., Wolfers, T., Xiao, Y., Eds.; Springer: Cham, Switzerland, 2024; pp. 57–67. [Google Scholar]
- Li, L.; Song, Y.; Quan, W.; Ni, P.; Wang, K. Lightweight intent recognition method based on diffusion model. Int. J. Comput. Intell. Syst. 2024, 17, 155. [Google Scholar] [CrossRef]
- Wang, C.; He, Z.; He, J.; Ye, J.; Shen, Y. Histology image artifact restoration with lightweight transformer based diffusion model. In Artificial Intelligence in Medicine; Finkelstein, J., Moskovitch, R., Parimbelli, E., Eds.; Springer Nature Switzerland: Cham, Switzerland, 2024; pp. 81–89. [Google Scholar]
- Luccioni, A.S.; Viguier, S.; Ligozat, A.-L. Estimating the carbon footprint of Bloom, a 176B parameter language model. J. Mach. Learn. Res. 2023, 24, 1–15. [Google Scholar]
- Carlini, N.; Hayes, J.; Nasr, M.; Jagielski, M.; Sehwag, V.; Tramer, F.; Balle, B.; Ippolito, D.; Wallace, E. Extracting training data from diffusion models. In Proceedings of the USENIX Security Symposium, Anaheim, CA, USA, 9–11 August 2023. [Google Scholar]
- Shan, S.; Cryan, J.; Wenger, E.; Zheng, H.; Hanah, R.; Zhao, B.Y. Glaze: Protecting artists from style mimicry by text-to-image models. In Proceedings of the USENIX Security Symposium, Anaheim, CA, USA, 9–11 August 2023. [Google Scholar]
- Bianchi, F.; Kalluri, P.; Durmus, E.; Ladhak, F.; Cheng, M.; Nozza, D.; Hashimoto, T.; Jurafsky, D.; Zou, J.; Caliskan, A. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. arXiv 2023, arXiv:2211.03759. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef]
- Wu, C.; Yin, S.; Qi, W.; Wang, X.; Tang, Z.; Duan, N. Visual ChatGPT: Talking, drawing and editing with visual foundation models. arXiv 2023, arXiv:2303.04671. [Google Scholar] [CrossRef]





| Category | Mechanism | Training Required | Key Technique and Description |
|---|---|---|---|
| Guidance | Classifier Guidance [9] Classifier-Free Guidance [6] | No | Modifies the sampling score using gradients from a classifier or by mixing conditional/unconditional noise predictions. (CFG requires joint training but no extra classifier.) |
| Structural Injection | ControlNet [11] T2I-Adapter [33] | Yes | Injects spatial conditions (edges, depth, pose) via side-branches connected by zero-convolutions to the locked main backbone. |
| Attention Control | Prompt-to-Prompt [12] Cross-Attention Control | No | Directly manipulates the cross-attention maps during the reverse process to preserve structure while changing style or specific objects. |
| Personalized Fine-tuning | DreamBooth [34] Textual Inversion [35] LoRA [36] | Yes | Adapts the model to specific concepts/identities. LoRA uses low-rank matrices to update weights efficiently; Textual Inversion optimizes word embeddings only. |
| Domain | Sub-Task | Representative Models | Key Innovation |
|---|---|---|---|
| 2D Image | Text-to-Image | Stable Diffusion [8], DALL-E 2 [49], Imagen [17], GLIDE [5] | Introduced Latent Diffusion Models (LDMs) for efficiency; Utilized Deep Language Understanding (T5/CLIP) for semantic alignment. |
| 3D | Text-to-3D & View Synthesis | DreamFusion [50], Magic3D [51], Zero-1-to-3 [52], Fantasia3D [42] | Score Distillation Sampling (SDS) transfers 2D priors to 3D; Coarse-to-fine optimization; Disentangled geometry and appearance. |
| Audio | Text-to-Audio, Music & Speech | AudioLDM [53], MusicLM [52], DiffWave [54] | Latent diffusion on Mel-spectrograms (“Spectrogram-as-Image”); CLAP embedding conditioning; Hierarchical generation for long-form consistency. |
| Restoration | Inpainting, Deblurring & Super-Resolution | DDRM [18], Palette [17], SDEdit [45], GDP [55] | Unified frameworks for multiple degradation tasks; Solves linear inverse problems using pre-trained priors without re-training. |
| Method | Core Strategy | Architecture | Task | Performance Gains |
|---|---|---|---|---|
| MobileDiT [74] | Hybrid Architecture (ViT + Conv) | MobileViT block + adaLN-Zero | General Image Generation | FID: 2.15 on ImageNet; Beat StyleGAN-XL. |
| L-DiffCOD [75] | Structural Pruning | PVTv2-B1 | Camouflaged Object Detection | FLOPs ↓ 47.45%; Params ↓ 75%; Real-time inference on edge devices. |
| Q-GESCO [76] | Quantization (PTQ) | 8-bit PTQ | Semantic Communication | Memory ↓ 75%; Robust reconstruction under channel noise. |
| Finger Vein Net [77] | Generation-Augmented Lightweight | Dual-branch Network, E-MHSA | Finger Vein Recognition | Only 2.15 M Params; High accuracy via synthetic data augmentation. |
| LWTDM [78] | Efficient Encoder-Decoder | Cross-Attention-based Lightweight Embedding | Remote Sensing Super-Resolution | Inference steps compressed to <200; Reduced Params and MACs. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ma, H.; Wong, H.-C. A Survey of Diffusion Models: Methods and Applications. Appl. Sci. 2026, 16, 2482. https://doi.org/10.3390/app16052482
Ma H, Wong H-C. A Survey of Diffusion Models: Methods and Applications. Applied Sciences. 2026; 16(5):2482. https://doi.org/10.3390/app16052482
Chicago/Turabian StyleMa, HaoYu, and Hon-Cheng Wong. 2026. "A Survey of Diffusion Models: Methods and Applications" Applied Sciences 16, no. 5: 2482. https://doi.org/10.3390/app16052482
APA StyleMa, H., & Wong, H.-C. (2026). A Survey of Diffusion Models: Methods and Applications. Applied Sciences, 16(5), 2482. https://doi.org/10.3390/app16052482

