fig2
Figure 2. Overview of Commonly Used Deep Learning Network Architectures The diagram categorizes four major architectural paradigms: (Top Left) Convolutional Neural Networks (CNNs), emphasizing local feature extraction through hierarchical convolution and pooling layers, which form the basis for models like U-Net, V-Net, and ResNet; (Middle) Transformers, utilizing self-attention mechanisms and multi-head attention (MHA) to capture global contextual dependencies, represented by Vision Transformer (ViT) and Swin Transformer; (Bottom Left) Generative Adversarial Networks (GANs), consisting of a generator and a discriminator in an adversarial training loop for image synthesis; and (Bottom Right) Diffusion Models, based on a two-step stochastic process involving the progressive addition of noise (forward process) and iterative denoising (reverse process) to generate high-quality images. All CT scan samples included are anonymized clinical data from the authors’ institution and are free of copyright restrictions. Created in Adobe Illustrator. CT: Computed tomography.






