Course Purpose
This course builds on prior knowledge of basic deep learning and convolutional neural networks. It explores advanced deep learning architectures, training strategies, and modern AI paradigms, you will be able to demonstrate applications in computer vision and other domains such as language, sequence modeling, multimodal learning, and intelligent systems.You will learn both the theoretical and practical applications of the advanced models, such as recurrent networks, generative models, attention mechanisms, transformers, metric learning, self-supervised learning, and deployment of AI systems.
Course Learning Outcomes
CLO 1: Identify the core components and mathematical foundations of advanced deep learning architectures
CLO 2: Compare different training strategies for deep learning models
CLO 3: Create models using advanced deep learning architectures
CLO 4: Evaluate the models created using deep learning architectures
Course Content
Module 1: Sequential Deep Learning Architectures: Recurrent neural networks, long short term memory (LSTM), Gated recurrent units
Module 2: Object detection and segmentation: Object detection fundamentals, Detection architectures: Faster R-CNN, SSD, YOLO, RetinaNet, Segmentation architectures: U-Net, Mask R-CNN, DeepLab
Module 3: Attention Mechanisms and transformer architecture: Self-attention and multi-head attention, Transformer architecture, Specialized Attention Variations
Module 4: Graph Neural Network: Introduction to graph structured data, graph neural networks, graph convolutional network, graph attention network, Sheaf Neural networks
Module 5: Generative modelling: Autoencoders , Variational autoencoders, GANs and GAN variants , Diffusion & Thermodynamic generative modelling
Module 6: Deep reinforcement learning: Introduction to reinforcement learning, Introduction to deep reinforcement learning, deep Q-networks, policy-based and actor-critic methods, advanced deep reinforcement learning methods, practical applications and challenges of deep reinforcement learning.
Module 7: Vision Transformers and Advanced Visual Representation Learning: Introduction to vision transformers, visual representation learning, patch embeddings, transformer encoder architecture, self-supervised and contrastive learning methods, multimodal and foundation vision models, applications and challenges of advanced visual representation learning, CLIP, Foundational Image Pretraining, Attention Rollout
Module 8: Mini project
