Course site
Welcome to STAT 940
This page will be updated throughout the term. Slides and assignments will be posted below as they are released.
Fall 2026
A graduate-level course on the foundations, architectures, optimization methods, and modern applications of deep learning.
Course at a glance
This page is the public home for the course topics, slide releases, and assignments. Enrollment-specific information, submissions, deadlines, and grades are maintained in Waterloo Learn.
From the course outline
Required textbook
Benyamin Ghojogh and Ali Ghodsi. Springer, 2026. ISBN 978-3-032-10738-1 (eBook).
View textbook| Component | Weight |
|---|---|
| Assignments | 45% |
| Paper presentation | 10% |
| Group final project | 45% |
| Total | 100% |
| Assessment | Release | Due |
|---|---|---|
| Assignment 1 | Monday, October 5 | Monday, October 19, 11:59 p.m. |
| Assignment 2 | Monday, November 2 | Monday, November 16, 11:59 p.m. |
| Assignment 3 | Monday, November 23 | Monday, December 7, 11:59 p.m. |
| Paper presentation | Eligible venues specified in the course outline | Scheduled presentation date |
| Group final project | To be announced | Monday, December 21, 11:59 p.m. |
All dates are tentative and may change. Changes will be announced.
Latest updates
Course site
This page will be updated throughout the term. Slides and assignments will be posted below as they are released.
Communication
Questions about lectures and course content should be posted on Piazza so that answers are available to the class.
Course schedule
This tentative schedule follows the lecture sequence in the course materials; textbook references provide supporting readings.
| Lecture content | Textbook alignment | Materials |
|---|---|---|
| Lecture 1 - Motivation, history, machine-learning foundations, and a neural-network overview | Ch. 1, Secs. 1-5; Ch. 2, Secs. 2.1-2.7 | Lecture 1 slides |
| Lecture 2 - Neural-network history, McCulloch-Pitts networks, the perceptron and linear separability, feedforward neural networks, and backpropagation | Ch. 1, Sec. 5; Ch. 2, Secs. 2.2-2.7; Ch. 3, Secs. 2-3 | Lecture 2 slides |
| Lecture 3 - Stochastic gradient descent, mini-batches, learning-rate tuning, momentum, optimizer selection, and first- versus second-order methods | Ch. 3, Secs. 2-5; the modern optimizer survey is supplemental | Lecture 3 slides |
| Lecture 4 - Model selection, empirical and true error, SURE, double descent, implicit regularization, and grokking | Ch. 5, Secs. 2-5; Ch. 6, Sec. 5; optional SURE proof: Ch. 5, Sec. 13 | Lecture 4 slides |
| Lecture 5 - Regularization: weight decay, augmentation and noise injection, manifold tangent classifier, early stopping, parameter sharing, label smoothing, bagging, and dropout | Ch. 5, Secs. 6-10.3; manifold tangent classifier, parameter tying/sharing, and general label smoothing are supplemental | Lecture 5 slides |
| Lecture 6 - Dropout for linear regression, batch normalization, optimization effects, and normalization alternatives | Ch. 5, Secs. 10.3-11; normalization alternatives and modern BatchNorm optimization analysis are supplemental | Lecture 6 slides |
| Lecture 7 - CNN foundations, convolution and pooling, ResNet/DenseNet, normalization, and architecture selection | Ch. 4, Secs. 1-5.8; Ch. 5, Sec. 11 | Coming soon |
| Lecture 8 - RNNs, BPTT, vanishing/exploding gradients, LSTM, and GRU | Ch. 8, Secs. 1-4.3 | Coming soon |
| Lecture 9 - Sequence-to-sequence models, attention, self-attention, and attention in vision | Ch. 9, Secs. 1-2.5 | Coming soon |
| Lecture 10 - Transformer encoder/decoder, multi-head attention, feedforward blocks, and positional encoding | Ch. 9, Sec. 3 | Coming soon |
| Lecture 11 - BERT, GPT, T5, and major large-language-model families | Ch. 11, Secs. 1-5 | Coming soon |
| Lecture 12 - LLM alignment, SFT, reward models, RLHF, DPO, and efficient Transformer variants | Ch. 11, Secs. 7-9; RL background: Ch. 18, Secs. 2 and 9; variants: Ch. 9, Sec. 4 | Coming soon |
| Lecture 13 - Performer; variational inference, ELBO, and variational autoencoders | Ch. 9, Sec. 4; Ch. 12, Secs. 1-3 | Coming soon |
| Lecture 14 - GAN foundations, conditional GANs, mode collapse, adversarial autoencoders, and applications | Ch. 14, Secs. 1-8; VAE comparison: Ch. 12, Sec. 3 | Coming soon |
| Lecture 15 - Diffusion models: DDPM forward/reverse processes, variational bound, and loss | Ch. 15, Secs. 1-2.5 | Coming soon |
| Lecture 16 - GNN foundations, graph Fourier transform, graph convolution, and ChebNet | Ch. 17, Secs. 1-3 | Coming soon |
| Lecture 17 - GCNs, neighborhood aggregation, graph learning tasks, GATs, and the Transformer connection | Ch. 17, Secs. 3-6.2 | Coming soon |
| Lecture 18 - PAC/VC generalization theory, Keras implementation, and deep reinforcement learning | Ch. 6, Secs. 2-3 and 5.1; Ch. 18, Secs. 1-10; coding context: Ch. 22, Secs. 2-3 | Coming soon |
Coursework
Assignment files and instructions will appear here. Submission links, deadlines, and grades remain in Waterloo Learn.
Course tools
Use Piazza for announcements, lecture questions, clarifications, and class discussion.
Go to PiazzaUse Learn for graded work, submissions, private course documents, deadlines, and grades.
Open Waterloo LearnSign-in may be required. Students should use the course-specific links available in their enrolled Piazza and Learn accounts.
How the course runs
Public slides are posted with the tentative topics. Additional or restricted materials are shared through Learn.
Public assignment instructions are posted on this page. Submission links, feedback, and official grades are provided through Learn.
Use Piazza for course-content questions. Use a private Learn message or email for personal matters.
Consult Learn and Piazza for authoritative dates, room information, schedule changes, and deadlines.