Every lesson in the Machine Learning slide course, in full text: 113 decks, 8508 slides.
Lesson 6: Gradient Intuition & PyTorch AutogradUSAAIO Lesson 6, from Week 2 on calculus, fully worked. It builds partial derivatives one axis at a time, assembles the gradient and reads it as a steepest-ascent compass, and derives the directional derivative from the projection. It then derives the three ML gradient identities - 2w, x, and 2Ax - and checks each one against autograd, covers the computational graph and the chain rule behind .backward(), and shows the accumulation and no_grad traps in real execution. It ends with a from-scratch gradient-descent project verified line by line against PyTorch. Every snippet runs as written with the seeds set, and every number in a trace was produced by real execution. The lesson runs to 60 slides.
Lesson 7: Linear Systems & the Normal EquationsUSAAIO Lesson 7, from Week 3 on linear algebra, fully worked. It writes out overdetermined systems equation by equation, then derives the normal equations XᵀXw = Xᵀy twice - once by projection and once by calculus - with no steps skipped. It solves the 5-student 2×2 system by hand, entry by entry, covers invertibility and rank, derives ridge regression from the penalized loss, and covers R². It ends with a from-scratch OLS implementation verified against sklearn. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 66 slides.
Lesson 8: Maximum Likelihood EstimationUSAAIO Lesson 8, from Week 3 on probability, fully worked. It states maximum likelihood estimation as an optimization principle and scores it on a real coin, then justifies the log trick with a live underflow demo. It derives the Gaussian mean AND variance one calculus step at a time on a fixed 8-point dataset, including the second-derivative concavity check, and covers the Bernoulli, Poisson, and Exponential MLEs. It then delivers the big reveal that MSE and cross-entropy ARE negative log-likelihoods, checking both against torch, and closes with MAP as MLE plus a log-prior, from which ridge regression falls out of a Gaussian prior. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 62 slides.
Lesson 9: The Chain Rule & BackpropagationUSAAIO Lesson 9, from Week 3 on calculus, fully worked. It builds the single-variable chain rule from local rates, then the multivariable sum-over-paths rule, then derives the Jacobian form and pins down the order in which the multiplications happen. From there it presents backpropagation as reverse-mode autodiff through a concrete two-layer sigmoid network: every forward and backward number is pushed through by hand, one beat at a time - z1, h, y-hat, L, then dy-hat, dW2, dh, sigma', dz1, dW1 - and checked against torch.autograd to machine precision. It quantifies vanishing gradients as (0.25)^depth, covers finite-difference gradient checking, and ends with a your-turn build of the whole backward pass. Every snippet runs as written, and every number came from real execution. The lesson runs to 64 slides.
Lesson 10: Eigenvalues & EigenvectorsUSAAIO Lesson 10, from Week 4 on linear algebra, fully worked. It reads the eigen-equation Av=λv geometrically, expands the characteristic polynomial det(A−λI)=0 term by term, and solves every eigenvector from (A−λI)v=0 by hand. It states the spectral theorem and verifies A=VΛVᵀ by reconstruction, derives power iteration from the eigenbasis and traces it iteration by iteration, and uses deflation to find the second eigenpair. It then builds PCA from scratch - center, covariance, eigh, sort - and matches it to sklearn. One 2×2 matrix and one 10-point dataset run through the whole deck, and every eigenvalue, trace row, and coefficient was produced by real execution. The lesson runs to 60 slides.
Lesson 100: BERT Fine-Tuning StrategiesUSAAIO Lesson 100, from Phase 3. It compares feature extraction with full fine-tuning, then covers layer-wise learning rates, two-stage domain adaptation, and catastrophic forgetting with elastic weight consolidation (EWC). A TinyBERT-like toy model - d=64, 4 layers, about 208K parameters - is used to trace all the strategies, and the EWC loss formula, the diagonal of the Fisher information, and the layer-wise learning-rate decay were verified analytically with numpy. The lesson runs to 27 slides.
Lesson 101: LoRA — Low-Rank AdaptationUSAAIO Lesson 101, from Phase 3, on LoRA - Low-Rank Adaptation - for parameter-efficient fine-tuning. It covers the low-rank decomposition delta_W=AB, initializing B to zero, the alpha-over-r scaling, which layers to apply LoRA to, and a full LoRALinear nn.Module implementation. All the parameter counts - 589,824 for a full 768×768 layer against 12,288 for rank-8 LoRA, a 48-fold reduction - and the forward-pass arithmetic were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 26 slides.
Lesson 102: RLHF & DPO — Instruction Fine-TuningUSAAIO Lesson 102, from Phase 3. It covers the full RLHF pipeline, running from supervised fine-tuning through a reward model to PPO with a KL penalty, along with the Bradley-Terry pairwise loss and the shaped reward r_env − β·KL. It then presents DPO as a closed-form alternative that eliminates the reward model entirely. Every loss value, KL divergence, and DPO margin was verified with torch 2.7.1+cpu at seed 42. The lesson runs to 26 slides.
Lesson 103: Seq2Seq Encoder-Decoder with Bahdanau AttentionUSAAIO Lesson 103, from Phase 3, on the seq2seq encoder-decoder architecture, in which a GRU encoder reads the source and a GRU decoder generates the target one token at a time. It covers Bahdanau additive attention through its energy scores, softmax alpha, and context vector, then teacher forcing for fast training convergence, exposure bias and the mismatch between training and inference, scheduled sampling to bridge that gap, and the copy mechanism as an attention-based pointer to source tokens. All the attention weights, copy probabilities, and decoder-step predictions were verified with torch 2.7.1+cpu at seed 99. The lesson runs to 31 slides.
Lesson 104: Text Decoding StrategiesUSAAIO Lesson 104, from Phase 3. It covers greedy decoding, beam search with width-k expansion, length normalization, temperature scaling, top-k masking, and top-p, or nucleus, sampling, all verified with real PyTorch log-softmax arithmetic. It proves by counter-example that greedy decoding is suboptimal, and builds a full beam-search trace table from scratch for k=2 over two steps. The lesson runs to 30 slides.
Lesson 105: BLEU, ROUGE, and BERTScoreUSAAIO Lesson 105, from Phase 3. It covers BLEU, built on n-gram precision with a brevity penalty; ROUGE, which is recall-oriented n-gram overlap, along with ROUGE-L via the longest common subsequence; and BERTScore, which uses the cosine similarity of contextual embeddings. It also covers where each metric fails. All the computations were implemented from scratch in Python and verified by real execution, giving p_1 = 5/6, p_2 = 3/5, and BLEU-4 = 0.4209 on the machine-translation example, and a ROUGE-L F1 of 0.8333 on the LCS example. The lesson runs to 25 slides.
Lesson 106: Named Entity Recognition — BERT + CRFUSAAIO Lesson 106, from Phase 3. It covers the BIO tagging scheme for named-entity recognition, a BERT encoder feeding a per-token linear head, and a CRF transition matrix that enforces valid label sequences. Viterbi decoding is traced on a three-token toy example, with the good path scoring 11.70 and the bad path 4.80, and the lesson contrasts token-level with entity-level F1 using a concrete span-mismatch example. The shapes of TinyBertNER - 81,991 parameters at hidden size 64 - were verified with torch 2.7.1+cpu. The lesson runs to 29 slides.
Lesson 107: Extractive QA — SQuAD, BERT, and Open-Domain RetrievalUSAAIO Lesson 107, from Phase 3, on extractive question answering with two linear heads on the BERT output, predicting a start and an end position. It covers why treating the two softmaxes as independent works, cross-entropy loss over token positions, and constrained span decoding, which requires start ≤ end. It then covers evaluation by exact match and token-level F1, generative QA with T5, and open-domain QA, which retrieves with DPR and then reads and extracts. All the numbers were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 28 slides.
Lesson 108: Summarization, BART, Perplexity & CalibrationUSAAIO Lesson 108, from Phase 3. It contrasts extractive with abstractive summarization, then covers BART's denoising pretraining - text infilling, token deletion, and sentence permutation - and Pegasus's sentence-level masking. It defines perplexity as exponentiated cross-entropy and traces the log-probabilities concretely on a three-token toy language model, giving 1.26 for the good model, 1.42 for the bad one, and 3 for the uniform one. A TinyBART encoder-decoder of 22,996 parameters is then fine-tuned on a copy task, taking perplexity from 11.02 to 1.42 over 50 epochs. It closes with calibration and expected calibration error, using a five-bin reliability table in which the calibrated model scores 0.0296 against 0.1862 for the overconfident one. All the numbers were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 31 slides.
Lesson 109: Object Detection & YOLOUSAAIO Lesson 109, from Phase 3. It covers how a bounding box is parameterized, in both corner and centre-offset formats, then builds IoU from scratch and covers non-maximum suppression. It works the YOLO grid and anchor arithmetic - S=13, B=5, C=80, giving 17,745 raw predictions - and decodes anchor offsets with sigmoid and exp. It then covers Feature Pyramid Networks for multi-scale detection and the calculation of mAP at 0.5. Every number was verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 27 slides.
Lesson 11: Convex OptimizationUSAAIO Lesson 11, from Week 4 on optimization, fully worked. It proves the chord definition of convexity on a running quadratic, then derives the first-order tangent condition and the second-order condition that the Hessian be positive semi-definite, step by step, and dissects a non-convex counterexample. It shows why every stationary point of a convex function is global, writes gradient descent as a contraction with the update factor derived from scratch, proves the stable-step bound eta < 2/L, and covers the speed limit set by the condition number along with the classes of convergence rate. It ends with a from-scratch gradient-descent experiment verified against real execution. Every number was produced by running the code. The lesson runs to 61 slides.
Lesson 110: Semantic Segmentation & U-NetUSAAIO Lesson 110, from Phase 3, on pixel-wise classification through semantic segmentation. It covers the U-Net encoder-decoder with its skip connections, which concatenate rather than add, and the upsampling methods available: bilinear interpolation, ConvTranspose2d, and pixel shuffle. It then covers Dice loss and Focal loss for class-imbalanced masks, and trains a TinyUNet from scratch. All the shapes, loss values, and training traces were verified with torch 2.7.1+cpu. The lesson runs to 32 slides.
Lesson 111: Transformers, NLP & CV — Phase 3 ReviewUSAAIO Lesson 111, the Phase 3 capstone review. It derives scaled dot-product attention and analyzes its complexity, covers sinusoidal positional encoding and BPE tokenization step by step, and implements and debugs multi-head attention. It then covers the IoU formula and its computation, the U-Net skip-connection architecture, and a comparison of ViT with CNNs. All the values were computed with torch 2.7.1+cpu and numpy 2.2.6 on synthetic data. The lesson runs to 28 slides.
Lesson 112: Instance Segmentation & Contrastive LearningUSAAIO Lesson 112, from Phase 3. Mask R-CNN adds a per-proposal mask head to Faster R-CNN, running RoI features through a conv stack and a ConvTranspose2d to produce per-class binary masks; panoptic segmentation unifies the semantic and instance predictions; and self-supervised contrastive learning, in the form of SimCLR with InfoNCE, trains representations without labels by pulling augmented views of the same image together while pushing different images apart. The InfoNCE loss, NT-Xent, the temperature tau, the design of the projection head, and linear-probe evaluation were all verified with torch 2.7.1+cpu on synthetic data. The lesson runs to 29 slides.
Lesson 113: CLIP — Contrastive Language-Image PretrainingUSAAIO Lesson 113, from Phase 3. It covers the symmetric InfoNCE contrastive loss, the dual-encoder CLIP architecture with its image and text towers, the learned temperature parameter, zero-shot transfer through text-based class descriptors, and the CLIP extensions ALIGN and BLIP-2. A TinyCLIP is built from scratch with torch 2.7.1+cpu, and all the logit values, softmax outputs, and loss numbers - InfoNCE at 1.2648, the perfect case at 0.0000, and the random case at log(4) = 1.3863 - were verified by real execution. The lesson runs to 25 slides.
Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer VisionUSAAIO Lesson 114, a Phase 3 review. It is a timed mock-exam simulation covering all the Phase 3 topics: deriving scaled dot-product attention, counting the parameters of multi-head attention, causal masking, sinusoidal positional encoding, the BERT architecture and how to fine-tune it, LoRA adapters, greedy against beam-search decoding, computing BLEU-4 and ROUGE-N, non-maximum suppression with IoU, and the cross-entropy and perplexity of a full transformer training run. Every value in a trace table was computed with torch 2.7.1+cpu, and each section comes with a worked answer-key walkthrough. The lesson runs to 27 slides.
Lesson 115: Transformer End-to-End ReviewUSAAIO Lesson 115, from Phase 3, a complete transformer review. It runs from scaled dot-product attention through multi-head attention, the encoder block - multi-head attention, feed-forward network, layer norm, and residual - and the decoder block, with its masked self-attention, cross-attention, and feed-forward network. It then covers all the major variants, BERT, GPT, T5, ViT, and GNNs, comparing their architectures, and analyzes the O(n²d) complexity, causal masking, sinusoidal positional encoding, and the role of the CLS token. It ends with a PixelBERT implemented from scratch and verified with torch 2.7.1+cpu on the sklearn digits dataset. The lesson runs to 27 slides.
Lesson 116: NLP + CV Review & Exam SimulationUSAAIO Lesson 116, a Phase 3 review. It is a comprehensive exam-simulation deck across NLP and computer vision, covering tokenization, word2vec, BERT, LoRA, seq2seq, beam search, BLEU, and ROUGE on the language side, and convolution, ResNet, ViT, U-Net, YOLO, IoU, NMS, and CLIP on the vision side. All the metrics were computed by real Python execution with torch 2.7.1+cpu and numpy 2.2.6. It emphasizes the cross-modal connections that transformers make, pattern recognition for exam pace, and spotting the gaps you need to close before Phase 4. The lesson runs to 30 slides.
Lesson 117: Phase 3 Gap Analysis & Phase 4 PreviewUSAAIO Lesson 117, at the end of Phase 3. It is a structured gap analysis across all nine Phase 3 topic clusters - transformers, BERT and GPT, ViT, GNNs, NLP, computer vision, seq2seq, BLEU and ROUGE, and positional encoding - with a scored self-assessment and a way to adjust your study plan for the weak areas. It then previews Phase 4, covering the VAE and its ELBO, the GAN minimax objective, and the DDPM noise schedule, and closes on exam strategy: allocating your time, approaching a derivation, and mindset. All the math was verified by real Python execution with numpy 2.2.6 and torch 2.7.1+cpu. The lesson runs to 29 slides.
Lesson 118: Autoencoders — Standard, Denoising & SparseUSAAIO Lesson 118, from Week 40 of Phase 4 on generative AI. It presents the autoencoder as an encoder feeding a bottleneck latent z feeding a decoder, trained on reconstruction loss alone, and covers undercomplete compression. It then covers the denoising autoencoder, which corrupts the input and reconstructs the clean signal, and the sparse autoencoder, which puts an L1 penalty on the activations and is the workhorse of mechanistic interpretability. It covers interpolation in latent space and ends on the key limitation that motivates the VAE: a standard autoencoder's latent space has no prior, so you cannot sample from it. You build a convolutional autoencoder yourself. The lesson runs to 26 slides.
Lesson 12: MAP, Priors & RegularizationUSAAIO Lesson 12, from Week 4 on probability, fully worked. It derives MAP from Bayes' rule as MLE plus a log-prior, showing every logarithmic step, then turns a Gaussian prior into L2 or ridge and a Laplace prior into L1 or lasso, term by term. It gives the exact reason L1 zeros out coordinates, through the subgradient and the soft-threshold, computed one coordinate at a time, proves Beta-Binomial conjugacy by proportionality, and covers dropout as approximate Bayesian inference. It ends with a ridge-against-lasso sparsity project verified against scikit-learn. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 61 slides.
Lesson 13: Positive (Semi)Definite MatricesUSAAIO Lesson 13, from Week 5 on linear algebra, fully worked. It builds the quadratic form xᵀAx entry by entry, defines positive semi-definite and positive definite, and connects both to eigenvalues by completing the square. The running matrix A = [[2,1],[1,2]] is tested five independent ways - by the definition, by the completed square, by its eigenvalues, by Sylvester's minors, and by Cholesky - and contrasted against the singular M = [[4,2],[2,1]] and the indefinite N = [[1,2],[2,1]]. It then proves that XᵀX and the covariance matrix are PSD from ‖Xx‖² ≥ 0, shows ridge regression lifting every eigenvalue, and ends with a from-scratch definiteness checker that is cross-validated. Every snippet runs standalone, and every number comes from real execution. The lesson runs to 60 slides.
Lesson 14: Information TheoryUSAAIO Lesson 14, from Week 5 on probability, fully worked. It builds surprise and entropy from a single axiom and computes the entropy of a running three-outcome weather distribution bit by bit. It then derives KL divergence as an extra coding cost and proves it asymmetric and non-negative, splits cross-entropy into H(P) + KL one algebra step at a time, derives mutual information from the joint table, and gets to the ELBO decomposition that underpins the VAE. Along the way it covers units, bits against nats, conditional entropy, and the closed form for the Gaussian KL. Every snippet runs standalone, and every number was produced by real execution. The lesson runs to 60 slides.
Lesson 15: Modern OptimizersUSAAIO Lesson 15, from Week 5 on optimization, fully worked. It starts by showing why plain gradient descent crawls on an ill-conditioned bowl with a condition number of 100, then derives momentum, RMSprop, and Adam and traces each of them step by step on ONE running problem, f(x) = ½(100x₀² + x₁²). Every update rule is walked out by hand, and bias correction is dissected at t=1, where the raw step is 3.16 times too big while the corrected one is exactly 1.0. All four optimizers are then raced with real loss numbers, and AdamW's decoupled weight decay and learning-rate warmup are derived. In the your-turn project you build momentum, RMSprop, and Adam from scratch. Every snippet runs standalone, and every number came from real numpy execution. The lesson runs to 61 slides.
Lesson 16: Singular Value DecompositionUSAAIO Lesson 16, from Week 6 on linear algebra, fully worked. It derives the SVD A=UΣVᵀ from AᵀA with no steps skipped, identifies the singular values as the non-negative roots of the eigenvalues of AᵀA, and proves the rotate-scale-rotate geometry on a running 2×3 example. It then computes low-rank approximation and Eckart-Young entry by entry, works a real compression on a decaying spectrum, re-derives PCA as the SVD of centered data, and builds the pseudoinverse A⁺=VΣ⁺Uᵀ by hand, matching it to np.linalg.pinv. Every snippet runs standalone, and every number came from real execution. The lesson runs to 63 slides.
Lesson 17: Loss Functions & Their GradientsUSAAIO Lesson 17, from Week 6, fully worked. It explains what a loss is, derives MSE and MAE term by term along with their gradients, and builds softmax and cross-entropy from scratch. It then proves the elegant result dL/dz = p - y TWICE, once by the softmax-Jacobian chain rule for the multi-class case and once directly for sigmoid with BCE, with no steps skipped. From there it explains why MSE stalls a classifier, through the vanishing p(1-p) factor, covers hinge loss and subgradients, gives a loss-to-gradient cheat sheet, and ends with a from-scratch build verified against torch autograd. One running example threads the whole deck - the regression y=[3,5,7] against f=[2,6,4], and the logits z=[1,2,3] with true class 2 - and every number was produced by real execution. The lesson runs to 60 slides.
Lesson 18: Full BackpropagationUSAAIO Lesson 18, from Week 6, fully worked. One tiny 3-2-2 network with fixed weights is carried end to end. The forward pass is built operation by operation - affine, ReLU, affine, softmax, cross-entropy - and then the complete backward pass is derived one gradient per beat, covering the softmax-and-cross-entropy shortcut, the outer-product weight gradient, and the ReLU mask that kills a dead unit. Every number is matched against PyTorch autograd and against a finite-difference gradient check. It then covers vanishing and exploding gradients with clip-by-norm and He initialization, and caps off by training an MLP to about 97% on the sklearn digits dataset, along with a from-scratch NumPy network. Every snippet runs standalone, and every number was produced by real execution. The lesson runs to 60 slides.
Lesson 19: PCA ImplementationUSAAIO Lesson 19, from Week 7, fully worked. PCA is derived and implemented end to end on one hand-computable five-point 2D dataset. Centering is shown row by row, the 2x2 covariance matrix is built entry by entry, its eigenvalues are found from the characteristic polynomial with the quadratic formula, and the top eigenvector is solved by hand. The SVD route is done in parallel and proven identical via sigma^2 = (n-1)*lambda. It then covers the explained-variance ratio and the cumulative curve, choosing k at a 95% threshold on the 64-feature digits data, which gives k=29, the 2D projection, reconstruction error, and the limitation that PCA is linear only. It ends with a from-scratch PCA class verified equal to sklearn. Every snippet runs standalone in a fresh interpreter, and every number came from real execution. The lesson runs to 60 slides.
Lesson 20: Hypothesis TestingUSAAIO Lesson 20, from Week 7 on statistics, fully worked. It covers H₀ and H₁ and Type I and Type II errors, derives the p-value as P(data|H₀), and dismantles the misreading of it as P(H₀|data). It then builds a one-sample t-test from the sample mean and the ddof=1 sample standard deviation, through the t₉ tail, to p = 0.0282, which rejects, and verifies it against scipy. From there it covers critical values, the t, chi-squared, and F family with a worked chi-squared test, power against sample size, the 85-against-83 two-proportion test, and multiple testing with Bonferroni on seeded noise. Every snippet runs standalone, and every number came from real execution. The lesson runs to 61 slides.
Lesson 21: Second-Order MethodsUSAAIO Lesson 21, from Week 7, fully worked. It builds the second-order Taylor model term by term, covers the Hessian and how its eigenvalues classify minima, maxima, and saddles, and derives Newton's update by setting the model's gradient to zero. It shows why Newton is exact on a quadratic, taking one step from [5,-3] to [0.2,0.4], and why it converges quadratically, with errors falling 5.9e-1, 3.4e-2, 1.3e-3, 1.7e-6, 3.3e-12. It then races IRLS Newton for logistic regression against gradient descent - 7 iterations against 1147, reaching the same w, matched to sklearn - and covers the O(d^3) per-step cost, the saddle trap that an indefinite Hessian creates, and L-BFGS and the natural gradient at scale. Every snippet runs standalone, and every number came from real execution. The lesson runs to 61 slides.
Lesson 22: Affine TransformsUSAAIO Lesson 22, from Week 8, fully worked. It defines the affine map T(x)=Wx+b and separates it into its linear and translation parts, tests the linearity axioms one at a time, and reads the geometric transforms - rotation, reflection, scaling, and shear - off W. It proves the affine invariants for lines, parallelism, and ratios, then derives the composition T2 of T1 move by move to (W2W1)x + (W2b1 + b2), and gives the homogeneous-coordinate trick that turns Wx+b into a single matrix. It then shows a deep linear stack collapsing into one matrix, three layers at a time, and why a ReLU breaks that collapse, since additivity fails on real numbers. It settles XOR once and for all: the best linear fit is stuck at 0.5 while a tanh MLP reaches 1.0, alongside a hand-built two-ReLU XOR. Every snippet runs standalone, and every number came from real execution with numpy 2.2.6 and torch 2.7.1, seeded. The lesson runs to 61 slides.
Lesson 23: The Multivariate GaussianUSAAIO Lesson 23, from Week 8, fully worked. It builds the N(mu, Sigma) density term by term from the one-dimensional bell curve and proves from Cov(a^T x) why Sigma must be symmetric and PSD - and positive definite if it is to be inverted. It then covers the covariance ellipse from the eigendecomposition, derives and hand-computes the Mahalanobis distance, and covers the normalizing constant and the determinant. It derives the marginals and the bivariate conditional entry by entry, proves that a diagonal Sigma implies independence by factorizing the density, and gives the X^2 counterexample. It closes with Cholesky sampling, x = mu + Lz, proving Cov(mu + Lz) = Sigma line by line, verifying against 200k samples, and tying it to the VAE reparameterization trick. Every snippet runs standalone, and every number was produced by real execution. The lesson runs to 63 slides.
Lesson 24: Constrained Optimization & KKTUSAAIO Lesson 24, from Week 8, fully worked, building constrained optimization from the ground up. One running toy problem - minimizing x² + y² on a line - carries the derivation of the Lagrangian gradient by gradient, and the multiplier is revealed as a shadow price. Inequality constraints then give the four KKT conditions, with complementary slackness proved on actual numbers. A four-point one-dimensional SVM is solved both as a primal problem and as a dual QP, giving α = [0, 0.5, 0.5, 0] with support vectors at x = 2 and x = 4, and the lesson closes strong duality and Slater's condition by matching the dual optimum to the primal. Every scipy and sklearn snippet runs standalone, and every number came from real execution. The lesson runs to 62 slides.
Lesson 25: Projections & the Hat MatrixUSAAIO Lesson 25, from Week 9, fully worked. It derives orthogonal projection matrices from P²=P and Pᵀ=P, then builds the hat matrix H=X(XᵀX)⁻¹Xᵀ from the normal equations one step at a time and proves that ŷ=Hy equals the OLS fit. It checks idempotence and symmetry algebraically, covers the eigenvalues of 0 and 1 and the fact that tr(H)=p, and introduces the residual projector I−H. It then derives leverage hᵢᵢ both from diag(H) and from the closed form 1/n+(xᵢ−x̄)²/Sxx, builds Cook's distance entry by entry, and runs a leave-one-out demo proving that influence is leverage times residual. Every snippet runs standalone, and every number came from real numpy execution. The lesson runs to 62 slides.
Lesson 26: Evaluation MetricsUSAAIO Lesson 26, from Week 9, fully worked on one running 12-email spam classifier and one six-house regression set. It builds the confusion matrix count by count, derives precision, recall, and F1 and computes them by hand, and explains why accuracy lies on imbalanced data. It sweeps the precision-recall trade-off across thresholds, derives ROC and AUC as a ranking probability - 31 of 35 pairs, by hand - and confirms it against sklearn, then covers AUC-PR and log-loss with a confident-but-wrong blow-up. It then computes MSE, RMSE, MAE, R2, and MAPE entry by entry, splits RMSE from MAE with an outlier, and gives each generation metric - perplexity, BLEU, ROUGE, and FID - a runnable toy example. It ends with a from-scratch metrics library verified against sklearn. Every number was produced by real execution. The lesson runs to 61 slides.
Lesson 27: Numerical StabilityUSAAIO Lesson 27, from Week 9, fully worked to olympiad depth. It covers the finite precision of IEEE-754 - overflow, underflow, and cancellation - with the exact float64 thresholds, then proves the shift-invariance of softmax algebraically. It derives a stable softmax and the log-sum-exp trick one move per beat and traces them by hand on z = [1000, 1001, 1002], then makes log-softmax and cross-entropy overflow-proof. It goes on to the condition number, an ill-conditioned Hilbert solve, and ridge or Tikhonov regularization lifting the smallest eigenvalue, then float32 against float64 and mixed precision. It ends with a build-it project verified line by line against numpy, scipy, and torch. Every snippet runs standalone in a fresh interpreter, and every number was produced by real execution. The lesson runs to 64 slides.
Lesson 28: Kernel MethodsUSAAIO Lesson 28, from Week 10, fully worked. It builds feature maps out on XOR, then gives the kernel trick k(x,z)=φ(x)·φ(z), deriving the polynomial kernel (x·z+c)² term by term and matching it to an explicit six-dimensional φ. It applies Mercer's theorem - a kernel is valid exactly when its Gram matrix is PSD - to accept the RBF kernel and reject a fake one by its negative eigenvalue, then shows that the RBF kernel corresponds to an infinite-dimensional feature space, with the series verified numerically. It covers the composition rules for sums, products, and positive scalings, and closes with kernelized ridge regression, where the dual equals the primal, and an RBF SVM that cracks XOR. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 61 slides.
Lesson 29: Mutual Information & Contrastive LearningUSAAIO Lesson 29, from Week 10, fully worked. It recaps entropy from scratch, then derives mutual information I(X;Y) THREE equivalent ways, one algebra move at a time: as H(X)+H(Y)-H(X,Y), in the chain-rule form H(Y)-H(Y|X), and in the KL form D(p(x,y)||p(x)p(y)). All three are computed by hand on a single running 2x2 joint distribution and confirmed by real code. It then covers pointwise mutual information, proves the bounds I>=0 and I<=min(H), and works the trap in which correlation is zero yet the variables are dependent, on Y=X^2. It closes with mutual-information feature selection in sklearn, the InfoNCE and CLIP loss traced on a tiny batch, the link to compression, and a from-scratch mutual-information project. Every snippet runs standalone, and every number came from real execution. The lesson runs to 61 slides.
Lesson 30: Support Vector MachinesUSAAIO Lesson 30, from Week 10, fully worked. It derives the geometric margin as 2/‖w‖, sets up the primal max-margin QP, and writes out the Lagrangian and the KKT conditions. It then derives the dual QP in α term by term and solves that dual BY HAND on a four-point set, giving α=[0,¼,¼,0], explains why only the support vectors survive, and rebuilds the decision function from the duals. It measures the soft-margin C trade-off and closes with the kernel trick on XOR. One running four-point dataset carries the whole deck; every snippet runs standalone, and every number came from real execution. The lesson runs to 63 slides.
Lesson 31: t-SNE & UMAPUSAAIO Lesson 31, from Week 11, fully worked. It starts from PCA's linear wall and the manifold hypothesis, then builds t-SNE from the ground up on a four-point running example: squared distances, the conditional similarities p_{j|i} with a per-point sigma, symmetrization into the joint P, perplexity as 2 to the power of the entropy, the Student-t low-dimensional kernel Q derived and contrasted with the Gaussian that caused crowding, and the KL(P||Q) objective with its gradient. Every P, Q, and KL number is recomputed by hand and reproduced from scratch in NumPy. UMAP is then contrasted as a graph method, PCA and t-SNE are run on the digits dataset - 0.61 against 0.98 by KNN - and the no-transform and geometry traps are proven in code. Every snippet runs standalone, and every number came from real execution. The lesson runs to 60 slides.
Lesson 32: Bias-Variance & Cross-ValidationUSAAIO Lesson 32, from Week 11, fully worked. It derives the bias-variance decomposition MSE = Bias² + Variance + Noise line by line, adding and subtracting the mean and showing every cross term vanish, then verifies it numerically on a toy averaging estimator and measures it empirically across polynomial degree on a fixed sin(1.5x) running example, so that the U-shaped test error emerges from real numbers. It then builds cross-validation from the ground up: k-fold written from scratch with a full 5-fold trace and matched to sklearn, along with the stratified, leave-one-out, and time-series variants, the leakage trap, and nested cross-validation. Every snippet runs standalone, and every number came from real execution. The lesson runs to 62 slides.
Lesson 33: Gradient Checking & DebuggingUSAAIO Lesson 33, from Week 11, fully worked. It derives the forward difference and its O(ε) error from Taylor's theorem, then the central difference and its O(ε²) error term by term, with the odd terms cancelling, and works the truncation-against-round-off trade-off that sets ε at about 1e-5. It builds the relative-error gradient test from scratch on one running logistic-regression model, deriving the analytic gradient Xᵀ(p−y)/N, gradient-checking it to a relative error of 1.9e-10, matching it against torch.autograd to 1e-16, and then breaking it two ways - a sign flip giving 0.46, and a missing 1/N giving exactly 0.60 - both of which the check catches. The second half is real-model debugging: broadcasting that turns (n,) − (n,1) into (n,n), reducing along the wrong axis, a forgotten zero_grad accumulating 2, 4, 6, the sigmoid derivative peaking at 0.25, per-layer gradient norms diagnosing vanishing and exploding gradients, clipping, the unstable-sigmoid overflow, and NaN hunting with clamping and detect_anomaly. It ends with a build-your-own gradient_check project. Every snippet runs standalone, and every number came from real numpy and torch execution. The lesson runs to 60 slides.
Lesson 34: Concentration Inequalities & GeneralizationUSAAIO Lesson 34, from Week 12 on generalization theory, fully worked. It derives Markov's inequality from the layer-cake, or indicator, trick with no steps skipped, obtains Chebyshev's by applying Markov to the squared deviation, and states Hoeffding's, proving its exponential decay in n empirically on a biased coin. It then solves the sample-complexity inversion n >= ln(2/delta)/(2 eps^2) by hand. From there it builds up shattering and VC dimension, from a one-dimensional threshold to a two-dimensional line, which shatters 3 points but fails XOR, managing at best 3 of 4. It covers Sauer's lemma and the VC generalization bound term by term, and demonstrates double descent with a real least-norm fit whose test error peaks at the interpolation threshold. One running example, a coin with p=0.3, threads the whole deck, and every inequality and every printed number was produced by real execution. The lesson runs to 62 slides.
Lesson 35: ROC, PR Curves & CalibrationUSAAIO Lesson 35, from Week 12, fully worked. It defines the confusion matrix and every rate derived from it, then builds the ROC curve threshold by threshold on one running 10-example dataset. AUC is computed three independent ways - by the trapezoid rule, by sklearn, and by counting concordant pairs - and all three land on 0.875. The precision-recall curve and average precision are derived the same way, and the lesson shows why ROC-AUC misleads under 2% imbalance while PR-AUC, at 0.111, tells the truth. It then defines calibration and the Expected Calibration Error bin by bin, and sweeps temperature scaling to its ECE minimum at T=3. Every snippet runs standalone in a fresh interpreter, and every number was produced by real execution. The lesson runs to 61 slides.
Lesson 36: Initialization & Gradient ClippingUSAAIO Lesson 36, from Week 12, fully worked, on why a deep network lives or dies at step zero. It presents exploding and vanishing gradients as a product of per-layer factors, derives gradient norm clipping and contrasts it with per-component clamping, and builds the forward-variance recursion for Var(a_L) one step at a time. It derives Xavier initialization from variance preservation and He initialization from the fact that ReLU zeros half its inputs, with no algebra skipped, then shows that the backward pass needs the SAME scale. It proves the zero-initialization symmetry problem by hand and ends with a 50-layer signal-propagation experiment in which only He survives. Every snippet runs standalone, and every number came from real execution with numpy 2.2.6. The lesson runs to 62 slides.
Lesson 37: Phase 1 Review & Mock PrepUSAAIO Lesson 37, from Week 13, a fully worked Phase 1 consolidation built around ONE running dataset, the five students from Lesson 7. OLS is derived and verified FOUR equivalent ways - by the normal equations, by Gaussian MLE, by the pseudoinverse and SVD, and by hat-matrix projection - with every step shown. Each of the four pillars is then re-derived on that same data: eigenvalues and PSD matrices, the SVD singular values, softmax with cross-entropy giving dL/dz = p - y, KL divergence and cross-entropy, bias and variance, ROC against PR-AUC under imbalance, gradient descent and Newton's method both converging to the closed form, and ridge regression lifting the zero eigenvalue. Every snippet runs self-contained, and every number was produced by real execution. The lesson runs to 62 slides.
Lesson 38: Mock Exam — Phase 1 TheoryUSAAIO Lesson 38, from Week 13, fully worked: a timed Phase 1 theory mock exam turned into an olympiad-grade walkthrough. Nine exam questions span linear algebra with symmetric eigendecomposition, MLE, Hessian curvature, KL divergence, AUC, kernels and Mercer's theorem, generalization, the softmax-with-cross-entropy gradient, and the p-value trap. Each is derived one move per beat with nothing skipped, every number was produced by real numpy, torch, and sklearn execution, and there is a from-scratch metrics-evaluator project verified against sklearn. The lesson runs to 69 slides.
Lesson 39: Mock Exam — Phase 1 CodingUSAAIO Lesson 39, from Week 13, fully worked: a timed from-scratch coding mock exam over the four Phase 1 algorithms. PCA is done by both the covariance and the SVD route, derived and traced on a four-point toy and then on digits; gradient descent on (w−5)² is traced step by step with the learning-rate stability bound; Gaussian MLE is derived from the log-likelihood, with the 1/n against 1/(n−1) split; and 5-fold cross-validation is built index by index. Each comes with a theory sub-question, every snippet is self-contained and verified against sklearn, and there is a code-review protocol and a your-turn rebuild. Every number was produced by real execution. The lesson runs to 64 slides.
Lesson 40: Linear Regression & the PyTorch Training LoopUSAAIO Lesson 40, from Week 14 of Phase 2, fully worked on ONE tiny dataset, x=[1,2,3,4] and y=[3,4,7,8]. Closed-form OLS is solved as a 2x2 system by hand, entry by entry, the MSE gradient is derived term by term, and gradient descent is stepped by hand until it lands on the closed-form w=[1.0,1.8]. The canonical PyTorch training loop is then dissected step by step - tensors, nn.Linear, MSELoss, autograd, and SGD - with one step hand-checked and a full loss trace to convergence, and the zero_grad and learning-rate traps are shown through real divergence. It closes with weight_decay and ridge, polynomial features, R2 and RMSE, and a match against sklearn. Every snippet runs standalone in a fresh interpreter, and every number came from real torch and numpy execution. The lesson runs to 63 slides.
Lesson 41: Logistic RegressionUSAAIO Lesson 41, from Week 14 of Phase 2. You derive the logistic-regression loss from scratch and prove the gradient dL/dw = X^T(p−y), then implement binary logistic regression two ways: by manual backpropagation in NumPy, and in PyTorch with nn.Linear feeding a sigmoid and BCELoss. It covers BCEWithLogitsLoss and why it is numerically preferable, softmax regression for the multi-class case with CrossEntropyLoss, the intuition behind the decision boundary and how L2 regularization sharpens or softens it, and a comparison with the linear SVM. The lesson runs to 31 slides.
Lesson 42: PyTorch Data Pipeline & Model SerializationUSAAIO Lesson 42, from Week 15. It covers writing a custom Dataset with __len__ and __getitem__, the DataLoader with its batching, shuffling, and num_workers, manual feature transforms, inspecting an nn.Module's parameters with .parameters() and .named_parameters(), and model serialization - saving a state_dict against saving the whole model, and saving and resuming from a checkpoint. You then build a complete pipeline: a 120-sample tabular dataset, a TwoLayerNet going 4→8→1, trained with Adam, saving a checkpoint at epoch 5, then resuming and reaching a loss of 0.1666 by epoch 10. The lesson runs to 31 slides.
Lesson 43: Support Vector MachinesUSAAIO Lesson 43, from Week 15 of Phase 2, fully worked. It derives the maximum-margin hyperplane from the signed-distance formula with no steps skipped, explains the ±1 canonical scaling, and proves that the margin is 2/‖w‖ on a hand-solvable four-point toy. It then sets up the hard-margin QP with its Lagrangian and KKT conditions and verifies w = Σαᵢyᵢxᵢ against sklearn's dual coefficients. From there it covers soft-margin slack and the exact role of C, the kernel trick with an RBF value checked by hand, grid search with cross-validation, one-versus-one against one-versus-rest for multi-class, and SVM against logistic regression, ending with a from-scratch Iris pipeline. Every snippet runs standalone, and every number came from real execution. The lesson runs to 64 slides.
Lesson 44: MLP Architectures — Depth, Width & ActivationsUSAAIO Lesson 44, from Week 15 of Phase 2, fully worked. One running 2→3→1 MLP is hand-traced neuron by neuron and confirmed against torch. It states the universal approximation theorem and demonstrates it on XOR, showing a single Linear layer fail, then covers depth against width as expressivity by counting linear regions. It derives the exact nn.Linear parameter formula and sums it layer by layer, computes all four activations - sigmoid, tanh, ReLU, and GELU - at five inputs with their gradient consequences, and rebuilds GELU from erf. It quantifies the vanishing-gradient and dying-ReLU failure modes, runs a fair wide-against-deep experiment at about 38k parameters on load_digits, and ends with a from-scratch build-it project. Every snippet runs standalone in a fresh interpreter, and every number came from real execution with torch 2.7.1 and numpy 2.2.6. The lesson runs to 63 slides.
Lesson 45: Learning Curves, Bias-Variance Diagnosis, and Double DescentUSAAIO Lesson 45, from Week 16, fully worked on ONE running dataset, y = x^2 - x + noise. It derives the bias-variance decomposition of the expected loss term by term, then builds loss-against-epoch and loss-against-training-size curves from scratch and reads them zone by zone, diagnosing high bias against high variance from the gap between them. It derives and measures the targeted cures - dropout, early stopping with patience, and changes in capacity - and produces double descent with a real minimum-norm interpolator that spikes exactly at P = n and then descends again. Every snippet is self-contained and runs in a fresh interpreter, and every number in every trace table was produced by real execution with fixed seeds. The lesson runs to 63 slides.
Lesson 46: Batch NormalizationUSAAIO Lesson 46, from Phase 2. It covers the batch-normalization forward pass, with mu, sigma^2, x_hat, and the gamma and beta parameters, then the full backward pass giving dL/dgamma, dL/dbeta, and dL/dx, and the running mean and variance used at inference. It then covers LayerNorm, which normalizes along the feature dimension for transformers, and GroupNorm for small batches. You implement BatchNorm1d from scratch and verify it against nn.BatchNorm1d. The lesson runs to 31 slides.
Lesson 47: Decision Trees — CART, Information Gain, Gini, and PruningUSAAIO Lesson 47, from Phase 2. It covers information gain from entropy, Gini impurity, the CART binary-split induction algorithm, the stopping criteria, and cost-complexity pruning through ccp_alpha. You build a DecisionTreeClassifier from raw splits up to a pruned final model, on iris and on make_classification. All the trace values were verified with scikit-learn and numpy in June 2026. The lesson runs to 32 slides.
Lesson 48: Dropout & Modern RegularizersUSAAIO Lesson 48, from Week 17 of Phase 2. It covers the theory and implementation of inverted dropout, the difference between training and evaluation mode, dropout read as averaging over an exponential ensemble, MC Dropout for Bayesian uncertainty estimation, and label smoothing from scratch, then compares several dropout rates on a classification network. All the numbers were verified with torch 2.7.1 and sklearn in June 2026. The lesson runs to 29 slides.
Lesson 49: Random ForestsUSAAIO Lesson 49, from Phase 2. It covers bootstrap aggregation, random feature subsets of size sqrt(d), out-of-bag error as free validation, and three kinds of feature importance - MDI, permutation, and SHAP - along with the trade-offs between a random forest, a single tree, and boosting. You implement RandomForestClassifier from scratch and compare its out-of-bag error against k-fold cross-validation. The lesson runs to 30 slides.
Lesson 50: Gradient BoostingUSAAIO Lesson 50, from Phase 2. It covers the gradient-boosting framework F_m = F_{m-1} + eta*h_m and pseudo-residuals as the negative gradients of the loss - y − F for MSE, and y − sigma(F) for classification - then builds a three-round GBM regressor from scratch using an sklearn DecisionTree. It goes on to the XGBoost improvements, namely L2 leaf regularization, column subsampling, and approximate splits, and to LightGBM's histogram-based leaf-wise growth, ending with a cross-validated hyperparameter sweep. You build GradientBoostingRegressorFromScratch on make_regression data and compare the GBM with a random forest. The lesson runs to 29 slides.
Lesson 51: Custom Loss Functions & AutogradUSAAIO Lesson 51, from Week 18 of Phase 3, on implementing custom differentiable losses in PyTorch. It covers focal loss for class imbalance, triplet loss with online hard-negative mining for metric learning, contrastive loss as the foundation of CLIP, and torch.autograd.Function for gradients that are not standard. Every formula was verified by real execution with torch 2.7.1 and sklearn in June 2026. The lesson runs to 30 slides.
Lesson 52: k-Means ClusteringUSAAIO Lesson 52, from Phase 3 on unsupervised learning. It covers k-Means as assign-then-update, k-Means++ initialization, the elbow method and WCSS, the silhouette score, the EM interpretation in which the E step assigns and the M step updates the centroids, and soft k-Means as a Gaussian mixture model. It then shows k-Means failing on non-spherical clusters and DBSCAN fixing it. You build k-Means from scratch with both random and k++ initialization, implement the elbow method and the silhouette score from scratch, and verify convergence on make_blobs. The lesson runs to 29 slides.
Lesson 53: Gaussian Mixture Models & the EM AlgorithmUSAAIO Lesson 53, from Week 19 of Phase 3. It covers the GMM density model as K weighted Gaussians, then derives the E-step soft assignments and the M-step parameter updates from the expected complete-data log-likelihood, along with the convergence guarantees of EM. It covers the covariance types and model selection by BIC, compares a GMM with k-Means on non-spherical data, and ends with GMM-based anomaly detection. You build a GMM from scratch with NumPy and SciPy and validate it against sklearn. The lesson runs to 32 slides.
Lesson 54: DBSCAN & Hierarchical ClusteringUSAAIO Lesson 54, wrapping up Week 19. It covers DBSCAN, with its core, border, and noise points and its epsilon and min_samples parameters, then agglomerative hierarchical clustering with dendrograms and the single, complete, average, and Ward linkage methods. It covers validating clusters without labels, by silhouette score, and with labels, by ARI and NMI, and ends with a practical guide to matching the geometry of a dataset to the right algorithm. It was verified with sklearn and scipy in June 2026. The lesson runs to 28 slides.
Lesson 55: 2D Convolutions & CNNsUSAAIO Lesson 55, from Phase 2. It builds 2D convolution from scratch, covering the output-shape formula, parameter sharing, and translation equivariance, then max pooling and the growth of the receptive field, and finishes with a complete CNN forward pass in PyTorch using nn.Conv2d, nn.MaxPool2d, and nn.ReLU. It was verified on the sklearn digits dataset, reaching 97% test accuracy after 100 epochs. The lesson runs to 30 slides.
Lesson 56: ResNet, Depthwise Separable Convolutions & EfficientNetUSAAIO Lesson 56, from Phase 3. It covers ResNet's identity shortcuts and how they fix the vanishing gradient, the BasicBlock against the Bottleneck design, and pre-activation against post-activation BatchNorm. It then covers depthwise separable convolutions, which give MobileNet roughly an eight-fold reduction in parameters, and EfficientNet's compound scaling of depth, width, and resolution together. It was verified with torch 2.7.1. The lesson runs to 30 slides.
Lesson 57: Transfer Learning & Fine-TuningUSAAIO Lesson 57, from Week 20. It compares feature extraction with fine-tuning, covers freezing the pretrained layers, and gives the four-quadrant decision rule that crosses data size with domain similarity. It then covers progressive layer unfreezing, differential learning rates - typically ten times lower for the pretrained layers - and the data-augmentation strategies mixup and cutmix. You build a frozen-backbone classifier and then progressively unfreeze it. The lesson runs to 30 slides.
Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP)USAAIO Lesson 58, from Phase 3. It builds Kernel PCA from scratch - the kernel matrix, centering, eigendecomposition, and projection - and covers the choice between an RBF and a polynomial kernel. It then covers the mechanics of t-SNE, including perplexity, the learning rate, the KL divergence, and why projecting a new point fails, and the theory of UMAP, with its fuzzy simplicial sets, cross-entropy optimization, and inductive mapping. It ends by comparing the methods on the Swiss roll. All the numbers were verified with sklearn 1.x and numpy 2.2.6. The lesson runs to 31 slides.
Lesson 59: Hyperparameter OptimizationUSAAIO Lesson 59, from Phase 3. It covers grid search, which is exhaustive and exponentially costly; random search, which samples on a log scale and matches or beats grid search at a fraction of the evaluations; Gaussian-process Bayesian optimization, with a surrogate model and an expected-improvement acquisition function; and successive halving with Hyperband, which eliminates candidates early. It closes with practical tips for working on a log scale. All the numbers were verified with sklearn 1.x on the digits dataset in June 2026. The lesson runs to 29 slides.
Lesson 60: sklearn Pipeline, ColumnTransformer & GridSearchCVUSAAIO Lesson 60, from Week 21. An sklearn Pipeline chains estimators together and prevents leakage; ColumnTransformer applies different transforms to different feature types; and FunctionTransformer wraps any function. Combining a Pipeline with GridSearchCV lets you tune the preprocessing and the model hyperparameters jointly, and joblib saves and loads versioned pipelines. You build a production-ready pipeline that imputes, encodes, scales, and fits a logistic regression, tune it end to end, and serialize it. The lesson runs to 31 slides.
Lesson 61: Multiclass Classification — Softmax, OvR, OvOUSAAIO Lesson 61. It covers softmax regression, giving P(y=k|x) and the cross-entropy gradient P − Y_one_hot, then the One-vs-Rest and One-vs-One multiclass strategies, multinomial logistic regression against a set of separate binary logistic models, and weighted cross-entropy for class imbalance. You implement multinomial logistic regression from scratch and benchmark One-vs-Rest, One-vs-One, and multinomial on an imbalanced 10-class dataset. The lesson runs to 29 slides.
Lesson 62: Learning Rate SchedulesUSAAIO Lesson 62. It covers linear warmup, cosine annealing, SGDR with cosine restarts, the one-cycle policy for both learning rate and momentum, and the inverse-square-root Transformer schedule. Each is derived from first principles and verified in PyTorch with LambdaLR and the built-in schedulers, and there is a project implementing LambdaLR from scratch. The lesson runs to 30 slides.
Lesson 63: Transfer Learning, Fine-Tuning & Few-Shot MethodsUSAAIO Lesson 63, from Week 22 of Phase 3. It covers the fine-tuning strategies, running from a linear probe through gradual unfreezing to a full fine-tune, then domain adaptation and distribution shift, few-shot learning with 5-shot tasks and prototypical networks, and zero-shot learning through textual descriptions, which is the foundation of CLIP. All the numbers were verified on the digits dataset with torch 2.7.1 and sklearn in June 2026. The lesson runs to 31 slides.
Lesson 64: Bag of Words, TF-IDF, and Text VectorizationUSAAIO Lesson 64. It covers bag-of-words count vectors and TF-IDF weighting, giving both the formula and a from-scratch implementation, then n-grams and the text-preprocessing pipeline - lowercasing, stopword removal, and stemming against lemmatization. It covers CountVectorizer and TfidfVectorizer, and ends with a full text-classification pipeline combining TF-IDF with logistic regression. The lesson runs to 30 slides.
Lesson 65: Naive Bayes ClassifiersUSAAIO Lesson 65, from Phase 3. It derives Multinomial, Bernoulli, and Complement Naive Bayes from first principles, covering the derivation of Laplace smoothing, log-space arithmetic to prevent underflow, Multinomial against Bernoulli Naive Bayes on count against binary features, and ComplementNB for imbalanced short texts. It also proves that Naive Bayes is a linear classifier in log space. You implement MultinomialNB from scratch and benchmark it against logistic regression. The lesson runs to 32 slides.
Lesson 66: Mixed Precision, Loss Scaling & Gradient AccumulationUSAAIO Lesson 66, from Week 23 of Phase 3. It compares the numerical ranges of FP16 and FP32, then covers mixed-precision training with torch.autocast and GradScaler, loss scaling to prevent FP16 underflow, gradient accumulation as a way to simulate a large batch, and the arithmetic of the memory budget. You build an AMP training loop and verify it against full precision. The lesson runs to 29 slides.
Lesson 67: Stacking & Blending EnsemblesUSAAIO Lesson 67, from Phase 3. It covers stacking, using a 5-fold out-of-fold meta-learner, and blending, using a 20% hold-out, then explains why diversity beats homogeneity and how to combine heterogeneous models such as logistic regression, a random forest, and kNN. It closes with the analytic variance-reduction proof for averaging. All the numbers were verified with sklearn 1.x and numpy 2.2.6. The lesson runs to 28 slides.
Lesson 68: AdaBoostUSAAIO Lesson 68, from Phase 3. It builds AdaBoost from scratch, covering the sample-weight updates, the alpha formula, minimization of the exponential loss, gradient descent in function space, and the convergence theory. You implement AdaBoost over decision stumps, verify that it matches sklearn, and compare it with a single deep tree. The lesson runs to 28 slides.
Lesson 69: kNN and the Curse of DimensionalityUSAAIO Lesson 69, from Week 24 of Phase 4. It builds kNN majority-vote classification from scratch, covers the L2, L1, and cosine distance metrics, and works the curse of dimensionality, where the ratio of the maximum to the minimum distance converges to 1. It then covers the KD-tree and ball-tree for efficient lookup, weighted kNN using 1/distance, and the bias-variance trade-off as k changes. Every number was verified with numpy and sklearn in June 2026. The lesson runs to 32 slides.
Lesson 70: PyTorch Hooks, Profiling & torch.compileUSAAIO Lesson 70. It covers forward and backward hooks for capturing activations and watching gradient flow, memory management with del and gradient checkpointing, torch.profiler for locating training bottlenecks, and the torch.compile modes for faster inference. Every number was verified with torch 2.7.1 and numpy 2.2.6 in June 2026. The lesson runs to 30 slides.
Lesson 71: Kaggle Competition PipelineUSAAIO Lesson 71, from Phase 4, on the end-to-end competition workflow. It covers exploratory data analysis - distributions, correlations, missing values, and outliers - then interaction and polynomial feature engineering, model selection across the linear, tree, and neural families with 5-fold cross-validation, stacking ensembles, and using cross-validation scores to predict where you will land on the leaderboard. It was verified on the sklearn diabetes dataset with numpy 2.2.6 and scikit-learn. The lesson runs to 30 slides.
Lesson 72: Phase 2 Algorithm ReviewUSAAIO Lesson 72, from Week 25, consolidating Phase 2. It is a timed review of every classification and regression algorithm, the ensemble methods of bagging, boosting, and stacking, and the clustering methods k-Means, GMM, DBSCAN, and hierarchical clustering, ending with a 30-minute PyTorch mastery check. Everything was verified on sklearn datasets with real execution numbers. The lesson runs to 30 slides.
Lesson 73: Metric Learning, Siamese Networks & Triplet LossUSAAIO Lesson 73, on learning embeddings in which distance reflects semantic similarity. It covers metric learning, Siamese networks, margin-based triplet loss, contrastive loss, online hard-negative mining, and the connection to attention. It implements a Siamese EmbeddingNet in PyTorch, where hard mining improves the ratio of inter-class to intra-class distance from 5.60× to 8.17× on a three-identity benchmark. The lesson runs to 30 slides.
Lesson 74: Neural Network Debugging ChecklistUSAAIO Lesson 74, from Phase 4. It gives a systematic five-step debugging protocol for a broken training loop: overfit a single batch, check gradient flow with per-layer norms, verify the initial loss against log(K), check the initial accuracy against 1/K, and start from hyperparameters taken from the literature. It includes an automated checklist function and a bug-injection exercise. The lesson runs to 30 slides.
Lesson 75: Mock Exam — Phase 2 Full ReviewUSAAIO Lesson 75, from Week 26: a timed Phase 2 mock exam. The two-hour theory paper covers ensemble methods, normalization, optimization, and evaluation, and the two-hour coding paper asks you to implement a random forest, a GMM, and a CNN training loop from scratch. It comes with a full answer-key walkthrough and a gap analysis to run before Phase 3 on transformers. All the numbers were verified against torch 2.7.1, sklearn, and numpy 2.2.6. The lesson runs to 32 slides.
Lesson 76: Algorithm Review & SelectionUSAAIO Lesson 76, a Phase 2 review. It is a flash review of seven core classifiers - logistic regression, the SVM, decision trees, random forests, GBMs, kNN, and Naive Bayes - covering their loss functions, gradients, hyperparameters, and bias-variance profiles, with whiteboard-style derivations. It also covers the USAAIO exam question patterns and gives an algorithm-selection framework. You implement and benchmark all seven on a synthetic dataset. The lesson runs to 28 slides.
Lesson 77: Deep Learning FoundationsUSAAIO Lesson 77, from Phase 4 on deep learning. It covers the MLP, CNN, and ResNet architectures; the ReLU, sigmoid, and tanh activations and the mechanics of the vanishing gradient; BatchNorm, dropout, and Xavier and Kaiming initialization; the PyTorch training loop line by line; backpropagation through the chain rule, with a concrete two-layer trace; the SGD, Momentum, Adam, and AdamW optimizers and learning-rate scheduling; and Dataset and DataLoader for mini-batch pipelines. You build a complete digits classifier with every one of these techniques applied. The lesson runs to 32 slides.
Lesson 78: Mock Exam — Phase 2 (Theory + Coding)USAAIO Lesson 78, from Week 26: a three-hour combined Phase 2 mock exam. It has 30 theory questions spanning every Phase 2 topic, plus two timed coding problems - logistic regression from scratch, and a CNN in PyTorch - each with a theory sub-question and a verified answer key. It is the Phase 2 coding and theory milestone. The lesson runs to 27 slides.
Lesson 79: Scaled Dot-Product AttentionUSAAIO Lesson 79, from Phase 3 on transformers and NLP. It starts from the sequence bottleneck in an RNN, then covers Bahdanau alignment scores and self-attention, where Q, K, and V all come from the same sequence. It derives Attention(Q,K,V) = softmax(QK^T/sqrt(d_k))V, proves by a variance argument that dot products grow with d_k, and explains why the scaling factor prevents the softmax from saturating. It was verified with torch 2.7.1+cpu in June 2026. The lesson runs to 28 slides.
Lesson 80: Scaled Dot-Product Attention from ScratchUSAAIO Lesson 80, from Phase 3, in which you implement scaled_dot_product_attention(Q, K, V, mask) from scratch in PyTorch. It analyzes the shapes of Q, K, and V, gives the rationale for the 1/sqrt(d_k) scaling, and covers causal lower-triangular masking and padding masks. It then covers numerical stability, including FP16 overflow and the max-subtraction fix, the O(T^2 * d_k) complexity, and verification against F.scaled_dot_product_attention. All the trace-table values were verified with torch 2.7.1+cpu in June 2026. The lesson runs to 32 slides.
Lesson 81: Multi-Head AttentionUSAAIO Lesson 81, from Week 28 of Phase 3, building multi-head attention from scratch. It covers projecting Q, K, and V into h parallel subspaces, running scaled dot-product attention in each head, and concatenating before the W_O projection. It proves that the parameter count is 4·d_model² regardless of h, and shows that h=1 reduces to single-head attention. You implement MultiHeadAttention as an nn.Module and verify every shape. The lesson runs to 28 slides.
Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder BlockUSAAIO Lesson 82, from Phase 3 on transformers. It builds LayerNorm from scratch, normalizing over the feature dimension and verifying against nn.LayerNorm, then compares pre-norm with post-norm for stability. It covers the feed-forward sublayer - two Linear layers with a GELU between them, at d_ff = 4·d_model - and residual connections for gradient flow, then assembles the complete pre-norm EncoderBlock as an nn.Module. All the numbers were verified with torch 2.7.1+cpu in June 2026. The lesson runs to 30 slides.
Lesson 83: Positional EncodingUSAAIO Lesson 83, from Phase 3. It explains why self-attention is permutation-equivariant and therefore needs positional information, then covers sinusoidal positional encoding, with its formula and the proof that relative position is linear in it; learned positional encoding via nn.Embedding; RoPE, which rotates Q and K by an angle set by position and so gives a dot product that depends on relative position; and ALiBi, which adds a linear bias to the attention scores and extrapolates to longer sequences. All the values were computed with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 32 slides.
Lesson 84: Transformer Encoder for ClassificationUSAAIO Lesson 84, from Week 29 of Phase 3. It covers stacking N encoder blocks, sinusoidal positional encoding, pooling on the [CLS] token, and constructing a padding mask, then assembles the full encoder classifier with AdamW, linear warmup, cosine decay, and label smoothing. You build TransformerEncoderClassifier from scratch, verify that it is permutation-invariant without positional encoding, and confirm all the encoder-block parameter counts. The lesson runs to 32 slides.
Lesson 85: Transformer Decoder BlockUSAAIO Lesson 85, from Phase 3. It covers the three sublayers of the transformer decoder block - causal masked self-attention, using an upper-triangular -inf mask; cross-attention, where Q comes from the decoder while K and V come from the encoder; and the feed-forward network - along with autoregressive generation. Every attention weight, shape, and row sum was verified with torch 2.7.1+cpu at seed 42. The lesson runs to 27 slides.
Lesson 86: Transformer Encoder-Decoder & Seq2SeqUSAAIO Lesson 86, from Phase 3. It covers the encoder-decoder architecture, teacher forcing against autoregressive inference, beam search with a k=2 trace, and cross-entropy loss with label smoothing at eps=0.1, then trains a toy seq2seq Transformer from scratch in PyTorch. All the attention scores, causal masks, label-smoothing probabilities, cumulative beam-search log-probabilities, and training losses were verified with torch 2.7.1. The lesson runs to 30 slides.
Lesson 87: Transformer InterpretabilityUSAAIO Lesson 87, from Phase 3. It covers visualizing attention heads with heatmaps, attention rollout across layers, linear probing classifiers on transformer representations, gradient-times-input attribution, integrated gradients, and layer-wise analysis, in which the early layers are syntactic and the late ones semantic. All the attention weights, the 92% probe accuracy, and the integrated-gradient scores were verified with torch 2.7.1+cpu and sklearn in June 2026. The lesson runs to 30 slides.
Lesson 88: BERT — Pretraining & Fine-TuningUSAAIO Lesson 88, from Phase 3 on transformers and NLP. It covers BERT's encoder-only architecture, Masked Language Model pretraining, Next Sentence Prediction, and bidirectional against causal attention, then fine-tuning on the [CLS] token for classification and a from-scratch TinyBertForClassification in PyTorch. The toy attention scores, the contrast with a causal mask, and the fine-tuning loss were all verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 32 slides.
Lesson 89: GPT & Decoder-Only TransformersUSAAIO Lesson 89, from Phase 3, on GPT as a decoder-only causal transformer. It covers the absence of cross-attention, the causal self-attention mask, learned positional embeddings, and the feed-forward network at 4× expansion, then the GPT-2 scaling family, running from 12 to 48 layers and 768 to 1600 dimensions, the Chinchilla scaling laws, the emergent abilities of few-shot prompting and chain-of-thought, and autoregressive inference. You build TinyGPT from scratch in PyTorch and trace a causal-language-modeling training step. It was verified with torch 2.7.1 in June 2026. The lesson runs to 30 slides.
Lesson 90: KV Cache & Attention EfficiencyUSAAIO Lesson 90, from Week 31 of Phase 3. It covers the KV cache for autoregressive inference, which costs O(n·d) memory and O(1) per step; Flash Attention's block-wise computation, which needs only O(n) SRAM; sliding-window attention, at O(n·w) rather than O(n²); the sparse attention patterns of BigBird and Longformer; and Group Query Attention along with Multi-Query Attention. All the complexity claims and memory numbers were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 30 slides.
Lesson 91: Flash Attention & Memory-Efficient TransformersUSAAIO Lesson 91, from Phase 3. It covers Flash Attention's IO complexity, which is O(N) in HBM traffic rather than O(N²), tiled softmax with an online running maximum, gradient checkpointing, the ALiBi position bias, and Multi-Query Attention. All the numbers were verified with torch 2.7.1+cpu: the tiled output matches standard attention to a maximum absolute error of 6e-8, and the MQA KV cache is eight times smaller than MHA at N=2048. The lesson runs to 32 slides.
Lesson 92: Vision Transformer (ViT)USAAIO Lesson 92, from Phase 3. It covers patch embedding through an nn.Conv2d with stride P, the CLS token, learned positional embeddings, applying a standard transformer encoder to the sequence of patches, and the classification head on the CLS token, then compares fine-tuning with training from scratch and the scaling behavior of ViT against a CNN. TinyViT8 is built from scratch and trained on load_digits to 96.4% test accuracy after 150 epochs, with its 26,538 parameters verified with torch 2.7.1+cpu. The lesson runs to 30 slides.
Lesson 93: BERT Variants — XLNet, RoBERTa, T5, DeBERTa, ALBERTUSAAIO Lesson 93, from Phase 3. It explains why bidirectional BERT is slow at inference and how each variant fixes a different limitation: XLNet's permutation language modeling, with two-stream attention and separate content and query masks; RoBERTa's dynamic masking, dropped NSP, and larger batches; T5's unified text-to-text seq2seq framing; DeBERTa's disentangled content and position attention; and ALBERT's cross-layer parameter sharing, whose six-fold reduction is verified in PyTorch at d=64. All the parameter counts, mask tables, and shape traces were verified with torch 2.7.1+cpu. The lesson runs to 31 slides.
Lesson 94: Graph Neural Networks (GCN & GAT)USAAIO Lesson 94, from Phase 3. It covers message passing on graphs and derives the normalized adjacency D^{-1/2} Ã D^{-1/2} from first principles on a 4-cycle, then works a GCN forward pass with real computed values. It covers Graph Attention Networks, with the attention weights verified to sum to 1, and global readout pooling by mean, max, or sum, ending with a TinyGCN of 48 parameters trained to 100% accuracy on a synthetic 8-node community graph. All the numbers were verified with torch 2.7.1+cpu. The lesson runs to 28 slides.
Lesson 95: GraphSAGE, Oversmoothing, and Scalable GNNsUSAAIO Lesson 95, from Phase 3. It covers GraphSAGE's fixed-size neighborhood sampling and mean aggregation, then oversmoothing in deep GCNs, where similarity rises from 0.806 at 2 layers to 0.999 at 10 on a six-node toy graph seeded with torch.manual_seed(0). It then covers skip connections through GCNII, where alpha=0.1 brings the 10-layer similarity back down to 0.902, the PNA multi-aggregator, which concatenates mean, standard deviation, max, and min into dimension 16, and Graph Transformer attention over graph neighborhoods. All the trace tables were verified with torch 2.7.1+cpu and numpy 2.2.6 on a fixed-seed six-node synthetic graph. The lesson runs to 27 slides.
Lesson 96: Graph Transformer & GraphormerUSAAIO Lesson 96, from Phase 3, on graph transformers. It replaces positional encoding with structural encoding, using degree centrality and shortest-path distance, then covers Graphormer's spatial bias and edge encoding, the VNode used for a global graph readout, full all-pairs attention against the sparsity of a GNN, and the applications to molecular property prediction and knowledge graphs. A four-node toy molecule was verified with torch 2.7.1+cpu, covering the shortest-path-distance matrix, the spatial bias matrix, the biased attention weights, masked attention in the sparse variant, and the edge-encoding projection. The lesson runs to 27 slides.
Lesson 97: Subword Tokenization — BPE, WordPiece, SentencePieceUSAAIO Lesson 97, from Phase 3. It sets out the out-of-vocabulary problem that character-level and word-level tokenization each run into, then covers the BPE algorithm, which iteratively merges the most frequent adjacent pair; WordPiece's likelihood-based scoring; SentencePiece's language-agnostic treatment of raw bytes; and the trade-off that vocabulary size forces between sequence length and parameter count. BPE is implemented from scratch on a toy corpus and verified with Python 3, so all the merge steps, token sequences, and sequence-length tables are real execution output. The lesson runs to 26 slides.
Lesson 98: Word2Vec, GloVe & Static Word EmbeddingsUSAAIO Lesson 98, from Phase 3. It covers the Skip-gram objective and the binary classifier that negative sampling turns it into, CBOW's mean-context prediction, and GloVe's factorization of the co-occurrence matrix under a weighted MSE. It then explains the linguistic regularity king − man + woman ≈ queen through PMI geometry, and gives the equivalence between Skip-gram with negative sampling and SPPMI factorization, following Levy and Goldberg (2014). All the objectives, losses, and trace-table values were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 32 slides.
Lesson 99: BERT Contextual Embeddings & SBERTUSAAIO Lesson 99, from Phase 3. It contrasts static with contextual word representations: Word2Vec gives "bank" a cosine similarity of 1.0 with itself regardless of context, while a contextual BERT-like encoder gives 0.5639. It covers BERT's layer specialization, where layers 1 to 4 are surface, 5 to 8 syntactic, and 9 to 12 semantic, and the sentence-pooling strategies of [CLS], mean, and max, of which mean wins on semantic-similarity tasks. It then trains an SBERT siamese network with a cosine and NLI objective, taking the loss from 1.6056 to 0.0008 in 50 steps and reaching a cosine similarity of 0.9896 for the same context against −0.9918 for a different one. It closes with semantic search by cosine retrieval, where the query scores 0.9945, 0.2380, 0.1835, −0.2640, and −0.2951. All the values were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 24 slides.
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.