The mean, variance, and confidence interval math underneath every eval score — and why sample size changes how much a "Model A beats Model B" claim is worth.
How the gradient that gradient descent needs actually gets computed, layer by layer, using the chain rule.
How training a model actually works — measure how wrong it is, nudge the weights downhill, repeat millions of times.