Guide home/Start learning
中文
Start learning

Which foundations do you need, and how much?

ML, deep learning, mathematics, and computing are worth building gradually. Filling a gap during a project and studying a course systematically can happen together.

Python and computing

First learn to read a small program: where data enters, what a function receives and returns, what a loop repeats, and where an error points. Use CS50P for functions, conditions, loops, exceptions, and files. If you can already code, start with the project in front of you.

Environments, paths, terminals, and Git often block progress before algorithms do. Missing Semester is useful here; begin with the shell and version control. For the roles of CPUs, operating systems, and networks, see Crash Course Computer Science.

After changing a program, save a version you can return to and explain how to run it. You will understand why another project's README needs those details.

ML: why models work on unseen data

Training, validation, and test sets; loss, overfitting, generalization, and baselines: these ideas determine how you interpret an experiment. Connect them through a classification or regression task.

An Introduction to Statistical Learning builds that understanding. Start with statistical learning, classification, resampling, and model selection. For Chinese videos, choose the relevant topics from Hung-yi Lee's course.

Return to your project and explain why repeated tuning on test results is a problem and why a high training score need not be good. For more systematic derivations, continue to CS229.

Deep learning: understand a training run

Dive into Deep Learning connects data operations, linear models, losses, gradients, optimization, and multilayer perceptrons. The companion course includes Chinese videos.

Compare it with the training loop in PyTorch's basics tutorial. Once you can follow a batch through the forward pass, loss, backpropagation, and update, continue the project. Return when a new architecture raises questions.

For an application-first approach, try fast.ai. To take automatic differentiation apart, try micrograd in Karpathy's Zero to Hero. Choose the explanation that addresses your current question.

Where mathematics appears

Vectors, matrices, and linear transformations appear in representations, network layers, and attention. Build visual intuition with 3Blue1Brown's linear algebra, then check dimensions in code. For systematic depth, use MIT 18.06.

Derivatives, gradients, and the chain rule connect loss to parameter updates. Use 3Blue1Brown's calculus for intuition, differentiate a simple function yourself, and compare with automatic differentiation.

Probability, expectation, and conditional probability recur in generative models, RL, and experimental analysis. Explore distributions interactively with Seeing Theory. Use Stat 110 when you want a full course.

Choose a concept you currently need, solve some exercises, and return to its role in the paper. Mathematics for Machine Learning helps connect these areas.

Going deeper into architectures

For Transformers, follow The Illustrated Transformer, then inspect inputs and outputs with Transformer Explainer. Use the Annotated Transformer to connect attention diagrams to implementation.

For derivations of positional encodings, diffusion, and related topics, browse Jianlin Su's Scientific Spaces. Start with a question, read until you can explain it, and return to your work.

Return to what you want to do

Small projects · LLMs · Vision · RL · Directions

Try another explanation

The course and textbook catalog preserves other choices in mathematics, computing, ML/DL, and programming. Mu Li's courses and paper readings connect foundations, code, and research reading.