Mu Li's material fits two stages. While learning training, use Dive into Deep Learning to connect concepts and code. When reading research, watch how he examines abstracts, figures, methods, and experiments.
Starting out: study and modify a training run
Use the Chinese Dive into Deep Learning textbook with the official course and video index. Start with data operations, linear models, losses, and optimization, then multilayer perceptrons. Keep code beside you: what shape is the input, where does loss come from, and how do parameters update?
If you need a project, try the classification route. For mathematics, return to foundations. Other explanations include fast.ai, 3Blue1Brown, Hung-yi Lee, and CS229.
First papers: read independently, then compare
Find “how to read a paper” in the paper-reading index, then a paper you are studying. For attention, try the Transformer close reading.
Write your understanding of the question, method, and experiments first. After listening, compare where he paused, which figure you overlooked, and which claims need experiments. That is how you develop a reading approach of your own.
Continue
Choose a paper offers specific starting points; finding and reading papers covers the process.