FIG 0 · The full curriculum
Twenty-seven chapters. Four acts.
Read straight through, or use the dependency map on the home page to jump in wherever you already have the prerequisites.
- 27
- Chapters
- 4
- Signature chapters
- 199k
- Words of writing
FIG 1 · CH. 00–08 · 9 CHAPTERS · Math, ML basics, classical algorithms
Act I · Foundations
Math & Python Prereqs
Linear algebra, calculus, probability, NumPy. The vocabulary every later chapter assumes.
The ML Landscape
Supervised vs unsupervised vs RL, the metric-choice axis, the workflow vocabulary you'll execute on for the next 26 chapters.
End-to-End ML Project
California housing, no-leakage pipelines, a working model in chapter two. The Géron classic.
Classification
Confusion matrix vocabulary, precision-recall tradeoffs, multi-label, the metric-choice that lies.
Training Models
Linear and logistic regression from first principles. Normal equation, GD, Adam, regularization.
Trees, SVMs, Kernels
CART splits, the max-margin objective, the kernel trick — and an honest 2026 verdict on when SVMs still win.
Ensemble Methods
Why ensembles work, voting → bagging → boosting → stacking, XGBoost in production.
Dimensionality Reduction
Curse of dimensionality, PCA from SVD, t-SNE done right, the UMAP→cluster pipeline.
Unsupervised Learning
k-means with the right init, DBSCAN for non-spherical, GMMs with EM derived in full.
FIG 2 · CH. 09–14 · 6 CHAPTERS · From MLP to NLP-with-attention
Act II · Neural Networks
Intro to Neural Networks
Backprop derived by hand, a 50-line MLP in pure NumPy, the activation-function evolution.
PyTorch Foundations
Tensors, autograd, nn.Module composition, the 4-line training loop, mixed precision, torch.compile.
Training Deep Networks
Init schemes, BatchNorm vs LayerNorm vs RMSNorm, AdamW vs Lion vs Muon, LR schedules, Karpathy's recipe.
CNNs & Computer Vision
Convolution as feature detection, LeNet → ResNet → ViT, the texture-bias problem, adversarial examples.
Sequences & Time Series
RNN, LSTM, GRU, why vanishing-gradient is structural, when to use RNNs in 2026 (rarely).
NLP with RNNs + Attention
Word embeddings, encoder-decoder, Bahdanau → Luong → self-attention. The historical pivot point.
FIG 3 · CH. 15–21 · 7 CHAPTERS · Transformers, multimodal, efficient inference, generative, RL
Act III · Modern AI
Transformers from Scratch
BPE → attention → multi-head → causal mask → RoPE → the full nanoGPT, every line earned by hand.
Multimodal Transformers
ViT, CLIP contrastive loss, LLaVA-style projectors, Stable Diffusion conditioning — modalities sharing a residual stream.
Efficient Inference
KV cache, paged attention, quantization, FlashAttention, speculative decoding, vLLM. The signature systems chapter.
Generative Models
Autoencoders → VAEs → GANs → Diffusion → flow matching. Latent diffusion explained from the math up.
RL + RLHF
MDPs, PPO, RLHF pipeline, DPO derived, GRPO + RLVR. Reward hacking taxonomy from the inside.
Agents & Tool Use
ReAct, function calling, MCP, multi-agent. The capability boundary as design discipline.
RAG & Vector Stores
Embedding models, ANN algorithms, chunking strategies, hybrid search, CAG vs RAG, eval discipline.
FIG 4 · CH. 22–26 · 5 CHAPTERS · Mech-interp, eval, safety, MLOps, reading papers
Act IV · Frontier
Mechanistic Interpretability
The residual stream view, induction heads, sparse autoencoders, attribution graphs. The obvix flagship.
Eval Science
Why evals are hard. LLM-as-judge failure modes, Elo arenas, custom-eval recipe, contamination detection.
AI Safety & Red-Team
OWASP LLM Top 10, the lethal trifecta, mech-interp for safety, a Level-1 prompt-injection CTF. Signature.
MLOps & Observability
The ML system lifecycle, drift detection, on-call discipline, eval-driven development. The 70% of the job.
Reading Papers
The three-pass method, citation-graph navigation, re-implementation as the deepest read.