
Intermediate · Python + PyTorch · 20 lessons · ~24 hours
How to Build an LLM from Scratch
Follow the path of the book 'Build a Large Language Model (From Scratch)': implement a byte-pair tokeniser, a data loader with sliding windows, self-attention from first principles, causal masking, multi-head attention, the full GPT block, a pretraining loop, decoding strategies, loading real GPT-2 weights, and two finetunes — a spam classifier and an instruction-following model. Every module ends with code you run.
What you walk away with
- Implement BPE tokenisation and a sliding-window data loader
- Derive and code scaled dot-product self-attention from zero
- Build a complete GPT-2 style model, block by block
- Run a pretraining loop and interpret the loss curve honestly
- Load real GPT-2 weights and finetune for classification and instructions
Syllabus
1. Module 1 · Working with Text Data
Turn raw text into tensors: tokenisation, vocabulary, special tokens, byte-pair encoding, sliding windows and embeddings.
- · From characters to tokens
- · Byte-pair encoding: why GPT never says <unk>
- · Sliding windows, input–target pairs and the DataLoader
- · Token embeddings + positional embeddings
2. Module 2 · Attention, Derived Not Memorised
Simplified attention with dot products → trainable Q/K/V → scaling → causal masking → dropout → multi-head.
- · Why attention replaced RNNs
- · Simplified self-attention with no trainable weights
- · Trainable Q, K, V and the √d scaling
- · Causal masking, dropout and multi-head attention
3. Module 3 · Implementing the GPT Model
Layer norm, GELU, the feed-forward expansion, shortcut connections, the transformer block, and the 124M-parameter assembly.
- · Layer normalisation and GELU
- · Shortcut connections and the transformer block
- · Assemble GPT-2 124M and generate (gibberish)
4. Module 4 · Pretraining on Unlabelled Data
Cross-entropy and perplexity, the training loop, train/val splits, temperature and top-k sampling, saving weights, and loading real GPT-2 weights.
- · Loss: cross-entropy and perplexity
- · The training loop
- · Decoding strategies: temperature and top-k
- · Load real OpenAI GPT-2 weights
5. Module 5 · Finetuning for Classification
Turn your GPT into a spam classifier: dataset balancing, replacing the output head, training only what matters, and honest metrics.
- · Dataset prep: the part everyone skips
- · Swap the head, freeze the body, train
6. Module 6 · Instruction Finetuning
Make the model follow orders: Alpaca-style prompt formatting, masking the prompt tokens, a custom collate function, and evaluating the result.
- · Format the data like Alpaca
- · Train, generate, evaluate
- · Where to go next: LoRA, RLHF, scaling