Module 10 of 12

Generative AI & LLMs

Lessons

About This Module

Generative AI is the family of models that create new content: text, images, audio, and video. Large language models (LLMs) like the ones behind modern chatbots are trained to predict the next word in a piece of text, and doing that at enormous scale turns out to produce surprisingly capable systems.

The lessons start with a short explainer of what an LLM is, then open up the transformer, the architecture underneath. Next is a one-hour talk on how LLMs are trained and used, and the module finishes with diffusion models, the technique behind AI-generated images and video.

Watch the lessons in order. If a video runs long, feel free to treat it as a reference you dip back into later rather than something to finish in one sitting.

Lessons

4 videos
01

Large Language Models explained briefly

3Blue1Brown's short animated explainer: an LLM is a giant function that predicts the next word, and its parameters are tuned by training on huge amounts of text.

Video 8min
02

Transformers, the tech behind LLMs

A visual walkthrough of the transformer, the architecture behind ChatGPT-style models: tokens, embeddings, and how data flows through the network to predict the next word.

Video 27min
03

Intro to Large Language Models

Andrej Karpathy's general-audience talk on what LLMs are, how they are trained (pre-training and fine-tuning), what they can do, and the security challenges they bring.

Video 1h
04

But how do AI images and videos actually work?

A guest video by Welch Labs on diffusion models and CLIP: how AI learns to turn a text prompt into an image or a video.

Video 37min