But How Do AI Images and Videos Actually Work? Diffusion Models Explained

This is a guest video on 3Blue1Brown, commissioned by Grant Sanderson during a deliberate break from his own channel. Rather than pause uploads entirely, Sanderson redirected Patreon support toward creators he personally respects, and the first of these guest episodes comes from Stephen Welch of Welch Labs, a channel Sanderson has described as a true kindred spirit to his own. That endorsement carries weight in a community that values rigor over hype.

Welch uses the video to go well beyond the standard one-line summary of diffusion models, walking through the actual mechanics of how CLIP connects text and images, and how the math of turning a written prompt into a coherent picture or video actually works. It's a deeper, more visual treatment of the technology sitting behind nearly every modern AI image and video generator, from Midjourney-style tools to text-to-video systems, aimed at viewers who already have some technical grounding but have never seen the "removing noise" explanation actually justified.

For anyone using or building on top of generative image and video tools, understanding what's happening under the hood changes how you think about their limits and their outputs. This video treats that curiosity seriously, offering the kind of rigor that turns a black box into something you can reason about.

Who is this for:

Intermediate tech audiences, AI enthusiasts, and creative professionals who use AI image or video tools and want to understand how they actually work.