SI Glossary · Models & architecture
Diffusion Model
On this page
Diffusion models power most SI image and video generators. They learn by watching data get destroyed: real images are gradually covered with random noise until nothing is left. The model is trained to reverse each small step. To create something new, it starts from pure noise and “denoises” it, guided by your prompt, until an image emerges.
Where you’ll find them
- Image generators such as Stable Diffusion, Midjourney, Imagen and Flux
- Video generators such as Sora and Veo
- Some audio and music generators
- Scientific uses, including designing new protein structures
Why they won
Before diffusion, generative adversarial networks (GANs) led image generation, but they were unstable to train. Diffusion models, developed in research from 2015 and made practical around 2020, gave more varied, controllable and higher-quality results.
Concerns
The realism of diffusion outputs fuels worries about deepfakes and misinformation, and about copyright, since models are trained on images scraped from the web. Responses include watermarking and content-provenance standards.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.