AI Frontiers, part 28: Diffusion models β how image generation actually works
Part 28from the AI Frontiers series · 65 parts in all
The image side of this field has a different intellectual history from the language side, and by mid-2023 it had also produced a different kind of product. Stable Diffusion ran on a consumer GPU, its weights were downloadable, and an ecosystem of interfaces, fine-tunes and extensions grew up around it in months. Meanwhile image generation had become a genuinely contested public issue in a way that text generation had not, which is a tell about how much less work is required to judge an image than to judge a paragraph.
This entry is about the mechanism, which is more elegant than it is usually given credit for, and about the specific engineering decision that turned a research curiosity into something a graphics card could run.
Denoising, which is the whole idea
The starting point is a training objective that requires no adversarial game and no likelihood estimation. Take an image. Add Gaussian noise to it in many small steps until nothing but noise remains. Now train a network to reverse one step: given a noisy image and the step index, predict the noise that was added. Do that across all steps and you have a model that can walk backwards from pure noise to a plausible image.
The idea traces to Sohl-Dickstein et al., who framed it thermodynamically as reversing a diffusion process. Ho et al. turned it into a practical training recipe (Denoising Diffusion Probabilistic Models), and the reason it took another two years to become dominant is the one that applies everywhere in this field: sampling was slow, because a naive implementation requires hundreds or thousands of network evaluations per image. The competing approach at the time, generative adversarial networks, produced an image in one forward pass.
A parallel formulation arrived from a different direction and produced the same algorithm. Song and Ermon approached generation as estimating the gradient of the data density β the score function β and sampling by following it, and Song et al. later showed that the score-based and diffusion formulations are two views of the same stochastic differential equation. This is more than a historical curiosity. It is why techniques from one line of work transfer directly to the other, and why the field stopped arguing about which framing was correct and started using both vocabularies interchangeably.
The practical advantages over GANs are worth stating because they explain the field's migration. The training objective is a regression, so it is stable β no mode collapse, no delicate balance between two networks, no oscillation. The loss is a meaningful number rather than a proxy for perceptual quality. And the sampling procedure is controllable in ways that matter: you can start from an existing image and add noise to a chosen degree, which gives you image-to-image editing and inpainting essentially for free.
Guidance, which is how text gets in
A diffusion model trained as described generates images from the data distribution. It has no idea what you asked for. The mechanism that connects a text prompt to the sampling process is guidance, and it comes in two flavors that were both in use by 2023.
Classifier guidance (Dhariwal and Nichol) trains a separate classifier to predict the label of a noisy image and uses its gradient during sampling to push generations toward a class. It works, and it requires a second model, and it is awkward because the classifier must itself be robust to noise. Classifier-free guidance (Ho and Salimans) removes the second model entirely by training the same network on both conditional and unconditional objectives, then combining the two predictions at sampling time: run the model once with the text, once without, and extrapolate in the direction that increases the difference. A guidance scale controls how hard you push.
That single parameter explains most of what users experienced in 2023. Low guidance produces diverse, sometimes unfaithful images. High guidance produces images that match the prompt more literally while becoming oversaturated, over-smooth and prone to artifacts, because extrapolating beyond the data distribution amplifies whatever the model is most confident about. The tradeoff is not a tuning artifact; it is the same proxy-optimization story as reward model overoptimization, expressed as a slider on a web interface. Every user who turned it up until the picture broke was running that experiment.
The text encoder nobody talks about
Guidance needs a conditioning signal, and the signal comes from a text encoder β which is where the language side of this field re-enters. Stable Diffusion uses a CLIP text encoder, the same contrastive image-text model described in the multimodal entry. DALLΒ·E 2 paired a diffusion decoder with CLIP latents explicitly (Ramesh et al., "Hierarchical Text-Conditional Image Generation"), and its predecessor generated images directly from a discrete token sequence (Ramesh et al., "Zero-Shot Text-to-Image Generation").
This explains a large fraction of what the models got wrong in 2023. CLIP is a contrastive matcher trained on 400 million caption-image pairs, which makes it excellent at broad semantic alignment and poor at the compositional and literalism that prompts often demand. It does not read text in images. It is not sensitive to word order in the way a language model is, and attribute binding β which color belongs to which object when there are two of each β is notoriously unreliable. When a system renders a red cube and a blue sphere as a blue cube and a red sphere, the failure is not in the diffusion process. It is in the encoder that never had to represent that distinction.
The fix arrived in stages: better encoders, larger text models as conditions (Imagen, Parti), and architectural choices that give the language side more influence. But the lesson held: the image model inherits the compositionality limits of whatever reads the prompt, and improving the generator cannot repair a representation that never captured the relationship.
The decision that made it affordable
Latent diffusion (Rombach et al.) is the reason any of this ran on a consumer GPU. Instead of running the denoising process in pixel space β where a 512Γ512 image is 786,432 numbers and every sampling step costs a full-resolution pass β it compresses the image with a variational autoencoder into a smaller latent representation and diffuses there. A factor of eight in each spatial dimension reduces the sequence length by roughly sixty-four, and the arithmetic becomes tractable.
Two details make it work rather than merely being cheaper. The autoencoder's task is only compression and reconstruction, so it can be trained once and reused. And crucially, the perceptual objective used for the autoencoder was chosen deliberately over a pixel-wise one, because a pixel-perfect reconstruction wastes capacity on high-frequency detail that the diffusion model can regenerate and that human perception barely weighs.
The consequence for the ecosystem was enormous. A model that fits in a few gigabytes and samples in seconds on a consumer card is a model that anyone can run, fine-tune and extend. The same latent compression also made fine-tuning cheap enough for the community to produce thousands of specialized variants, and made LoRA β the parameter-efficient method described in part 3 β the standard way to teach a base model a new subject, style or face from a handful of images (Ruiz et al.).
Control, which is where it stopped being a toy
Text is a lossy interface for spatial intent. Describing a composition precisely is hard, and the model routinely satisfies the words while defying the arrangement. The 2023 answer was to add conditioning channels that carry structure directly: a depth map, a pose skeleton, an edge map, a segmentation mask.
ControlNet (Zhang et al.) is the clean implementation of that idea. Copy the encoder of a pretrained diffusion model, train the copy on a structural condition while freezing the original, and connect them with zero-initialized convolutions so that training starts as a no-op and degrades nothing. The result is a control channel you can bolt onto an existing model without retraining it, which is why it spread through the ecosystem in weeks and became the basis of most professional workflows involving image generation.
The broader pattern is worth noting: the techniques that produced the most adoption in this period were not new models. They were interfaces β a control net, a low-rank adapter, a compression scheme β that let people put their own intent and their own data into a capability that already existed. That is a much better predictor of what gets used than benchmark quality, and it recurs everywhere in this series.
Twenty steps instead of a thousand
The reason diffusion was impractical before 2022 was sampling cost. The original formulation needs as many as a thousand network evaluations to go from noise to image, at full resolution, per image. A generation that takes ten minutes on a research cluster is not a product.
Two lines of work fixed it. Deterministic samplers replaced the stochastic reverse process with a non-Markovian one that can skip steps, cutting a thousand evaluations to twenty or fifty with minimal quality loss β which is the single largest practical improvement in the area, and the least discussed outside the field. Distillation approaches then taught a model to emulate a slower sampler in fewer steps (Salimans and Ho), and consistency models (Song et al.) made the extreme version possible by training a network that maps any point on the noise trajectory directly to the output.
This is the same pattern as speculative decoding and every other serving optimization: the capability is decided by the research, and whether it becomes a product is decided by the latency. A model that takes ten minutes per sample is a demo. The same model at two seconds is a business, and nothing about the model's quality changed in between.
What the images were still bad at
It is worth recording the failure modes of this generation honestly, because they are the reason the next year's work looked the way it did. Faces and hands at small scale were unreliable, which is a high-frequency-detail problem and was partly a resolution problem. Compositional prompts with multiple objects and bound attributes failed regularly, for the encoder reasons above. Text rendering was close to impossible, because a latent space compressed for perceptual similarity discards exactly the fine structure that glyphs consist of. And long prompts were not followed proportionally better than short ones, which surprised people who assumed the text encoder was reading the way a language model reads.
The common thread is that each failure sat at a different layer β resolution, encoder, compression objective, training data β and the fixes came from the corresponding layer rather than from a bigger generator. This is worth internalizing as a general debugging habit for any multimodal system: when output is wrong, ask which component had to represent the thing that is missing before you ask for a bigger model of the last stage.
The licensing problem underneath
None of the technical elegance addresses the thing that made image generation contentious. The training data is scraped from the open web. LAION-5B (Schuhmann et al.) is a dataset of billions of image-text pairs assembled from links, released as an open resource for research, and it is the provenance of the models that became household names.
The questions that follow are not technical and do not have technical answers. Whether training on copyrighted work is fair use is being litigated. Whether a model that can reproduce a specific artist's style, or occasionally an almost-exact copy of a training image, is a derivative work is unsettled. Whether the licenses attached to model weights actually bind downstream users in the way their authors intend is unclear, and the fact that a model card says "for research use" does not obviously determine what a person may do with an image they generated.
What is clear is that the industry's response was mostly to add a compliance layer rather than to change the data: provenance metadata, content credentials, output filters, opt-out mechanisms for rightsholders. Whether that is adequate is a question for courts and legislatures rather than for engineers, and the honest engineering posture is to know exactly what your training data contains, to be able to say so, and to design as if the answer will eventually be demanded of you. That last habit is the one durable lesson of the period, and it applies with equal force to every model in this series.
Works Cited
Dhariwal, Prafulla, and Alexander Nichol. "Diffusion Models Beat GANs on Image Synthesis." arXiv, 2021, arxiv.org/abs/2105.05233. Accessed 15 June 2023.
Ho, Jonathan, and Tim Salimans. "Classifier-Free Diffusion Guidance." arXiv, 2022, arxiv.org/abs/2207.12598. Accessed 15 June 2023.
Ho, Jonathan, et al. "Denoising Diffusion Probabilistic Models." arXiv, 2020, arxiv.org/abs/2006.11239. Accessed 15 June 2023.
Nichol, Alex, et al. "GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models." arXiv, 2021, arxiv.org/abs/2112.10741. Accessed 15 June 2023.
Radford, Alec, et al. "Learning Transferable Visual Models From Natural Language Supervision." Proceedings of the 38th International Conference on Machine Learning, 2021, arxiv.org/abs/2103.00020. Accessed 15 June 2023.
Ramesh, Aditya, et al. "Hierarchical Text-Conditional Image Generation with CLIP Latents." arXiv, 2022, arxiv.org/abs/2204.06125. Accessed 15 June 2023.
---. "Zero-Shot Text-to-Image Generation." arXiv, 2021, arxiv.org/abs/2102.12092. Accessed 15 June 2023.
Rombach, Robin, et al. "High-Resolution Image Synthesis with Latent Diffusion Models." arXiv, 2021, arxiv.org/abs/2112.10752. Accessed 15 June 2023.
Ruiz, Nataniel, et al. "DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation." arXiv, 2022, arxiv.org/abs/2208.12242. Accessed 15 June 2023.
Saharia, Chitwan, et al. "Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding." arXiv, 2022, arxiv.org/abs/2205.11487. Accessed 15 June 2023.
Schuhmann, Christoph, et al. "LAION-5B: An Open Large-Scale Dataset for Training Next Generation Image-Text Models." arXiv, 2022, arxiv.org/abs/2210.08402. Accessed 15 June 2023.
Salimans, Tim, and Jonathan Ho. "Progressive Distillation for Fast Sampling of Diffusion Models." arXiv, 2022, arxiv.org/abs/2202.00512. Accessed 15 June 2023.
Sohl-Dickstein, Jascha, et al. "Deep Unsupervised Learning Using Nonequilibrium Thermodynamics." arXiv, 2015, arxiv.org/abs/1503.03585. Accessed 15 June 2023.
Song, Yang, and Stefano Ermon. "Generative Modeling by Estimating Gradients of the Data Distribution." arXiv, 2019, arxiv.org/abs/1907.05600. Accessed 15 June 2023.
Song, Yang, et al. "Consistency Models." arXiv, 2023, arxiv.org/abs/2303.01469. Accessed 15 June 2023.
---. "Score-Based Generative Modeling through Stochastic Differential Equations." arXiv, 2020, arxiv.org/abs/2011.13456. Accessed 15 June 2023.
Zhang, Lvmin, et al. "Adding Conditional Control to Text-to-Image Diffusion Models." arXiv, 2023, arxiv.org/abs/2302.05543. Accessed 15 June 2023.