How AI Image Generators Work

Diffusion: turning noise into pictures. Many AI image generators use diffusion: they learn to remove noise step by step until a picture appears that matches your text prompt.

AI-generated street scene in Spain created with Stable Diffusion

An image made with Stable Diffusion: Stable Diffusion/Tullius Detritus, Public domain

Start with static

Imagine taking a photo and adding a little random noise, like TV static, then a little more, over and over until nothing is left but static. A diffusion model learns to run that film backward. Shown a noisy image, it predicts what noise was added so it can be removed, revealing a slightly cleaner picture.

To generate something new, the model starts from pure random static and repeats that cleanup many times. With each step, rough shapes appear, then colors, then fine details. The number of steps is a setting in many tools, and more steps usually means more detail up to a point.

Video: But how do AI images and videos actually work? | Guest video by Welch Labs (3Blue1Brown), embedded from YouTube.

Where the words come in

A model that only cleans up noise would produce random pictures. To follow a prompt, the system also uses a text encoder, a network trained on huge numbers of images paired with captions, that turns your words into numbers capturing their meaning. During every denoising step the model is steered toward images that fit that meaning.

Many tools also use a technique called guidance, which compares what the model would draw with and without your prompt and pushes the result further toward the prompt. Turn it too high and images look harsh or overcooked, too low and they ignore your words.

Working in a compressed space

Cleaning up every pixel of a large image is expensive. Latent diffusion, the approach behind Stable Diffusion, released publicly in 2022, first squeezes images into a smaller compressed form, does the denoising there, then expands the result back into full pixels. That made high-quality generation fast enough to run on a good home computer. Newer systems mix diffusion with Transformer networks and extend the same ideas to video.

Using them well and fairly

Image generators are useful for mood boards, concept art, illustrations and quick mockups. They still struggle at times with hands, readable text and exact counts, though each generation improves. Be specific in prompts about subject, style, lighting and composition.

There are real questions to respect. Models are trained on large image collections, and artists have raised concerns about consent and credit. Laws on copyright and AI output differ by country and are still developing. Never use these tools to make deceptive images of real people, and label AI-made images when it matters to your audience.

Media credits
  • An image made with Stable Diffusion: Stable Diffusion/Tullius Detritus, Public domain
  • More denoising steps, more detail: Stable Diffusion WebUI - Automatic1111, Public domain

Text written by Strawberry Lemonadai.

← Large Language Models ExplainedGPUs: Why AI Runs on Graphics Chips →

Shop

The Strawberry Lemonadai collection

Tees and hoodies in the Strawberry Lemonadai colors, printed to order in the USA. Use code FIRST15 from the newsletter for 15% off your first order.

Watch

Strawberry Lemonadai Shorts

Quick, accurate, under a minute. Every Short is original and made only for this channel.

YouTube @StrawberryLemonadAI
Channel launching soon. Subscribe on the newsletter to hear first.
Get notified
A computer beat the world chess champion in 1997Photos: IBM Deep Blue at Computer History Museum (9361685537).jpg - Anton Chiang from Cupertino, CA, USA (CC BY 2.0) via Wikimedia Commons | Garry Kasparov (37097592314).jpg - Gage Skidmore from Peoria, AZ, United States of America (CC BY-SA 2.0) via Wikimedia Commons | Chess game Staunton No. 6 perfil view 8.jpg - Wilfredor (CC0) via Wikimedia Commons | Chess game Staunton No. 6.jpg - Wilfredor (CC0) via
An AI solved a 50 year biology puzzle and won a Nobel PrizePhotos: Human Oxy-Hemoglobin Protein.jpg - PDB code 2DN1 1.25 a resolution crystal structures of human (CC0) via Wikimedia Commons | Ribbon diagram of the DED.jpg - BQUB16-Oibanez (CC BY-SA 4.0) via Wikimedia Commons | Laboratory pipettes.jpg - J.N. Eskra (CC BY-SA 4.0) via Wikimedia Commons | Use of a Multichannel Pipette for High-Throughput Liquid Handling in a Biosafety Cabinet.jpg - Siduduziwe Nxumal
ChatGPT hit an estimated 100 million users in 2 monthsPhotos: BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Diverse people using phones.jpeg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Crowd of people with phones.jpg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Sam Altman CropEdit James Tamim.jpg - TechCrunch (CC BY 2.0) via Wikimedia Commons | Ilya Sutskever and Sam Altman in TAU.jpg - Eladkarmel (CC B
The T in ChatGPT comes from one 2017 paperPhotos: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg
The word robot was invented for a 1920 playPhotos: Karel Čapek podepisuje první výtisky Povětroně, Pestrý týden 27.1.1934.jpg - Unknown authorUnknown author (Public domain) via Wikimedia Commons | Karel Čapek 30.léta.jpg - re-photo by David Sedlecký (Public domain) via Wikimedia Commons | Plakat za predstavo R.U.R Rossmus Universal Robots v Narodnem gledališču v Mariboru 28. oktobra 1933.jpg - Unknown authorUnknown author (Public domain) via Wikim
Newsletter

The weekly AI digest

The AI stories that actually matter this week, explained in plain English, plus one useful tool tip. No hype, no jargon, unsubscribe anytime.