The Useful Question
English

How does a diffusion model turn noise into a red mug?

A diffusion model learns how to move noisy data toward images, then uses that learning to make new pictures. Training adds noise to known examples; generation starts from a random state and repeatedly applies a trained model. The finished picture is not already hiding inside the static. Separating those two stages is the useful first step.

When does the mug become a mug?

Imagine requesting “a red mug, handle on the right.” You have not specified a square or rounded base, a wooden or painted table, or the direction of the light. A usable image still needs to settle those details.

That request is our invented example, not a prompt we ran. It helps identify the puzzle: randomness supplies a starting point, but learned image patterns and the text condition shape what happens next. Noise alone does not contain instructions for drawing a mug.

Training makes a problem with a known answer

In a representative diffusion setup, researchers take an image, choose a noise level, and mix in random noise. The model receives the noisy result and information about that level. One common training objective asks it to predict the added noise. Because the training process knows what was added, it can measure the error and adjust the model.

Repeating this across examples and noise levels teaches useful predictions for different stages of corruption. Training can construct an intermediate noisy state directly; it does not have to play the entire corruption sequence every time. The DDPM paper’s training and sampling algorithms make that distinction explicit.

Think about our mug at different levels of visibility. If its outline is mostly intact, repairing local detail is a different problem from working with barely recognizable structure. This comparison is an intuition aid, not a claim that the model literally names a rim or follows a human drawing lesson.

An upper row adds noise to an example image for training; a lower row updates fresh noise toward a different invented mug. Neither row is a recorded model run.
Original explanatory diagram created for this article
Expand the diagram

Scroll within the diagram horizontally; arrow keys work when focused.

An upper row adds noise to an example image for training; a lower row updates fresh noise toward a different invented mug. Neither row is a recorded model run.
Open the diagram at full size

This original schematic uses drawn mugs and patterned squares. It does not reproduce model outputs, latent values, an actual noise schedule, or a required number of sampling steps.

Generation uses what was learned

To create a new image, a typical diffusion sampler begins with random noise and updates that state using the trained model’s predictions. It repeats the process toward an image. “Removing noise” is a helpful shorthand; the exact update rule, including whether it introduces additional randomness, depends on the method.

For text-conditioned generation, an encoder converts the prompt into numerical representations the system can use. “Red,” “mug,” and “handle on the right” help guide the calculation rather than becoming labels pasted onto the canvas. Google’s Imagen paper describes one text-conditioned diffusion system and the role of its language representations.

Generation is not normally trying to recover one particular training photograph. It is producing a sample consistent with what the model learned and the conditions supplied. Our underspecified tabletop still has to acquire a color. That is one reason the request does not define a single correct picture. Whether the handle actually lands on the right remains something to inspect in the result.

Often, the work happens in a smaller representation

Repeatedly processing every pixel of a large picture is expensive. The Latent Diffusion paper describes encoding images into a smaller learned representation, doing diffusion there, and decoding the result into a visible image.

This “latent” is numerical data, not a list containing entries such as “one handle” and “one table.” Nor is latent diffusion simply enlarging a finished thumbnail: generative computation happens within that representation. Our visible mug sketches stand in for a process that may operate on values you could not directly view as an ordinary photograph.

The names describe different choices

“Diffusion” is not a complete explanation of every image generator. Training targets, model structures, and sampling methods vary.

Stable Diffusion 3, introduced in 2024, provides a useful historical example. It uses a Rectified Flow formulation related to flow matching: training connects noise and data along a path and learns how the state should move. That differs from explaining every system as predicting an added noise pattern at every stage. See Stability AI’s technical explanation and the SD3 research paper. SD3 is an example here, not a claim about the newest model available.

Autoregressive image models provide another approach: predict successive image-representation tokens, conditioned on what came before. Taming Transformers demonstrates that route.

These terms are not all competing boxes. “Latent” describes a representation in which computation happens; diffusion and flow matching concern the learning and generation process. A transformer can also appear in more than one kind of system. A product name alone does not settle the architecture.

A generated picture does not reveal its whole history

Describing diffusion solely as finding and pasting together existing pictures misses the learned generation process. But computation does not guarantee that training examples cannot be memorized. Researchers have extracted training images from particular diffusion models. That finding does not mean every generated image is a duplicate.

For someone using a picture, visual quality and provenance are separate checks. A convincing red mug does not disclose the full training collection or establish permission for every intended use. Review source-material permissions and the service’s applicable terms rather than treating “AI-generated” as a complete rights statement.

The key distinction is simple: training creates noisy learning problems; generation uses learned predictions to construct a new sample. If you also wonder why convincing AI text can lack adequate evidence, see the related article on confident wrong answers.

Sources and limits

Primary sources checked October 11, 2026, Japan time. This is a principles explainer, not a hands-on generation test or a current-model ranking. The diagram is original explanatory artwork. Recheck if cited research is corrected or new model-specific claims are added.