ScruTool
Technology

How AI Image Generators Actually Work in 2026

Learn how AI image generators work in 2026, from diffusion and prompts to autoregressive and hybrid models, plus why AI still struggles with text and hands.

Aug 24, 2026 14 min read

Type a sentence. Wait four seconds. Receive a photograph of something that has never existed.

Most explanations of that process stop at one line: the model starts with random static and cleans it up. That line is correct, and it explains almost nothing. It does not tell you why the machine understood the word "melancholy," why six fingers used to appear on every hand, why letters came out as gibberish for years, or why the model currently sitting at number one on the public leaderboards does not use that process at all.

This article follows a single prompt through the machine, stage by stage. No equations. By the end you will know what happens between your keystroke and the picture, and why the answer changed in 2026.

The 30-Second Version

Four stages, and the rest of this article is a slow walk through each one.

•      The model spends months studying billions of image-and-caption pairs until words and pictures share a single mathematical map.

•      It then learns to destroy images by burying them in static, and, more usefully, to reverse that destruction.

•      To generate, it starts with fresh static and reverses the destruction it never actually performed, guided at every step by your prompt.

•      A decoder converts the result from a compressed internal format back into pixels you can see.

Stage one is where the strangeness begins, so start there.

Stage One: Teaching a Machine What Words Look Like

A neural network has no concept of "melancholy." It has numbers. The first problem anyone building an image generator has to solve is translation: turning a word into a set of coordinates that also happen to describe pictures.

The solution arrived in 2021 with CLIP, a model OpenAI trained on 400 million image-and-text pairs scraped from the open web. CLIP was given a deceptively simple job. Shown a batch of images and a batch of captions, it had to work out which caption belonged to which image. Nothing else.

Doing that job well requires building a shared space where the caption "a melancholy dog in the rain" and an actual photograph of a melancholy dog in the rain end up as nearly identical sets of numbers. That shared space is the foundation everything else sits on. Your prompt does not instruct the model. It supplies a destination coordinate.

The scale that made it possible

The open-source community answered CLIP with LAION-400M, then with LAION-5B: 5.85 billion image-text pairs, of which 2.32 billion carry English captions. Stable Diffusion was trained on a subset of that second dataset.

Scale here is doing something specific. A model that has seen four hundred photographs captioned "golden hour" has memorised four hundred photographs. A model that has seen four million has extracted a concept: warm light, low sun angle, long shadows, softened edges. The concept is what survives, and the concept is what you are addressing when you type.

Stage Two: Why Everything Starts as Static

Having a destination coordinate does not produce an image. Something has to build the picture, and the method that dominated from 2022 onward came from an unlikely place: thermodynamics.

Picture a drop of ink in a glass of water. It disperses. Given enough time, the water is uniformly grey and the original shape of the drop is unrecoverable. The 2015 paper that introduced diffusion to machine learning asked what would happen if a network learned to run that dispersal backwards.

Training: learning to break things

During training the model takes a real photograph and adds a measured amount of random noise. Then more. Then more again, across hundreds of steps, until the image is indistinguishable from television static. This part is trivial, because destroying information is easy.

The valuable part is the record. At every step the system knows exactly which noise it added. So it can train a network on a narrow, learnable question: given this noisy mess, what noise was added to it?

Generation: running the tape backwards

Now flip it. Hand the trained network pure static that was never an image at all. Ask the same question: what noise is in here? The network answers, the system subtracts that noise, and what remains is a fraction less random than what went in. Repeat twenty to fifty times.

Here is the misconception worth killing. The model is not uncovering a picture hidden inside the noise. Nothing is hidden there. Each step is a prediction, and the accumulated weight of thousands of small predictions is what produces an image. A different starting noise pattern, called the seed, produces an entirely different picture from the identical prompt. That is why reusing a seed is the closest thing these tools have to a repeat button.

Stage Three: The Compression Trick That Put This on Your Laptop

Denoising works. Denoising a full-resolution photograph fifty times over is also ruinously expensive, and in 2021 that cost confined the technology to research labs with large GPU clusters.

The fix came from Rombach and colleagues in the latent diffusion paper, and it is the single reason you can run one of these models on a gaming PC. Their observation: most of the data in an image is perceptual detail nobody would miss. Compress the image first with an autoencoder, run the whole noisy business in that compressed space, then decode back to pixels at the final step.

Roughly 48 times fewer values pass through the network on every one of those fifty steps. The paper measured a speedup of at least 2.7 times in training and sampling throughput, with image quality that improved rather than degraded.

That compression is the reason "Stable Diffusion" carries the word latent in its formal name, and it is why an entire hobbyist ecosystem exists. The technique moved image generation from a data-centre problem to a consumer-hardware one.

Compression solved the cost. It did not yet explain how your sentence steers the outcome, which is the next stage.

Stage Four: How Your Prompt Grabs the Wheel

A denoiser left alone will produce something. It will not produce what you asked for. Your prompt enters through a mechanism called cross-attention, and this is the point where the coordinate from stage one finally does its work.

At every denoising step, the network compares the patch of image it is currently working on against the encoded prompt, and asks which words are relevant right here. Painting the region where a face is forming, it attends heavily to "melancholy" and "dog." Painting the background, those words fade and "rain" takes over. The prompt is consulted continuously, not once at the beginning.

Two dials that explain most prompt behaviour

ControlWhat it actually does
SeedSelects the starting noise pattern. Same seed plus same prompt reproduces the same image. Change it and you get a different composition from identical words.
Guidance scale (CFG)Sets how hard the model is pushed toward your prompt at each step. Low values drift and invent. High values obey rigidly and produce oversaturated, brittle images. Most tools default to a middle value for good reason.

Cross-attention also explains a frustration you have probably hit. Ask for "a red cube and a blue sphere" and you may receive a blue cube. Attention distributes prompt influence across the canvas without a strict binding between an adjective and the object it modifies. Word order and emphasis shift that distribution, which is the entire mechanical basis of prompt engineering.

The 2026 Plot Twist: Diffusion Lost Its Monopoly

Everything above describes how a diffusion model works, and until recently that description covered the entire field. It no longer does. This is the part most explainers have not updated.

First crack: the U-Net was replaced

The denoising network in early Stable Diffusion was a U-Net, an architecture borrowed from medical image segmentation. In 2023 researchers swapped it for a transformer, the same family of architecture that powers language models. The Diffusion Transformer, or DiT, scaled more cleanly with added compute. Nearly every serious 2026 model uses a transformer backbone, including Krea 2, a 12-billion-parameter model released in June 2026 that merges text and image tokens into a single stream.

Second crack: some models stopped diffusing entirely

In March 2025 OpenAI shipped image generation inside ChatGPT and had, by its own published figures, more than 130 million users creating over 700 million images in the first week. The model behind it was not DALL-E and was not a diffusion model. It was autoregressive: it predicts image tokens one after another, the way a language model predicts words. That lineage became GPT Image 1, then 1.5, then GPT Image 2 in April 2026. DALL-E 2 and DALL-E 3 were retired outright.

Why the switch matters in practice: an autoregressive model reads your instruction with the same machinery it uses to read a sentence, so it follows long, conditional, awkwardly worded prompts far better. It also holds a conversation about the image across turns.

Third crack: hybrids

The two approaches are now being combined. GLM-Image, open-sourced by Zhipu AI and Huawei in January 2026 under an MIT licence, pairs a 9-billion-parameter autoregressive transformer that decides layout and meaning with a 7-billion-parameter diffusion decoder that renders fine detail. Tencent's HunyuanImage 3.0 pushes further, using a unified autoregressive framework that handles generation and editing without swapping pipelines, inside an 80-billion-parameter mixture-of-experts model with roughly 13 billion parameters active per token.

ApproachHow it builds the imageCharacteristic strength
DiffusionRefines the whole canvas at once, repeatedlyPhotorealistic texture, mature open-source tooling
AutoregressivePredicts image tokens in sequenceInstruction following, conversational editing
HybridAutoregressive layout, diffusion renderingDense text and information-heavy layouts

 

GPT Image 2 added something none of the earlier stages required: a reasoning pass before generation begins. It plans the composition, verifies object counts, checks the constraints in your prompt, and inspects its own output before returning it. That step is aimed squarely at the failures described next.

Where the Four Stages Still Fail

Every failure mode these tools are mocked for follows directly from the four stages above. None of them are random.

Hands

A diffusion model refines the entire canvas simultaneously. The region rendering one finger has no reliable channel to the region rendering the next, and no counter tracking how many have been drawn so far. Hands also appear in training data at wildly varying angles, folded, gripping, partially hidden. The model learned an average of hand-shaped things rather than a rule stating five.

Text

Letterforms are unforgiving. A texture that is 95 percent correct reads as fine; a letter that is 95 percent correct reads as gibberish. Early models treated text as decorative shape rather than symbol, which is exactly what you would expect from a system optimised on perceptual similarity.

This one is largely solved, and the fix is instructive. GLM-Image attaches a dedicated glyph encoder to the diffusion decoder and reports 91.16 percent word accuracy on the CVTG-2K text-rendering benchmark. Fixing text required a purpose-built component, not more training data.

Counting and spatial placement

OpenAI's own documentation for GPT Image 1 listed the model's weak spots plainly: non-English text, small type, rotated type, counting, and spatial localisation such as the position of pieces on a game board. "Seven birds" is a numeric fact; nothing in the denoising loop tallies. GPT Image 2's planning pass exists because the only durable fix was to add an explicit checking step outside the generation process.

What This Means the Next Time You Open One

The mechanics are only worth knowing if they change what you type. Four things follow from them.

•      A vague prompt is not a small prompt. It is an ambiguous destination coordinate, and the model resolves the ambiguity using the most statistically common interpretation in its training data. Specificity is not politeness toward the machine, it is a narrower target.

•      If an image is nearly right, change the seed before you rewrite the prompt. Composition often comes from the noise pattern rather than your wording.

•      Match the architecture to the job. Long conditional instructions and multi-turn edits suit an autoregressive model. Texture-led photorealism and heavy customisation still favour the open diffusion ecosystem.

•      Winning at generation does not mean winning at editing, and the leaderboards separate the two for a reason.

GPT Image 2 tops the text-to-image board by a wide margin and sits third on editing, behind models that trail it badly at generation. Generating from noise and surgically altering an existing image are different problems, and no single architecture currently owns both.

Price has decoupled from quality along the same lines. GPT Image 2 runs at roughly 211 dollars per thousand images through the API against 67 dollars for Nano Banana 2, a model that renders in four to six seconds and holds character consistency across a workflow. The expensive model is not three times better. It is better at a specific class of complex prompt, and worse value for high-volume work.

The label most people never look at

One consequence of the pipeline rarely gets mentioned. Because generation happens inside a controlled system rather than a camera, the major vendors can stamp provenance data into the file on the way out. Output from OpenAI's image models carries C2PA metadata identifying it as machine-generated. Google's Nano Banana 2 ships every image with SynthID alongside C2PA by default.

That metadata survives until someone screenshots the image or strips the file, which is roughly the first thing that happens to anything shared socially. Reading the file properties of an original download is more reliable than squinting at fingers, and it costs nothing to check.

The eleven-year arc from stage two to today has one direction. Each generation replaced a piece of learned statistical intuition with something more like deliberate structure: a shared embedding space where raw pixels used to be, a transformer backbone where a U-Net used to sit, a glyph encoder where letterforms used to smear, and a planning pass where object counts used to drift. The static is still there at the start of every image. There is now considerably more architecture standing between it and what you receive.

Community

Discussion

Join the discussion and share your perspective.

Related Articles