In short

A diffusion model does not draw — it denoises. It starts from a cloud of random noise and refines it, step by step, until a coherent image emerges that matches the provided text. This inversion of a degradation process is the core mechanism behind DALL-E, Stable Diffusion, Midjourney, and FLUX. It produces visually impressive results, but comes with persistent structural limitations and an unresolved legal dispute over training data.

In short: imagine a sculptor who does not carve marble but restores it. They are presented with a block damaged by acid rain — unrecognizable. With each chisel stroke, they “remove” the erosion to bring forth the statue. Except the statue never existed: it is born from the restoration process itself, guided by the sentence that was whispered to the sculptor. This is the principle of diffusion: generating by denoising, rather than by building.


Denoising as a generative principle

Imagine a photograph progressively buried under snow. Flake by flake, the image disappears. At the end, only Gaussian noise remains — a random texture with no signal.

A diffusion model learns to reverse this process. During training, it observes thousands of intermediate stages of this degradation and learns to predict, at each step, which direction removes the most noise. At inference time, it is given a cloud of pure noise and applies this inversion step by step. After 20 to 100 steps depending on the implementation, an image emerges.

This process is called reverse diffusion or denoising. The final quality depends on two factors: the precision of the neural network guiding each step, and how the provided text conditions the direction taken.

The role of text

The text prompt arrives in the model as a numerical vector — an embedding produced by a language encoder (CLIP or T5 depending on the architecture). This embedding steers each denoising step: the neural network does not denoise in a vacuum; it denoises toward something consistent with the description.

The CFG scale (classifier-free guidance) parameter controls the intensity of this steering. A high value (7 to 12) produces images that closely match the text but are less varied. A low value (1 to 3) gives the model more creative freedom at the expense of prompt fidelity.

The latent space: why it matters

Denoising pixel by pixel on a 1024×1024 image would be computationally prohibitive. The dominant solution is the Latent Diffusion Model (LDM): the process operates in a compressed space, called latent space, produced by a variational autoencoder (VAE). The image is first compressed into roughly a 64×64 representation, denoising takes place there, and a decoder then reconstructs the final image at full resolution. Stable Diffusion is the emblematic implementation of this paradigm, and this is what made generation accessible on consumer-grade hardware.

In short: going from 1024×1024 (≈ 1 million pixels) to a 64×64 latent space (≈ 4,000 values) is a 256× compression. Without this trick, generating an image in 50 steps would take hours on a consumer GPU; with latent diffusion, the same image is produced in 5-15 seconds on a €500 card. This is the innovation that democratized generation in 2022.

The architectural evolution: from U-Net to DiT

Until 2022, the standard backbone for diffusion models was the U-Net: an encoder-decoder architecture with skip connections between levels, well-suited to spatial data. SD 1.x and SDXL rest on this foundation.

Since 2023, a transition toward Diffusion Transformers (DiT) has taken shape. The idea: replace the U-Net with a pure transformer. The latent image is divided into small blocks (patches) that become tokens — exactly like words in a language model. The transformer then processes these tokens in parallel.

Results published by Peebles & Xie (ICCV 2023) are clear: DiT-XL/2 achieves an FID of 2.27 on ImageNet 256×256, surpassing all previous diffusion models, at a computational cost under 20% of that of U-Net architectures at the same performance level. The decisive property is log-linear scalability: the more compute invested, the more predictably FID improves. This property enabled text LLMs to scale — it now applies to images.

FLUX.1, Midjourney V6 and V7, and SD3 all use DiT or DiT-hybrid architectures.

Flow matching: an alternative to Gaussian denoising

Alongside DiTs, another approach is gaining traction: flow matching. Instead of learning to reverse a Gaussian noising process, the model learns a velocity field that directly transports a noise distribution toward an image distribution. FLUX.1 and Stable Diffusion 3 use rectified flow matching, which produces more direct trajectories and requires fewer inference steps — translating into faster generation.

The ecosystem: four distinct strategies

The image generation market today rests on four very different positioning strategies.

OpenAI has integrated image generation directly into GPT-4o. Since March 2025, gpt-image-1 is no longer a separate model but a native capability of the multimodal system. The advantage: the underlying VLM carries world knowledge that improves understanding of complex prompts and the legibility of text within images. The trade-off: a fully closed system with API pricing.

Midjourney remains the aesthetic reference for a community of designers and illustrators. The V7 model (April 2025) produces visually premium results. Its internal architecture has never been publicly documented — DiT is presumed but unconfirmed. Distribution exclusively via Discord and a web interface, no open weights.

Stability AI built the dominant open-source ecosystem with the public weights of SD 1.x, SDXL, SD3, and SD3.5. The community has produced tools such as ControlNet (conditioning on pose or depth maps), thousands of LoRA adaptations (for lightweight style or character customization), and interfaces like ComfyUI or Automatic1111. This ecosystem is what makes Stable Diffusion indispensable for customization — though Stability AI has progressively tightened its licensing terms since SD2.

Black Forest Labs is the actor to watch. Founded in 2024 by the creators of Stable Diffusion (Robin Rombach, Andreas Blattmann, Patrick Esser), backed by a16z, it releases FLUX.1 as open weights under the Apache 2.0 license. FLUX.2 (November 2025) pushes the architecture to 32 billion parameters by coupling a rectified flow transformer with the Mistral-3 24B VLM — an approach close to OpenAI’s but with open variants. Quality rivals closed systems across several benchmarks.

Adobe Firefly occupies a distinct niche: exclusively licensed training data, native integration in Photoshop and Premiere, and contractual guarantees of commercial use without legal risk. Firefly 5 (October 2025) supports 4 megapixels. It is the reference choice for teams operating under strict legal constraints.

Worth noting: Ideogram stands out for text generation within images, an area where all models still struggle.

Decision matrix — which model for which need

Usage contextRecommendationWhy
Commercial production, strict legal constraintsAdobe FireflyLicensed data, contractual guarantee of commercial use without lawsuit.
Strong customization (ControlNet, LoRA, custom style)Stable Diffusion / FLUX.1 (open weights)Mature open source ecosystem, thousands of LoRAs available, local deployment.
Premium aesthetic quality, visual brand identityMidjourney V7Established reputation among designers, refined artistic rendering, but closed system.
Legible text in image (logo, slogan, capture)Ideogram, GPT-4o image, FLUX.1Multimodal architectures with improved text encoding; others remain fragile.
Multimodal generation integrated in LLM workflowGPT-4o image, FLUX.2Underlying VLM understands complex prompts; no separate model to orchestrate.
Research, prototyping with local GPU compute budgetStable Diffusion 3.5 + ComfyUIOpen source, huge tools community, free except electricity.

Persistent limitations

Anatomy and physics

Diffusion models are generators of statistical distributions. They have no internal model of anatomy, physics, or perspective. Errors involving hands — extra fingers, aberrant proportions — remain frequent even on 2025–2026 models. Complex medical structures are particularly poorly rendered. Scenes containing more than twenty distinct objects pose coherence difficulties.

Multimodal architectures that couple a VLM with the generator (GPT-4o image, FLUX.2) deliver real improvements because the VLM encodes world knowledge — but they do not fully correct these structural flaws.

Text within images

For a long time, generating legible text inside an image was nearly impossible for diffusion models. Progress has been notable since 2024 (Imagen 3, GPT-4o image, Ideogram), but errors persist: missing characters, language mixing, distortion in corners. The issue stems from the very nature of denoising: letters are treated as visual shapes, not as meaning-carrying symbols.

Temporal coherence

Extending diffusion models to video runs into the absence of coherence between frames. An image generator produces each frame independently — it does not know that the character in frame 42 must have the same face as in frame 41. Video DiT architectures (Sora, Veo, Wan) add a temporal dimension to the tokens, but this problem is not entirely solved.

The copyright dispute is the most structurally significant debate in the sector, and the least resolved.

The central question: does training a model on copyrighted works without a license constitute fair use under US law, or infringement? The Andersen v. Stability AI case — brought by visual artists against Stability AI, Midjourney, and DeviantArt — is moving toward the discovery phase following Judge Orrick’s decision in August 2024, with a trial scheduled for September 2026. Google has faced a similar action since April 2024.

AI-generated works, for their part, receive no copyright protection in the United States: the Copyright Office reaffirmed in 2023–2024 that only significant human creative contribution to the process is protectable. An image produced entirely by prompt without documented human creative intervention is not protectable.

The Generative AI Copyright Disclosure Act passed in 2024 imposes transparency on training data, but had not yet become law at the time this article was written.

The situation creates an asymmetry: models trained on licensed data (Adobe Firefly) offer legal coverage; others operate in a gray zone that will likely be resolved through case law rather than legislation.

In short: for professional use, the criterion “was the model trained on data with clear rights?” sometimes becomes more important than raw aesthetic quality. The 2026 Andersen trial could retroactively impose a licensing obligation on models trained on scraped web data — that is, the majority of the current market.


Key takeaways

  • Image generation by diffusion rests on learning to invert noise: the model learns to denoise, not to “draw.”
  • The transition from U-Net to Diffusion Transformers marks a scalability breakthrough, with documented computational efficiency gains.
  • Four strategies coexist: closed multimodal integration (OpenAI), closed premium artistic quality (Midjourney), open-source ecosystem (Stability AI / Black Forest Labs), and legal certainty (Adobe Firefly).
  • Structural limitations — anatomy, text, temporal coherence — persist across all current models.
  • The legal framework on training data copyright is not settled. A 2026 trial could reshape the rules for the entire sector.