Stable Diffusion 3: how an AI builds an image

Duration 4:29

Stability AI released Stable Diffusion 3 in March 2025, built on an architecture called rectified flow transformers. This film explains how diffusion actually works: starting from an image of pure noise, then removing that noise step by step under the guidance of text. It covers what's new in version 3.0 — generating 3D models from flat images, presented at the ICML conference in July 2025 — along with the key players in the field: OpenAI's DALL·E, Midjourney, and Adobe's Firefly. Two challenges remain unresolved: the copyright status of generated images, and quality that is still imperfect.

Chapters

Every chapter below is clickable.

Transcript

In March 2025, Stability AI released Stable Diffusion three, an image generator. It rests on a specific architecture: rectified flow transformers. These tools turn a written sentence into an image. You type a description, and the model produces a matching visual. Fidelity to the text was long the real challenge. Stability AI is not alone in this field. OpenAI offers DALL·E, Midjourney offers its model, Adobe offers Firefly. Four competing approaches to the same idea. Stability AI chose open source from the start. The model is open; anyone can inspect and improve it. That decision accelerated its adoption. How is an image actually born? It all starts from an image of pure noise, entirely random. No shape is visible yet, just grain. In parallel, the written prompt is encoded. The text becomes a representation the model can handle. It will guide everything that follows. The model then removes a little noise. Then a little more, step by step. With each pass, the image grows sharper. At each step, the prompt guides what is removed. The noise does not vanish at random. It fades in the direction of the requested text. This cycle repeats for enough steps. Then the final image appears, sharp. The starting noise has become a coherent composition. Two settings really matter here. The number of steps, and the strength of the guidance. The stronger the guidance, the closer the image sticks to the prompt. This is where the architecture of Stable Diffusion three comes in. Rectified flow transformers shorten the path between noise and image. The result: more fidelity, and more speed. Version three brings a notable new feature. It generates three-dimensional models from flat images. An image becomes a manipulable object. This method was presented in July 2025, at the ICML conference. It extrapolates dimensions absent from the source image. Industrial design and animation benefit from it. The uses are already concrete in entertainment. Film and video game studios generate sets and concepts. Production times are shortened as a result. Personalization opens another path. Anyone can produce visuals tailored to their exact preferences. Custom content becomes quick to obtain. The recent integration reaches virtual reality. The model no longer creates only still images. It composes entire environments you can move through. These advances raise an unresolved question. Who holds the rights to an image produced by an AI? Regulators and industry are still seeking clear rules. Quality remains a second obstacle. The result does not always match the intended request. Technical flaws remain, despite the progress. Two forces push these models forward. The available computing power, and the refinement of algorithms. Collaboration between researchers and industry does the rest. In the longer term, a shift is taking shape. We will no longer only consume digital content. We will create entire worlds, personalized and immersive. Image creation is not Koaee's business. Yet we follow these generative architectures very closely. They remind us what a fast and faithful model demands. Our work is conversational assistants for businesses. We build them with that same technical standard. The same rigor, applied to conversation rather than images. The full article is in the video's comments. The film skims; the text details everything. Subscribe if AI interests you, and like this video. See you soon on koaee.ai.

Read the full article

Conversational assistants for business