OpenAI shipped ChatGPT Images 2.5 on 2026-09-08 and put an Image prompting guide in the developer docs the same day. The launch post covers the product features, the guide covers how to write prompts. This page merges both into a plain-language version you can follow with no coding background, with a copy-and-paste prompt attached to every section.
Six product changes matter to ordinary users.
Sketch canvas: draw a rough sketch inside ChatGPT and use it as a reference.
Templates: the examples OpenAI names are Poster and Merch, plus common formats like flyers and product photos. Pick a template, then add the message and the style you want.
Annotations on the image: mark notes straight onto the picture to make a local edit.
Faster: OpenAI says generation latency drops by up to 50% compared with Images 2.0.
Stable across turns: editing repeatedly inside a long conversation drifts less, and earlier edits are easier to keep.
Prompt sharing: when you share an image you can share the prompt with it, so someone else can swap in their own photo and details and run their own version.
The model itself differs in three places.
Lighting and materials look more natural.
The main subject in a reference photo keeps its likeness more reliably.
It changes only the spot you point at, leaving the detail next to it alone.
OpenAI also says it handles complex layouts better, transparent backgrounds included. Images 2.5 is live in ChatGPT, ChatGPT Work and Codex on every plan, on desktop, mobile and web.
One note about Sketch: the launch post says typing @Sketch in ChatGPT opens the canvas, but the feature looks like a batched rollout (OpenAI did not publish the order, that part is my guess), and my own account still didn't have it on 2026-09-10. If you don't see it either, draw the sketch on paper or in any drawing app, photograph or screenshot it and upload it as a reference image. The guide's "sketch to finished image" recipe is written for exactly this case, and the "sketch to finished image" section in part two below is the one to use.
Last updated: 2026-09-10
Part one: the six-slot prompt order
Six-slot order: purpose, composition, detail, people, text, exclusions. Write them in that order. (screenshots in Chinese)
The official guide's Prompting fundamentals section lists eight principles (define the result, choose a format that is easy to maintain, describe visible details, say what people are doing, give text verbatim, separate changes from constraints, assign a role to each reference, change one thing at a time). The "six slots" below are my rearrangement of the first few into an order that is easy to fill in with no coding background. It is not an official format.
Expand for the full walkthrough and pasteable prompts
Slot one: purpose
Say what the image is and what it is for. The guide asks you to define the result first, naming the subject and the use, for example a product photo, an ad, a diagram. When the content gets complex, break the prompt into sections, scene, subject, detail and constraints, with a heading on each.
Make a product photo. It will be the main listing image on Shopee. The subject is a 200ml bottle of handmade shampoo.
Slot two: composition and constraints
Place things first: the whole product in frame, empty space at the top. Spell out proportions and margins.
Next, say how the frame is arranged: composition, aspect ratio, and what has to sit where. This is the slot I drop most often, and the result is the subject cropped away, empty space in the wrong place, and text squeezed against the edge.
Square 1:1, product centred and slightly low, leave the top third empty for text, the whole product in frame with no crop on the cap.
Slot three: visible details
The guide calls this one Describe visible details: state the material, the lighting, the colour and the visual medium. For a real photo, say "photorealistic" or "real photograph" outright and describe the framing and the texture. Camera specs (a 50mm lens, say) are only a hint about the look, not a promise that the lens is physically simulated. For wide-angle, cinematic, low-light, rain or neon scenes, the guide suggests writing out scale, atmosphere and colour, because mood adjectives alone are not enough.
Photorealistic, a frosted glass bottle with a wooden cap, soft natural window light from the side, shallow depth of field, a warm off-white countertop behind it, natural colour with no heavy retouching.
Slot four: people and action
If there are people in the frame, say how much of the body is in shot, how big the person is next to the object, where the eyes are looking, and how the hands touch things. The guide's own examples are sentences like "full body in frame, including feet", "looking down at an open book", "both hands resting naturally on the handle".
On the right of the frame, a female shop assistant from the waist up, looking at the shampoo in her hand, right hand holding the bottle naturally, left hand supporting it from below, the person smaller in scale than the product.
Slot five: text in quotes, verbatim
Text gets its own lines: give the wording verbatim, write placement and font separately.
For words that appear in the image, the guide asks you to put them in quotes verbatim and describe placement and font separately. For brand names or unusual words, spell them out letter by letter when needed. End with a line saying no other text, then check the spelling and legibility yourself. For small type, dense information, or several fonts at once, the guide suggests running quality at medium and at high and comparing the two.
The image shows only this line, copied verbatim: "Anniversary Sale, Buy Two Get One Free"
Place it in the middle of the empty space at the top, bold sans-serif, even letter spacing, high contrast.
This line appears exactly once, clear and readable.
No other text, no watermarks, no unrelated logos.
Official example (the Render exact text section):
Render the tagline exactly once, clearly and legibly, integrated into the ad layout.
No extra text, no watermarks, no unrelated logos.
Slot six: exclusions
The last slot is what you do not want. The guide uses the same pattern across its examples: list the unwanted elements as a row of short phrases at the end. Common exclusions are watermarks, logos, extra text, checkerboard backgrounds, shadows, over-retouching, a stock-photo feel, gradients and ornament.
No watermarks, no logos, no extra text, no cartoon illustration style, no over-retouching, no stock-photo feel.
Part two: when editing, write what changes and what stays separately
This is the most valuable line in the whole guide for me. It is called Separate changes from constraints. When you edit, say "change only X" first, then list the details that must survive: the person's face, the geometry of an object, the layout, the lighting, the words on a label. Name the exclusions (extra text, logos, watermarks) in the same breath. For fine local retouching, also name the things that must not move, saturation, contrast, arrows, camera angle, surrounding objects.
Expand for the full walkthrough and pasteable prompts
Swap one thing, keep everything else
Change and keep, separated: only the chair changes, background and lighting are locked.
In this room photo, replace only the white chair with a wooden one.
Preserve the camera angle, the indoor lighting, the floor shadows and every surrounding object.
Keep everything else unchanged.
The wood needs real contact shadows and a real material texture.
Official example (the Change furniture in a room section):
In this room photo, replace ONLY the white chairs with chairs made of wood.
Preserve camera angle, room lighting, floor shadows, and surrounding objects.
Keep all other aspects of the image unchanged.
Photorealistic contact shadows and fabric texture.
Give each reference image a role
Give every reference image a role: image one is the setting, image two is the subject, and say where it goes.
When you drop in several reference images, the guide asks you to number each one and say what it is for: subject, style, clothing, background. Then say how the inputs combine and which element moves where. Without that step the model has to guess what each image is for.
Image 1 is the setting, image 2 is the dog.
Put the dog from image 2 into the setting of image 1, next to the woman.
Keep the lighting, composition and background from image 1.
Do not change anything else.
Official example (the Combine references section):
Place the dog from the second image into the setting of image 1, right next to the woman, use the same style of lighting, composition and background. Do not change anything else.
Change the clothes, keep the person
Portrait editing lives or dies on what you pin down. The guide's clothing example lists it in detail: face, features, skin tone, body shape, pose, identity, expression, hairstyle, proportions. Only the clothing changes, and the lighting, shadows and colour temperature have to match the original so it does not look pasted on.
Use the clothing reference image to change what the person in the photo is wearing.
Do not change her face, features, skin tone, body shape, pose or identity. Keep the exact same likeness, expression, hairstyle and proportions.
Change only the clothing. It should follow her current pose and body naturally, with realistic fabric drape.
Lighting, shadows and colour temperature must match the original photo so it looks genuinely worn.
Do not change the background, camera angle, framing or image quality, and do not add accessories, text, logos or watermarks.
Sketch to finished image
Sketch to finished image: keep the placement, the proportions and the perspective.
Turning a sketch you already have into a finished image comes down to two things: preserve the original layout, proportions and perspective, then fill in plausible materials, lighting and surroundings. The guide specifically asks you to add "do not add new elements or text" so the model does not improvise. With no Sketch canvas, draw on paper, photograph it, upload it, and this section still works.
Turn this hand-drawn sketch into a photorealistic image.
Preserve the original layout, proportions and perspective exactly.
Choose materials and lighting that make sense for what the sketch intends.
Do not add new elements or text.
Official example (the Turn a drawing into a realistic image section):
Turn this drawing into a photorealistic image.
Preserve the exact layout, proportions, and perspective.
Choose realistic materials and lighting consistent with the sketch intent.
Do not add new elements or text.
Transparent cutouts
After cutting out, check the edges on a light and a dark background
A cutout needs two things at once: the prompt has to ask for the subject in isolation, and on the API side you set background to transparent and output PNG or WebP. The guide has one very practical warning: a drawn checkerboard is not transparency. Real transparency means the file carries an alpha channel, so open the file you get and check hair, glass, shadows and object edges. On later edits, repeat "keep the transparent background". The ChatGPT web app has no parameter to set, so write the requirement into the prompt in full and check the edges yourself once you have the image.
Extract the product from the input image and place it on a fully transparent background.
Output: product centred, crisp edges, no halos or colour fringing.
Preserve the product's shape and the text on the label exactly.
Only light polishing. Do not add a solid backdrop, a checkerboard, scenery or shadows.
Do not restyle the product. Just remove the background and keep a clean transparent channel.
Official example (the Create a transparent product cutout section):
Extract the product from the input image and isolate it on a fully transparent background.
Output: centered product, crisp silhouette, no halos/fringing.
Preserve product geometry and label legibility exactly.
Add only light polishing. Do not add a solid backdrop, checkerboard, scenery, or shadow.
Do not restyle the product; remove the background and preserve clean alpha transparency.
Remove an object
Removal is the shortest pattern: name the thing to remove, then tell it to leave everything else alone. The guide's version is two sentences.
Remove the flower from the man's hand. Do not change anything else.
Official example (the Remove an object section):
Remove the flower from man's hand. Do not change anything else.
Put a person into a new scene
For compositing a person into a new scene, the guide asks for four things to be spelled out: natural lighting, believable detail, how much of the body is in frame and where they look, and how the person interacts with the scene. Then name the features and proportions that must not change. If you want it to look like a casual snapshot, say no cinematic lighting, no dramatic grading, no stylised composition.
Put this person into a realistic scene: she is browsing at a night market stall.
It should look like a snapshot a passer-by took, not the over-polished look of a movie poster.
She is slightly left of centre, from the waist up, looking at the goods on the stall, right hand picking one up.
Do not change her features, face shape or body proportions.
It is early evening, natural light and true colour, with plausible detail on the stall.
No cinematic lighting, no dramatic grading, no stylised composition.
Part three: how to revise over several turns without making it worse
The guide's Iterate deliberately line is about rhythm: feed the last output back in as the next input, ask for one change at a time, and repeat the details you want kept. "Same style as before" carries some context, but as soon as the result starts drifting, restate the constraints that matter. Compare results before adding new instructions, do not stack them up.
Expand for the full walkthrough and pasteable prompts
Step one: get the starting image right
The starting image has to nail the text and the subject in one go. The guide's example wraps the billboard copy in "EXACT, verbatim, no extra characters" and spells out font, contrast, centring and letter spacing, then asks for the text to appear once.
Make a photorealistic highway billboard scene with this shampoo, at sunset.
Text on the billboard (verbatim, no extra characters):
"Fresh and Clean"
Font: bold sans-serif, high contrast, centred, even letter spacing.
Make sure this line appears exactly once and is completely clear and readable.
No watermarks, no logos.
Step two: change one condition at a time
One change per round: check, then continue.
Send the image from the last step back in, then name only the one thing that changes. The guide's demo for this step is a single sentence.
Make it a winter evening with snowfall.
Official example (the Change one condition section):
Make it look like a winter evening with snowfall.
Step three: restate what must be kept, every round
The editing section of the guide has an important warning: repeated edits can still wipe out the detail you meant to keep, so restate those constraints and check every round. If an area must not move by a single pixel, the more reliable route is compositing the approved edit back onto the original, not relying on the prompt alone.
Character consistency: build a character sheet first
Lock the character sheet first, then swap scenes, same face, same proportions, same outfit
For a picture book, an illustration series or serialised content, the guide's approach is to generate a character reference first, fixing the look, proportions, clothing and temperament, then repeat those defining details on every new scene and change only the setting and the story.
Continue this story with the same character.
Scene:
The same little forest hero, gently freeing a startled squirrel from a fallen tree, crouching beside it to calm it down.
Character consistency:
The same green hooded tunic
The same features, proportions and colour palette
The same gentle, brave temperament
Style:
Children's book watercolour, soft light, a forest after snow, a warm and reassuring mood.
Constraints:
Do not redesign the character
No text
No watermarks
Part four: six examples you can use in Taiwan
Start from a template: pick the template, then add the product, the background and the empty space. All six examples follow that order.Expand for the full walkthrough and pasteable prompts
One: product photo
For products, OpenAI focuses on material, packaging and print legibility, and asks for original designs that do not infringe.
Make a product photo: a 150g pack of hand-drip coffee bags, kraft paper packaging with a dark green label.
High-end product photography, real paper and print texture, studio lighting, shallow depth of field, sharp and readable text on the label.
Square composition, product centred, clean light grey background.
The packaging shows only this line (verbatim): "Natural Yirgacheffe"
No watermarks, no logos, no extra text.
Two: flyers and text posters
For text-heavy images, the guide's approach is to write the prompt like a spec: state the message, the design elements and the style, put the words in quotes verbatim, and say outright that the layout has to be readable.
Make an A4 portrait event flyer for a neighbourhood craft market.
Layout: headline at the top, an illustration of the market in the middle, three lines of information at the bottom, generous white space, clean and readable typesetting.
Copy the following text verbatim, with nothing added:
Headline "Zhongxing Street Craft Market"
Information "October 18, 2026 (Sat)" "2 PM to 8 PM" "The plaza next to 12 Zhongxing Street"
Style: warm hand-drawn illustration, beige background with dark green and brick red, clear sans-serif type.
Avoid type too small to read, unnecessary ornament, and a stock-photo feel.
No watermarks, no other text.
Three: logo
The guide's logo example stresses three things: simple shapes, legible at every size, delivered on a transparent background. It also puts "checkerboard" straight into the exclusions, because the guide makes a point of it: a drawn checkerboard is not real transparency.
Design an original, non-infringing logo for a neighbourhood bakery called "Wheat and Field".
It should feel warm, simple and long-lasting. Use clean vector-style shapes, a clear outline and balanced white space.
Favour simplicity over detail so it reads clearly at small and large sizes. Flat design, minimal strokes, no gradients unless essential.
Fully transparent background. A single logo, centred, with generous padding and clean edges.
No solid backdrop, no scenery, no checkerboard, no watermark.
Official example (the Design a reusable logo section):
Favor simplicity over detail so it reads clearly at small and large sizes. Flat design, minimal strokes, no gradients unless essential.
Fully transparent background. Deliver a single centered logo with generous padding, clean alpha edges, and no solid backdrop, scenery, checkerboard, or watermark.
Four: process diagram
For flowcharts and infographics the guide adds one more warning: beyond the look, verify the labels and the factual relationships yourself. A wrong name or a reversed arrow makes a beautiful diagram unusable.
Make an infographic explaining a process: "an online order from checkout to shipping".
Show, in order: customer places the order, system takes payment, warehouse picks, packing, handover to the courier, customer receives. Connect each step with arrows.
Every step needs a clear label and a simple icon.
Make it look like a class handout or a slide: white background, simple icons, clear labels, readable type.
Avoid type that is too small, unnecessary ornament, or anything that makes the diagram hard to follow.
Five: four-panel comic
Turning a story into a comic comes down to splitting it into one visual beat per panel, described concretely and driven by action, so the model can lay out panels that read with rhythm.
Make a portrait four-panel comic.
Panel one: the owner pulls down the shutter on her way out, and the shop cat peers out from behind the counter, eyes wide.
Panel two: the shutter closes and the shop goes quiet. The cat slowly turns to look at the empty shop, its eyes starting to gleam.
Panel three: the shop is a mess. The cat is sprawled next to the till like it owns the place, torn snack packets scattered around.
Panel four: the shutter opens. The cat sits neatly by the door, looking up at the owner as if nothing had happened.
Six: slide graphic
The guide suggests a landscape size for slides, and writing the numbers and labels straight into the prompt instead of letting the model invent them. When the image has small type, a legend, axes or a footnote, set quality to high. The guide adds its own reminder: the market figures in its example are fictional design inputs, and you should replace them with verified data before real use.
Make a landscape slide titled "Three Years of Revenue Growth".
White background, modern sans-serif, a clean and spare layout. The page contains:
A bar chart showing revenue growth from 2023 to 2026, trending up
The figure above each bar: 3,200,000 in 2023, 5,800,000 in 2024, 9,100,000 in 2025, 14,000,000 in 2026
A small footnote in the bottom left: "Source: internal company figures"
Space reserved in the bottom right for a logo
Text must be very readable, the data hierarchy clear, the spacing tidy, so it looks like a slide someone would actually take into a meeting.
Avoid clip art, stock photos, gradients, shadows and unnecessary ornament.
The model draws whatever numbers you give it in slides and diagrams. The guide asks you to supply real text and real data instead of letting it fill in the gaps. So type the numbers yourself, then check them one by one on the image you get.
Part five: the checklist for after the image comes out
Check the result item by item: text, relationships, things kept, edges.
The guide ends with a section called Check the result, four questions. Run through them after every generation, before you use the image:
Expand for the full walkthrough and pasteable prompts
One: text
Are the words that should appear correct and legible? Are the labels on a diagram, and the relationships between them, right?
Two: detail preserved
Did the person's face, the product's shape, the labels and the details from the reference images all survive intact?
Three: scope
Did this edit change only the thing you asked for, and did anything else get touched?
Four: transparent background
If you need transparency, does the file really carry an alpha channel, or is the background just painted on?
The guide adds one more line: when you change prompts or change models, run representative inputs and compare quality, response time and cost, all three together.
Part six: the developer section
This part is for people calling the API. If you do not use it, skip straight to the FAQ.
Expand for the full walkthrough and pasteable prompts
How the two models differ
Model
Official positioning
gpt-image-2.5-flare
The small model, optimised for speed. The guide says quality is comparable to GPT Image 2, while the launch post says it is above GPT‑Image‑2. The two sources word it differently, so both are listed here. The launch post says it carries the same improvements in quality, editing and speed, with 50% lower latency than GPT-Image-2. Good for creator content, social assets, product experiences, visual search, fast prototyping and high-volume generation.
gpt-image-2.5-sunburst
The base model, optimised for quality. Image quality is above GPT Image 2 and generation takes longer. Good for high-end visual pipelines that need precise control across many edits, such as marketing assets that go live as they are, or finely retouched product images.
How to choose: on a new pipeline, try Flare first when speed matters and Sunburst first when quality is strict. Once Sunburst passes, test Flare with the same prompts and inputs, and switch if it also passes and is faster. Keep Sunburst only when the quality gap genuinely calls for it. One reminder from the guide: measure speed on your own workload, because a figure from someone else's workload does not transfer.
Parameters
Parameter
Accepted values
model
gpt-image-2.5-flare or gpt-image-2.5-sunburst
quality
auto (default), low, medium, high, xhigh, max
size
auto or a custom resolution. Common ones: 1024x1024, 1536x1024, 1024x1536, 2048x2048, 2048x1152, 3840x2160, 2160x3840
background
auto, opaque, transparent
Custom resolutions have four limits: no side above 3,840 pixels, both sides multiples of 16, the long side no more than 3:1 against the short side, and total pixels between 655,360 and 8,294,400. Output above 3,686,400 total pixels (the equivalent of 2560x1440) is experimental.
The way to use quality is to pick the model first and tune quality after. On a first comparison, hold the quality setting that both models support fixed, and hold the prompt, the reference images and the output size fixed too. The same quality label on a different model does not mean the same picture quality or the same response time. Step quality up when it falls short, then step back down once it passes, to see whether you can get faster while quality stays acceptable. Use xhigh and max only when they genuinely fix a quality requirement that is failing and still fit your time budget.
For transparent assets, set background to transparent explicitly, use PNG or WebP, and check the decoded alpha channel once you have the file, especially hair, glass, shadows and object edges. output_compression applies to JPEG and WebP only, not PNG.
Three things that matter when migrating an old pipeline
The guide's migration flow has six steps. Here they are boiled down to three points.
Save a baseline first: collect representative production prompts and reference images, including hard edits, exact text, faces, product geometry and transparent assets, and record the model, the request settings and the results you use today.
Change only the model in the first comparison: hold the prompt, reference images, size and output format fixed. Compare instruction following, preservation of identity and product, text accuracy, whether anything got changed that should not have been, and transparency. Repeat each request several times to see how stable it is, and test editing pipelines as a whole chain, not one step at a time.
Change one setting at a time: compare quality levels before touching the prompt. Measure typical response time, slow responses, failures and retries, plus the cost per usable image. The guide specifically warns against assuming a faster model is cheaper, so check the current pricing. When you go live, send a small slice of traffic first, watch the same metrics, then scale up, and keep the old model as a fallback while it is still supported.
FAQ
Expand for the full walkthrough and pasteable prompts
I can't find Sketch in my ChatGPT
The launch post says typing @Sketch in ChatGPT opens the canvas, but it looks like a batched rollout, and my own account still didn't have it on 2026-09-10. If you don't see it, take the other route: draw the sketch on paper or in any drawing app, photograph or screenshot it, upload it as a reference image, then paste the "sketch to finished image" prompt above. It still keeps your original layout and perspective.
The text in my image keeps coming out wrong
Do the four moves from slot five: put the words in quotes verbatim, say how many times they appear and where, describe font and placement separately, and end with a line saying no other text. For small type, dense information or several fonts, move quality from medium to high and compare again. For brand names or unusual words, spell them out letter by letter when needed. Always check spelling and legibility yourself once the image is out.
The ChatGPT web app does not have those parameters. What now?
quality, size and background are API parameters, and I did not find matching settings in the web app. What you can do there is write the requirement into the prompt: for a transparent background, spell out "fully transparent background, no solid backdrop, no checkerboard", and for landscape just say landscape composition, then check the image yourself. The guide is written for the API. The prompt patterns work fine for me in the web app, and the parameter layer is the part the web app does not have.
How much faster is it really
The launch post says generation latency drops by up to 50% compared with Images 2.0. The API docs say GPT-Image-2.5 Flare has 50% lower latency than GPT-Image-2. The guide also warns that the real figure has to be measured on your own prompts, reference images, output sizes and quality settings, because numbers from another workload do not transfer.
The combined version: one paragraph to paste
Join it into one paragraph: fill the six slots, join them into one prompt, save it and reuse it.
If you do not want to assemble it slot by slot, use the block below as a template. Replace the text in brackets with your own and paste it into ChatGPT. The first six lines are the six-slot order for generation, the four lines after that are the change-versus-keep split for editing.
Purpose: make a [product photo / flyer / process diagram], to be used for [where].
Composition: [aspect ratio], subject placed [where], empty space at [where] for text, no crop on [what must not be cut].
Detail: [photorealistic / illustration / vector] look, material is [material], lighting is [lighting], colours [colours], [depth of field or texture].
People: [who] is [where] in the frame, [how much of the body], looking at [where], hands [doing what], scale [relative to what].
Text: the image shows only these words, verbatim: "[the words]". Placed [where], font [font], appearing once, clear and readable. No other text.
Exclusions: no watermarks, no logos, no extra text, no [the style you do not want].
If you are editing an existing image, add this on the end instead:
Change only [the one thing to change].
Must stay exactly as is: [the person's face / the product's shape / the layout / the lighting / the text on the label / the camera angle].
Do not touch anything else.
Do not add new elements, text, logos or watermarks.
Reading the whole guide, the two lines that actually help most if you have no coding background are these: put text in quotes verbatim, and when editing, write "what changes" and "what stays" on separate lines. Start with those two and you will redo images noticeably less often (that is my experience, not an official figure). Come back for the rest when you actually get stuck.