GPT Image 2.5 prompt — combine two references into one scene
Number the images, say which element moves, say where it lands, then close the door.


Also attached
| Mode | Model | Quality | Size | Ratio | Credits |
|---|---|---|---|---|---|
| Image to Image | Think | Standard | 1K | 2:3 | 10 cr |
Published by OpenAI as an example for this model — not generated on this site. Model: gpt-image-2.5-sunburst. Source: OpenAI The published example ran at 1024x1536, a legal size we do not currently offer; the frame here is the same shape at our nearest size.
A dog from the second photograph, placed next to a woman in the first. Twenty-six words, and they contain all four things a composite prompt needs: which image is which, what is being taken, where it goes, and an instruction that nothing else may move. Numbering the inputs is the part most people skip, and it is the part that makes the sentence unambiguous.
The prompt, nothing cut
Place the dog from the second image into the setting of image 1, right next to the woman, use the same style of lighting, composition and background. Do not change anything else.
This exact prompt, at these exact settings
The mode and all five settings are locked to this template, so there is nothing to line up by hand. Press Run this prompt above and the panel opens filled in — prompt, mode, quality, size and frame — and your first image needs no account. If your plan sits below what this ran at, the line under the panel says so.
This example was published at Think · Standard · 1K. A free run is Fast · Standard · 1K, so expect less fine detail and softer small text. The composition is the same; the finish is not.
What is safe to change
Most libraries publish this list and stop, which is why so many copied prompts come back worse than the original: the swappable nouns are the safe half. The three sections under it are the other half — the clauses that are doing the work, the ones that break it if you touch them, and why it is written in this order.
The image numbers
"the second image" and "image 1" — each input identified by position. With two references you can sometimes get away without it; with three you cannot, and the habit is worth forming on two.
The placement phrase
"right next to the woman" is the destination. Vague placement is the most common reason a composite comes back with the element floating, oversized, or behind something. Say where.
The match clause
"use the same style of lighting, composition and background" is the integration instruction in compressed form. It is the same job the clothing template does at length.
Three ways to break it
Each one is a real failure with a reason attached. A rule without a reason is not usable.
01Leaving the inputs unnumbered
With two images the model may guess right. With three it will not, and the failure is silent — you get a coherent picture assembled from the wrong roles, which takes a moment to even notice.
02Not saying where it goes
Scale and occlusion are decided by placement. "Put the dog in this scene" leaves both open, and the usual result is a dog that is the wrong size for where it ended up standing.
03Combining subjects with incompatible light
A subject shot in hard midday sun composited into an overcast scene needs its shadows rebuilt, not matched. The model will try, and the seam is usually visible at the contact points — feet, ground shadow, rim light.
Inside the Put something from one photo into another prompt
Twenty-six words that contain a complete composite brief, and the reason it is complete is that it answers four separate questions rather than three.
Which images are which: "the second image", "image 1". What moves: the dog. Where it lands: right next to the woman. What must not change: anything else. Drop any one of the four and you get a characteristic failure — a merged scene if the roles are unclear, a floating subject if the placement is missing, a redrawn background if the prohibition is missing.
Numbering is the habit worth building even when it feels unnecessary. OpenAI's fundamentals put it as identifying each input by number and purpose, and the reason it matters more than it seems is that references are positional. With two images, a model has a fair chance of inferring which is the scene and which is the subject from the grammar of your sentence. With three or four it is guessing, and a wrong guess does not produce an error — it produces a confident, coherent image built from the wrong assignment, which is much harder to spot than a broken one.
The placement phrase deserves more respect than it usually gets. "Right next to the woman" is doing geometry: it fixes the dog's distance from camera, which fixes its scale, which fixes what it occludes and what occludes it. Without it the model chooses, and its choice is frequently a subject that is subtly the wrong size — the single most common tell in an AI composite.
The honest limitation is lighting, and it is a physical one rather than a prompting one. If your two sources were lit differently, matching them means rebuilding shadows that were never photographed. The model does a reasonable job and the seam usually shows where the subject meets the ground. Pick sources with compatible light and this prompt is nearly free; pick incompatible ones and no amount of rewording fixes it.
Put something from one photo into another — common questions
- Do I have to number the images?
- With two you can sometimes get away with it. With three or more you cannot, and the failure is a coherent image built from the wrong roles — much harder to spot than an obvious break.
- Why is my composited subject the wrong size?
- Because the placement was not specified. Distance from camera determines scale, so naming the destination — "right next to the woman" — fixes the size for free.
- How many references can I combine?
- Up to sixteen at the endpoint, but the practical limit is lower. Past three or four elements the model starts blending them rather than placing them.
- Why is there a visible seam at the ground?
- Mismatched source lighting. Contact shadows have to be invented rather than matched. Choose sources with similar light direction and quality.
Written and maintained by Andy SwiftPublished Last updated






