GPT Image 2.5 prompt — insert a person into a new scene and keep their face
The scene is invented from scratch. The face is not. Say which is which.


| Mode | Model | Quality | Size | Ratio | Credits |
|---|---|---|---|---|---|
| Image to Image | Think | Standard | 1K | 2:3 | 10 cr |
Published by OpenAI as an example for this model — not generated on this site. Model: gpt-image-2.5-sunburst. Source: OpenAI The published example ran at 1024x1536, a legal size we do not currently offer; the frame here is the same shape at our nearest size.
A woman photographed in a museum, relocated to a campsite in Yosemite at dusk with a bear behind her. Everything about the image is new except the person, and the prompt spends most of its length on two things: making the new scene specific enough to be believable, and repeatedly refusing the cinematic treatment a model reaches for when you ask for drama.
The prompt, nothing cut
Generate a highly realistic action scene where this person is running away from a large, realistic brown bear attacking a campsite. The image should look like a real photograph someone could have taken, not an overly enhanced or cinematic movie-poster image. She is centered in the image but looking away from the camera, wearing outdoorsy camping attire, with dirt on her face and tears in her clothing. She is clearly afraid but focused on escaping, running away from the bear as it destroys the campsite behind her. The campsite is in Yosemite National Park, with believable natural details. The time of day is dusk, with natural lighting and realistic colors. Everything should feel grounded, authentic, and unstyled, as if captured in a real moment. Avoid cinematic lighting, dramatic color grading, or stylized composition.
This exact prompt, at these exact settings
The mode and all five settings are locked to this template, so there is nothing to line up by hand. Press Run this prompt above and the panel opens filled in — prompt, mode, quality, size and frame — and your first image needs no account. If your plan sits below what this ran at, the line under the panel says so.
This example was published at Think · Standard · 1K. A free run is Fast · Standard · 1K, so expect less fine detail and softer small text. The composition is the same; the finish is not.
What is safe to change
Most libraries publish this list and stop, which is why so many copied prompts come back worse than the original: the swappable nouns are the safe half. The three sections under it are the other half — the clauses that are doing the work, the ones that break it if you touch them, and why it is written in this order.
The anti-cinematic refrain
The prompt refuses movie-poster treatment three times, in three different wordings. That is not redundancy — "dramatic action scene" is a strong pull toward film stills, and one refusal does not hold against it.
The body direction
"Centered but looking away from the camera", "clearly afraid but focused on escaping". Framing, gaze and intent, stated separately. OpenAI's fundamentals recommend exactly this for people.
The named location
"Yosemite National Park" pulls in a specific geology, tree line and light. Any real place name does the same work as the historical-date trick, and for the same reason.
Three ways to break it
Each one is a real failure with a reason attached. A rule without a reason is not usable.
01Asking for drama without refusing the grade
An action scene with a bear is a film premise. Without explicit refusals you get a movie poster — teal shadows, orange highlights, impossible rim light — because that is what the training data pairs with the premise.
02Leaving out the wear details
"Dirt on her face and tears in her clothing" is what makes the scene read as mid-event. A clean subject in a destroyed campsite looks composited, because she has not been through what the background says she has been through.
03Expecting a portrait-grade likeness
This edit relocates a face into new lighting at a new angle in a new environment. It stays recognisable; it is not identical. The more the pose and light differ from the source, the more it moves.
Inside the Put a person into a completely different scene prompt
This is the most ambitious edit in the guide. Every pixel except the face is invented, and the face has to survive being relit, reposed and re-angled. Two things make it work, and neither is what people expect.
The first is how hard the prompt fights its own premise. "A large, realistic brown bear attacking a campsite" is a film. That premise sits in the training data next to movie stills, posters and trailer frames, all of them heavily graded. So the prompt refuses the grade three separate times, in three different wordings — "not an overly enhanced or cinematic movie-poster image", "grounded, authentic, and unstyled, as if captured in a real moment", "avoid cinematic lighting, dramatic color grading, or stylized composition". One refusal would not hold. Three, spread through the prompt so they are read alongside each part of the description, do.
That is a technique worth naming: when your subject matter has a strong genre attached, the refusal has to be repeated at roughly the same strength as the pull. A single "not cinematic" at the end of a paragraph about a bear attack loses.
The second thing is the physical continuity detail. "Dirt on her face and tears in her clothing." Six words, and they are what make the person belong in the scene. A subject inserted into a destroyed campsite while looking as she did in a museum reads as a composite, not because the edges are wrong but because she has not experienced the event the background depicts. Wear is continuity. Any insertion into a scene with weather, action or mess needs its own version of this line.
The body direction follows OpenAI's recommendation for people directly — describe framing, relative scale, gaze and interaction. "Centered in the image but looking away from the camera" fixes composition and gaze in one clause; "clearly afraid but focused on escaping" fixes the expression by giving it an intention rather than an emotion label, which is the same trick the comic template uses.
The honest limit: recognisable is not identical. This is the largest transformation in the library and the likeness moves the most. It is a good enough for a concept, a story beat or a personal image, and not for anything where the person has to be provably themselves.
Put a person into a completely different scene — common questions
- Why does my dramatic scene look like a movie poster?
- Because the premise is a film premise, and the genre pull is strong. Refuse the grade explicitly and more than once — the official prompt does it three times.
- How well is the face preserved?
- Recognisable, not identical. This is the biggest transformation here — new light, new angle, new environment — and the likeness moves proportionally.
- Why do I need the dirt and torn clothing?
- Continuity. A clean subject in a wrecked campsite reads as pasted in, because she has not been through what the background says happened.
- Does naming a real park help?
- Yes. It pulls in a specific geology, tree line and quality of light, the same way naming an exact date pulls in a period.
Written and maintained by Andy SwiftPublished Last updated





