GPT Image 2.5 vs GPT Image 2: What Actually Changed
OpenAI kept all three of GPT Image 2's price points and slid two new tiers in between them. The cheapest tier you'd actually ship went from 5.3¢ an image to 1.3¢ — and that, not picture quality, is the generational change.

We sell access to GPT Image 2.5, so the bias is declared. What keeps this honest: every figure below is either published by OpenAI or measured here, and the two are labelled apart. The places 2.5 is no better than GPT Image 2 are named before the places it is. And three rows further down say unconfirmed, because we have not run the test yet and would rather show you the gap than fill it with a guess.
Most write-ups about this release argue about sharpness, which is subjective and moves with the prompt. The thing that actually decides whether you migrate is in OpenAI's own pricing table, and it is not a quality claim at all: all three of GPT Image 2's price points survived into 2.5, to the tenth of a cent — and two new tiers appeared between them.
The cheapest tier you'd actually ship went from 5.3¢ an image to 1.3¢.
GPT Image 2.5 vs GPT Image 2: the release in two columns
| Changed | Did not change | |
|---|---|---|
| Models | Two — Flare and Sunburst | Token rates: $5 / $8 / $30 per million |
| Quality tiers | Two new ones, and a 4× cheaper usable tier | The checkerboard artifact |
| Editing | Precise editing and subject preservation, on both models | Text placement still slips |
| Latency | Flare: up to 50% lower | Element placement in structured layouts |
| Your side | One string in the request body | Your prompts — keep them unchanged for the first comparison |
The right-hand column is the half nobody publishes, and for anyone with GPT Image 2 already in production it is the more valuable one. Most of what a migration costs is not rewriting the call; it is discovering, one job at a time, which of your assumptions survived.
The upgrade nobody wrote about: the price ladder
Put the two generations' published per-image rates side by side and something odd shows up. Read the two columns across and the story is nothing changed — five prices that barely moved. Read them as a ladder and the story is different.
- GPT Image 2:
low
$0.006
GPT Image 2.5:low
$0.0059
- GPT Image 2:no tier at this priceGPT Image 2.5:
medium
New$0.0132
The tier OpenAI reaches for in almost every example in its own prompting guide.
- GPT Image 2:
medium
$0.053
GPT Image 2.5:high
$0.0527
- GPT Image 2:no tier at this priceGPT Image 2.5:
xhigh
New$0.0937
Above what everyday work used to cost, well under the old top rate.
- GPT Image 2:
high
$0.211
GPT Image 2.5:max
$0.2107
On GPT Image 2, medium at 5.3¢ was the first tier most people considered shippable — low was for thumbnails and little else. On 2.5 the tier OpenAI recommends across nearly every example in its prompting guide is medium, and that one costs 1.3¢. Same job, same place in the range, a quarter of the price. That is the largest thing in this release for anybody paying per image, and no launch coverage mentions it, because it only exists when you read the table as a ladder.
Price alignment is not quality alignment
One caution before you budget against this, and OpenAI is explicit about it: "The same quality label does not imply the same image quality or response time across models." 2.5's high costs what 2's medium cost. Whether it looks like 2's medium is a question only your own prompts can answer.
Token rates themselves did not move at all — see measured API cost per image for the full matrix. What changed is the number of rungs you can stand on.
Only Sunburst is rated higher. Flare is rated the same.
The launch announcement says "sharper details", and most coverage flattened that into 2.5 is better than 2. The developer documentation is more careful, and the distinction it draws is the one that decides which model you should be testing:
GPT Image 2.5 Flare is the small model, optimized for speed, with image quality comparable to GPT Image 2. GPT Image 2.5 Sunburst is the base model, optimized for quality, with higher image quality than GPT Image 2. Both models offer improvements in precise editing and subject preservation.OpenAI — image prompting guide
| vs GPT Image 2 | |
|---|---|
| Flare | Same quality, roughly half the latency |
| Sunburst | Higher quality ceiling, longer generation |
| Both | Better at precise editing and subject preservation |
That last row is worth sitting with. The thing both models genuinely improved on is neither speed nor resolution — it is not changing what you didn't ask them to change, which is the complaint every image editor in this class collects and the one no benchmark reports. If your work is multi-turn editing rather than one-shot generation, that row is the release.
Which of the two to run is a separate question with a separate answer: Flare or Sunburst.
What's genuinely new, and three things we haven't confirmed
| Status | |
|---|---|
| Two models: Flare and Sunburst | ✅ Confirmed |
Two new quality tiers: xhigh and max | ✅ Confirmed |
| Sketch, Templates, comment-based edits, shareable prompts in ChatGPT | ✅ Confirmed |
| Up to 50% lower latency on Flare | ✅ Confirmed — stated as a maximum, not per request |
| Transparent backgrounds | ⚠️ Unconfirmed as new. The migration docs tell GPT Image 2 users to keep their existing transparency requirements when comparing, which reads like 2 had them too |
| Native 4K | ⚠️ Unconfirmed as new. Third-party coverage credits GPT Image 2 with 2K output and multiple ratios; we have not pinned down where 4K starts |
| Reference image limit | ⚠️ Unconfirmed. GPT Image 2 is widely reported at up to 16 per edit; no limit is documented for 2.5 |
OpenAI model and announcement pages, read 9 September 2026
Every round-up we have read lists transparent backgrounds as a 2.5 feature. It might be one. Until we have run background=transparent against GPT Image 2 ourselves, the row says so — a page whose whole argument is check this yourself does not get to fill its own gaps with the reading that suits it.
What GPT Image 2.5 kept from GPT Image 2
Half the value of a migration guide is knowing what you don't have to retest. Four things carried over unchanged, and one of them is the reason some jobs should not move yet.
Token rates. Identical.
$5 per million text input, $8 image input, $30 image output. OpenAI's own wording is "Token rates match GPT Image 2." Your cost model does not need rebuilding — only your tier choices do.
The checkerboard artifact. This is the one that matters.
GPT Image 2.5 still produces a faint grid across clouds, grass, hair and behind dense text. Nine different people reported it across four subreddits within a day of launch, and one went looking specifically:
Have they fixed the checkerboard artefacts? Edit: No they have not. It's still really noticable in clouds and plants. The more it's edited the more obvious it becomes — once you notice it you'll see it everywhere.Two reports from the launch-day threads
gpt-image-2.5-sunburst · medium · 1536×1024If that artifact is what currently keeps you off GPT Image 2 for a particular job, 2.5 will not fix it, and a single-step test will not show it to you — it accumulates across an edit chain. Test your own soft textures before you move that work.
Text placement, and elements in structured layouts
Both are still listed as limitations in OpenAI's own documentation. Text accuracy improved; knowing exactly where to put it did not. If your job is a poster on a strict grid, budget the same number of retries you budget today.
gpt-image-2.5-sunburst · medium · 1536×1024Your prompts
Don't rewrite them. OpenAI's migration steps are explicit about keeping the prompt, references, dimensions and output format unchanged for the first comparison. You cannot tell what the model changed if you changed things too.
What the pre-launch write-ups got right, and what they got wrong
For about three weeks before launch, most of what you could find about GPT Image 2.5 was extrapolated from two anonymous Arena checkpoints. Several of those pages still rank for this comparison, and some of them still say the model has not been announced. Here is how the speculation held up.
| Pre-launch claim | Verdict |
|---|---|
| Two models codenamed Flare and Sunburst | ✅ Correct — shipped under exactly those names |
| Reduced noise after repeated edits | ❌ Not what users report. The checkerboard artifact is still there and gets more visible with each edit |
| Better typography at small sizes | ⚠️ Partly. OpenAI recommends high for small text and still lists placement as a limitation |
| Higher quality across the board | ❌ Only Sunburst. Flare is explicitly "comparable to GPT Image 2" |
| A newer knowledge cutoff | ⚠️ Unconfirmed. Real-world detail is said to be more accurate; no date published |
| "No official announcement yet" | ❌ Obsolete. Announced 8 September 2026 |
| Up to 16 reference images | ⚠️ Unconfirmed for 2.5. That is GPT Image 2's documented figure |
One of those pre-launch pages ended with genuinely good advice: build your evaluation set now, save the prompts and their current outputs, "and when the new model arrives you will have a same-prompt comparison on day one instead of relying on other people's Arena screenshots." That is what this page is trying to be.
Migrating from GPT Image 2 to GPT Image 2.5, in five steps
Still on gpt-image-1.5?
That model is marked deprecated with a scheduled shutdown date. For you this is not an optional upgrade — find the shutdown date on OpenAI's model page before you plan anything else on this list.
The procedure below is OpenAI's, out of the prompting guide, and it is the most useful thing in there that nobody has reprinted. What we can add sits beside each step rather than replacing it.
Step 1Save a baseline
Collect representative production prompts and reference images — difficult edits, exact text, faces, product geometry, transparent assets. Record the current model, request parameters and results.
What we'd addTake that list literally and use it as your test set. It is the closest thing there is to an official statement of what this model class is expected to find hard, which makes it a better baseline than anything you would assemble under time pressure.
Step 2Choose the first candidate
If GPT Image 2 already meets your bar, start with Flare and test for latency. If it falls short, start with Sunburst. Keep the prompt, references, dimensions and output format unchanged for the first comparison.
What we'd addThe unchanged clause is the one that gets skipped. Move the prompt and the model in the same run and you have two variables and one result — whatever you see, you cannot attribute it to either.
Step 3Check the complete result, not the first frame
Instruction following, identity and product preservation, text accuracy, unwanted changes, transparency. Repeat requests to measure consistency. For editing workflows, test the complete sequence of edits rather than single steps.
What we'd addA single edit comes back clean on both generations. The checkerboard artifact is the thing that accumulates across a chain, so a one-step test is exactly the test that will not show it to you.
Step 4Only then measure latency
Test for a latency gain after quality passes.
What we'd addFlare is documented at up to 50% lower latency — a maximum, not a per-request figure. Measure your own typical and slow responses; a published ceiling is not a number you can put in a budget.
Step 5Tune one setting at a time
Compare quality levels before rewriting the prompt. Measure typical and slow responses, failures, retries, and cost per accepted image. "Confirm current pricing rather than assuming the faster model costs less."
What we'd addThat closing line has an answer here, and it is not the one it warns against: Flare and Sunburst cost exactly the same per token. The faster model is not the cheaper one. What you are trading is latency, not money.
In your code, all of that comes down to one string:
quality gained two tiers above high, so every value you already send is still valid, and the rest of the request body stays exactly as it is. The parts that will actually cost you time are organisation verification, worst-case latency, and rate-limit errors that arrive inside an HTTP 200 — none of which are new in 2.5.
So should you switch from GPT Image 2 to GPT Image 2.5?
"Is it better" has six different answers depending on which tier you are on today and what the output is for.
| If you're… | Then |
|---|---|
Running gpt-image-2 at medium | ⭐ Switch, for the money. 2.5's high costs what you pay now, and there is a tier below it at a quarter of the price |
Running gpt-image-2 at high for finished work | Go to Sunburst and test the ceiling |
| Mostly doing multi-turn editing | Switch. Precise editing and subject preservation are the real shared upgrade |
| Doing large soft textures — sky, foliage, hair | ⚠️ Test the checkerboard first. 2.5 did not fix it |
Still on gpt-image-1.5 | ⭐ You have to. There is a shutdown date |
| Only using ChatGPT, never the API | You are already on 2.5 — it rolled out to every tier, free included |
Nothing on that table needs a decision today except the last two rows. gpt-image-2 has no announced shutdown date, so a migration you start this week and finish next month costs you nothing but the testing.
Where each one sits on the Image Arena
Both 2.5 models rank above GPT Image 2 on the public Image Arena. This is the one claim on the page you cannot check from your own account, and the size of the move is the part worth reading rather than the ordering.
| Image Arena | |
|---|---|
| GPT Image 2.5 Sunburst | First |
| GPT Image 2.5 Flare | Second |
| Elo, previous release → this one | 1,381 → 1,421 |
Arena standings move · figures read 9 September 2026
Going from 1,381 to 1,421 isn't 'destroying' the benchmark they previously set — it's an incremental gain.A commenter on the announcement thread
Which squares with everything above it. The generational change here is in the price ladder and in precise editing, not in a leaderboard position.
How this page was checked
Three of the pages currently ranking for this comparison still say GPT Image 2.5 has not been announced. That is not a writing-quality problem to out-write; it is a sourcing problem. So here is the sourcing.
- Prices come off the API, not off a blog. Every figure in the ladder is the published per-image rate for a 1024×1024 request at that tier, read for both generations on the same day.
- Quotes are verbatim. Where OpenAI's own wording decides something — Flare being "comparable to" GPT Image 2 rather than better than it, or a quality label not meaning the same thing across models — the sentence is reproduced word for word. A paraphrase is where a comparison quietly becomes an argument.
- Community reports are counted, not sampled. The checkerboard finding is nine separate people across four subreddits inside a day of launch. One report is an anecdote and nine is a pattern; the difference is the only reason that claim is here.
- What we haven't checked stays marked. Three rows above say unconfirmed, and each says why rather than leaving a blank.
One thing is still owed: a same-prompt gallery running both generations through a single code path — same prompt word for word, same size, same output format, three runs each with the median shown. The eight tests in it are built from the list OpenAI's own migration guide names. It goes up when those runs exist, and not before.