Skip to main content

Images 2.5: $527 vs $2,110 for 10,000 Images

Images 2.5: $527 vs $2,110 for 10,000 Images

September 9, 2026 · by opensourcefactory1 · all model claims below are vendor-reported unless stated otherwise

I opened OpenAI's Images 2.5 announcement expecting the usual sharper-faster-better bingo card. Then I hit the pricing calculator and stopped scrolling, because the same "high" quality setting now costs roughly a quarter of what it did last week. Same pixels, new labels. That pricing twist turned out to be the most concrete thing in the whole launch.

I cover open image and video models on this blog, and this one is closed weights, so why write about it? Because the pricing relabel is a genuinely consumer-friendly move wrapped in confusing labels, and because Sketch, the doodle-to-image feature, is the first ChatGPT image tool in a while that made me want to try it myself. Short version up front: the model looks like a solid step forward, the tools are the fun part, and the price story needs a decoder ring, which is what the rest of this post is.

So what actually shipped?

On September 8, OpenAI released ChatGPT Images 2.5, five months after Images 2.0 from April. It is rolling out to every ChatGPT, ChatGPT Work, and Codex user on all tiers, desktop, mobile, and web, which includes the free tier. For developers there are two new API models, GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst. Some scale context from the announcement: people now make more than 3 billion images a week across ChatGPT Images and the API models.

The model claims, all vendor-reported: sharper details, more natural lighting and richer textures, better preservation of subjects from your reference photos, and editing instructions that hold up better across multiple turns. Latency is down by up to 50% versus Images 2.0. That "up to" is doing real work, as always, but there is at least one outside data point suggesting the speedup is real, more on that below.

AI-generated 80s-style portrait example from the official Images 2.5 announcement

The kind of reference-faithful stylized portrait OpenAI is showing off: an 80s neon headshot. Image: OpenAI official announcement.

Sketch is the part I would actually use

Sometimes the clearest way to explain an idea is to draw it badly, and OpenAI finally agrees. Sketch lets you draw directly inside ChatGPT and use the doodle as a visual guide for the final image. Room layout, an outfit contour, a funny stick figure. You type "@Sketch" in chat, scribble, add a style description, and the model turns your rough art into a finished image. No artist skills required, which describes me perfectly.

Product lead Adele Li told Axios the sketch tool and templates are "an attempt to make people more engaged in the process of creation, rather than them being a passive force," and that she doesn't want to be able to "see ChatGPT in the world." I like that framing. It also reads as an honest admission that AI image output has a recognizable house style problem, and this is one attempt to fix it.

The other new tools: templates for popular formats like Poster and Merch so you don't start from a blank canvas, comments you can pin directly onto an image for focused edits ("lighter HERE" with a circle, exactly like Axios did with a tattoo design), and shareable prompts so someone else can rerun your idea with their own photos. Axios' hands-on testing is worth reading here: they turned a photo of a cat into a stained-glass window and a joke anatomical diagram, superimposed tattoo additions next to an existing arm tattoo across several edit rounds, and built a merch mockup for a logo. The tattoo case is the telling one, because multi-turn edits where earlier changes survive is precisely what OpenAI claims got better.

ChatGPT Images 2.5 editing interface with markup comment and remove background tools

Axios' hands-on: the new editing screen with Markup, Comment, Remove BG, Erase, and Resize tools. The anatomical-cat diagram shows text rendering holding up inside a fussy vintage style. Screenshot: Axios, used with credit.

AI-generated merchandise mockup with logo applied to shirts tumbler tote and stickers

Same session: a logo applied across shirts, a tumbler, a tote, and stickers. Single-element edits that preserve the rest are the stated API use case. Image: Axios, used with credit.

Flare vs Sunburst: same price, different wait

For API users the choice is Flare or Sunburst. Per-token rates are identical across Flare, Sunburst, and gpt-image-2 Standard: $8 per 1M image input tokens ($2 cached), $30 per 1M image output tokens, $5 per 1M text input ($1.25 cached). So the models don't differ in price at all. They differ in time. Flare is the default for most apps, higher quality than gpt-image-2 at 50% lower latency, aimed at creator content, product experiences, visual search, rapid prototyping, high-volume generation. Sunburst trades longer generation times for tighter control across edits: campaign creative, polished product imagery.

Early-customer quotes in the launch post back this up directionally. Adobe's Matt Chotin says both models are now in Firefly. Manus' evaluation team says Flare ran two to four times faster than gpt-image-2 in their tests with better transparent backgrounds. Higgsfield's head of product says the model preserves a source image's character and composition through edits. All vendor-selected quotes, so read them as direction, not proof. But they agree with each other, which is at least consistent.

The quality ladder got relabeled, and "high" now costs a quarter

Here is the part worth a decoder ring. Apidog worked through OpenAI's pricing calculator and found the per-image costs moved a lot even though per-token rates didn't. GPT-Image-2.5 has five quality levels now: low, medium, high, xhigh, max. The token budget sitting under the label "high" is 1,756 output tokens at 1024x1024, which is what gpt-image-2 spent at "medium." And "max" at 7,024 tokens is the old "high" budget. So keeping quality: high in your config gets you the old medium render at about a quarter of the old high price, and matching the old high look means moving to max at effectively the old price. Xhigh is a genuinely new middle rung with no predecessor.

Bar chart of per-image cost across the five GPT-Image-2.5 quality levels

Per-image cost at 1024x1024 across the five 2.5 quality levels, from OpenAI's calculator estimates. High sits on the old medium budget; max matches the old high. Vendor-reported, redrawn.

2.5 labelTokens (1024x1024)CostSame budget on gpt-image-2
low196$0.00588low ($0.006)
medium439$0.01317no equivalent
high1,756$0.05268medium ($0.053)
xhigh3,122$0.09366no equivalent
max7,024$0.21072high ($0.211)

Two things to get right before anyone quotes "half price" or "double price" at you. First, the "2x pricing" rumor in early coverage came from comparing 2.5 Standard rates against the gpt-image-2 Batch row, which carries a 50% discount. Standard versus Standard, every rate is unchanged. Second, reference images and prompt text bill on top: a reference photo sent to the edits endpoint costs image input tokens, and OpenAI doesn't publish a per-image input count, so read usage.input_tokens from your own responses.

Bar chart comparing monthly cost for 10000 images across three quality settings

10,000 square images a month, output tokens only: $526.80 at 2.5 high, $2,110.00 at old high, $2,107.20 at 2.5 max. Same render budget, same bill; a smaller budget, a quarter of it. Author math from calculator estimates.

What are outsiders actually measuring?

Two non-OpenAI signals exist so far, both with caveats. The community Arena leaderboards (votes, updated September 7) rank Sunburst first in text-to-image at 1421 and Flare second at 1399, with gpt-image-2 at 1381 and Nano Banana 2 at 1261. On the image-edit board the gap widens: Sunburst 1520, Flare 1491, gpt-image-2 1461. Crowd preference, not a benchmark, but note the pattern matches OpenAI's own framing: the edit gap (29 points) is bigger than the generation gap (22 points), which is exactly where Sunburst is supposed to earn its longer wait.

Two-panel bar chart of Arena text-to-image and image-edit scores

Arena community votes: Sunburst leads both boards, and the edit lead is the wider one. Crowd preference, not an OpenAI benchmark. Vendor-reported release, redrawn.

On speed, one HN user running about 50,000 images through gpt-image-2 for an AI UI design tool reports average latency holding around 104 seconds on gpt-image-2 with 2.5 images coming in around 35 to 40 seconds. One workload, one data point, but it lands in the same zone as OpenAI's "up to 50%" latency claim and Manus' "two to four times" figure. If your pipeline is latency-bound, Flare at high is the obvious first experiment.

Should you switch? Who is each model for?

Default to Flare. That is OpenAI's own guidance and the numbers support it: same bill as before per token, faster, better quality at the same label, and a new cheaper default if you leave quality at high. Switch to Sunburst for multi-turn edit workflows where a reference subject has to survive several rounds of instructions, and only after a side-by-side on your own prompts. If the 22-to-29-point Arena gap doesn't show up on your images, you're paying latency for nothing. Keep gpt-image-2 around for two reasons: the 50% Batch discount tab currently lists only gpt-image-2, and a pinned snapshot like gpt-image-2-2026-04-21 that someone already visually reviewed is a feature, not technical debt.

The migration itself is small but the relabel is the trap: old medium becomes new high, old high becomes new max, low stays low, and do it in config rather than per call site. Also check whether the Batch endpoint accepts the 2.5 ids before cutting over a batch workload, and re-measure usage.output_tokens for a week because the calculator figures are estimates by OpenAI's own warning.

My take

This one landed a week after the GPT-6 Astra launch, and the contrast is funny: last week was "welcome to the AGI era," this week is "your doodles can become posters and your 10k-image bill just fell 75% if you read the labels right." I know which one I had more fun reading. The honest caveats are the usual ones. Every quality claim is vendor-reported with no independent benchmark yet, the safety numbers (1.09% unsafe generations for Sunburst, 1.41% for Flare, versus 1.64% for 2.0) come from OpenAI's own adversarial set that the system card admits isn't representative of production traffic, and the sharper realism makes the deepfake concern sharper too, which OpenAI answers with C2PA metadata plus SynthID watermarking. Still, a faster default model, a genuinely cheaper everyday tier hiding under a familiar label, and a sketch tool that lowers the floor for who gets to make images: that's a good week for an image model. I'll post again when independent benchmarks land or when someone measures Flare against Nano Banana 2 properly.

Sources: OpenAI announcement (Sep 8, 2026), OpenAI pricing page + calculator, Images 2.5 system card, Axios hands-on, 9to5Mac, Unite.AI, Apidog pricing breakdown (figures via their calculator read, Sep 9). Real-content images credited in captions (OpenAI, Axios). Charts: my own redraws. Related on this blog: the GPT-6 Astra launch post and the Astra vs Fable 5.1 pricing comparison.

Comments