Key Highlights
- ChatGPT Images 2.5 was released on September 8, featuring enhanced detail, improved editing capabilities, and latency reductions of up to 50% compared to its earlier version.
- In a comparison across six testing categories, Nano Banana 2 secured victories in three, while ChatGPT Images 2.5 claimed the remaining three.
- The outcomes hinged on minor mistakes rather than glaring issues.
OpenAI introduced ChatGPT Images 2.5 on September 8, emphasizing improvements such as sharper detail, richer textures, more natural lighting, and an editing process that respects user instructions.
The company reports that image generation latency has improved significantly, with reductions of up to 50% when compared to Images 2.0. Two new models—GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst—are now available in the API, with Flare designed for speed and Sunburst tailored for high-precision edits.
Myriad: Which company will IPO next? Click to make your prediction.This latest release follows a previous evaluation conducted in May, where GPT Image 2 and Nano Banana 2 competed across eight categories, with GPT Image 2 winning more overall but exhibiting a notable oversharpening issue under complex prompts. This led to a lingering question: Would the new model rectify this flaw?
To find out, we conducted a fresh comparison using the same evaluator, adapting six categories from prior evaluations and applying identical prompts to both models, with Google’s Nano Banana 2 being the version known as Gemini 3.1 Flash Image, not the slower Nano Banana Pro.
This matchup features OpenAI’s fast and precise model against Google’s equivalent offering.
Changes in GPT-Image 2.5
OpenAI's image models have historically had a signature flaw. The original GPT Image 1 was infamous for a persistent warm yellow tint dubbed the "piss filter", a problem that OpenAI never fully addressed.
With GPT Image 2, this issue was replaced by another: prompts with excessive constraints would lead to oversharpening, resulting in images that appeared overly processed and artifact-laden.
This time, however, neither of these flaws was present. Each image generated by Images 2.5 maintained its color balance and intricate detail, free from the yellow tint or excessive sharpness that plagued its predecessors. This improvement was particularly evident in tests involving steampunk and portrait images, showcasing the most photographically coherent outputs we've seen across three generations of models.
OpenAI has also introduced new features that enhance workflow, such as Sketch for drawing rough layouts directly within ChatGPT, prompt sharing, inline comments for specific image areas, and format templates for posters and merchandise. On the API side, quality tiers now extend from low to a new high and max level, surpassing the capabilities of Images 2.0.
Here’s how OpenAI’s latest model compares to Google’s top offering.
Text Clarity: The Kellerman's Hardware Scene
This challenging test involves a gritty 2 a.m. intersection with nearly every surface featuring readable text, including ghost signs, graffiti, storefront letters, concert posters, curb stencils, and stickers on payphones.
Nano Banana 2 managed to render most elements clearly, with a minor slip involving a payphone sticker that duplicated its text in a confusing manner.
Nano Banana 2: Kellerman's Hardware sceneIn contrast, ChatGPT Images 2.5 included a detail that Nano Banana omitted: a lamppost adorned with overlapping, weathered flyers, precisely as described in the prompt.
ChatGPT Images 2.5: Kellerman's Hardware sceneHowever, the model’s street-art tag was incorrectly rendered as "STILLL HERE" with an extra L, and the apostrophe in "KELLERMAN'S" on the ghost sign was either missing or illegible. For a model that emphasizes precise text rendering, these two legibility issues are significant misses.
Aesthetically, OpenAI’s model produced a more realistic scene, but it fell short in text accuracy compared to Google’s output.
Winner: Nano Banana 2.
Spatial Awareness: The Steampunk Clock Tower
This complex aerial composition test features a five-plane depth scene with a large clock tower displaying different times in legible Roman numerals, alongside six other text elements spread across the image.
ChatGPT Images 2.5 delivered a more atmospheric image, complete with visible steam rising from rooftops and a river in the mid-ground, along with a richer tonal range across the depth planes. The letters in this generation were also easier to read.
ChatGPT Images 2.5: The steampunk clock towerNano Banana 2's atmosphere appeared flatter, but it successfully rendered both clock faces with legible Roman numerals—XII, III, VI, IX—at similar hand positions, which didn't align with the prompt's request.
Nano Banana 2: The steampunk clock towerWinner: GPT Images 2.5, for better adherence to instructions without compromising realism.
Illustration: The Anime Spirit Medium
This prompt called for a Studio Ufotable-style key visual: a girl transforming into spiritual energy at a torii gate, accompanied by a nine-tailed kitsune fox and a twilight sky reminiscent of Makoto Shinkai's work.
ChatGPT Images 2.5 provided the most impressive sky of any test in this series, featuring a visible sun disc, water reflections, and a mountain silhouette that truly captures the essence of Shinkai’s style. The asymmetric eyes were also well-executed. However, the model deviated slightly from the prompt regarding the "dissolving into energy" aspect, interpreting it as an electric effect in her hair instead of the flowing dissolve described.
ChatGPT Images 2.5: The anime spirit mediumNano Banana 2 presented a wispy blue-white energy trail that more closely matched the specific request. Neither model convincingly rendered the nine-tailed fox.
Nano Banana 2: The anime spirit mediumWinner: ChatGPT Images 2.5, based on overall visual impact, despite the interpretative divergence.
Realism: The Rooftop Architect
This cinematic portrait required a variety of specific elements: beige trench coat, round glasses, blueprints in the left hand, golden-hour lighting, shallow depth of field, and film grain.
ChatGPT Images 2.5 produced stunning lighting, with the sun disc visible behind the subject and excellent skin micro-texture. Although the skin appeared overly smooth, the overall scene felt very realistic, akin to a shot taken with an analog camera.
ChatGPT Images 2.5: The rooftop architectInterestingly, adding commands typically associated with poor quality enhanced its realism. One generation included keywords like “realistic, highlights, crushed shadows, uneven flash, blown out skin tones, candid moment, shot on a phone camera.”
ChatGPT Images 2.5: The rooftop architectNano Banana 2 maintained a fuller composition, with the blueprints in the right hand rather than the left, and included a visible blueprint label—"PROJECT: 124 DUANE ST"—that many renders omitted.
Nano Banana 2: The rooftop architectWinner: Nano Banana 2, for one-shot accuracy, while GPT excels in repeated iterations.
Agentic Research: The Bitcoin Timeline
Both models can perform research on a topic before generating an image, so we requested a widescreen Bitcoin history timeline in a child-drawing style, with an emphasis on factual accuracy. This category proved crucial for the comparison, revealing the most significant gap.
ChatGPT Images 2.5 created a two-row infographic with specific dates, but one date was incorrect, mistakenly labeling 2023 as the year Bitcoin ETFs were approved in the U.S. The SEC actually approved the first spot Bitcoin ETFs on January 10, 2024, a year later, although futures ETFs were approved that year, suggesting a potential interpretation issue.
ChatGPT Images 2.5: The Bitcoin timelineNano Banana 2 produced a less structured output but included similar events, reflecting the model's consistency on repeated prompts. However, it cleverly hedged the ETF approval and the fourth halving into a "2023–2024" timeframe, avoiding any outright inaccuracies.
Nano Banana 2: The Bitcoin timelineWinner: Nano Banana 2. A confidently incorrect date is more significant in a category specifically assessing the accuracy of "agentic reasoning" outputs.
Abstract Concepts: A Prompt of Invented Words
This prompt involved a scenario composed of nonsensical terms: "A woman eating shmfiyxl in Lyxin. Next to her, her Lymglsushing plays Lakishkark." With no dictionary definitions for these words, each model had to create an interpretation and visually represent it.
ChatGPT Images 2.5 addressed this by incorporating the nonsense directly into the scene as text. "Lyxin" appeared as glowing signage above a sleek, futuristic restaurant, and "Lakishkark" was featured on a board game box held by the woman’s alien companion. The previously abstract terms became concrete when represented as readable labels.
ChatGPT Images 2.5: A prompt made of invented wordsNano Banana 2 took a different approach, crafting a warm, culturally rich scene featuring a Guatemalan market stall, a woman in a traditional huipil, and a creature resembling an orc playing a hybrid string-and-pipe instrument. However, it interpreted "plays" as making music rather than playing a game, resulting in a valid but divergent interpretation. Nevertheless, none of the invented terms were represented in the scene.
Nano Banana 2: A prompt made of invented wordsWinner: ChatGPT Images 2.5. Transforming undefined concepts into visible labels is a more direct response to a prompt offering no concrete guidance.
In conclusion, Nano Banana 2 won three out of six categories in this evaluation, making it nearly a tie. The outcome largely depends on individual expectations and interactions with the models. ChatGPT Images 2.5 has successfully addressed the oversharpening issue that affected its predecessor, and its illustrative output in this round could be considered the most visually impressive single image produced by either model throughout two rounds of testing.
The distinctions between the two are not about overall quality, but rather specific, verifiable issues: minor errors in lettering, a mistaken year in a research infographic, etc. Overall, their aesthetic and quality levels appear closely aligned.
