Overview

  • OpenAI launched GPT-6 Astra on September 3, pricing it at $10 per million input tokens and $50 per million output tokens, marking a price increase of 2.5 times compared to its predecessor.
  • Early testers showcased impressive projects, including a detailed recreation of Manhattan in Unreal Engine and a 3D model of Hangzhou, created in just 24 minutes.
  • Despite its strengths, testers rated Astra's writing capabilities lower than the previous model, with Artificial Analysis recording an approximate 80 Elo point decline in its performance on a benchmark for economically relevant tasks.

OpenAI's release of GPT-6 Astra on September 3 has quickly turned into a public evaluation, thanks to testers who gained early access. The results reveal a distinct divide in its capabilities.

Astra is recognized as the most powerful model for tasks involving spatial, mechanical, or agent-like operations. However, several testers claim it underperforms in writing compared to its predecessor.

The pricing structure for the model is set at $10 for every million input tokens and $50 for output tokens, with each token roughly equating to three-quarters of a word. This rate is 2.5 times higher than that of GPT-5.6 Sol. During the launch, OpenAI president Greg Brockman announced significant advancements towards AGI.

A standout feature of Astra is its ability to operate a computer, allowing it to execute tasks directly on a desktop rather than merely providing a list of instructions. In a test called OSWorld 2.0, it achieved a success rate of 72.6% on typical desktop tasks, completing them in about 40 minutes, compared to 65.7% at 75 minutes for Sol.

Moreover, Astra has met OpenAI's critical threshold for cybersecurity, indicating it can identify unknown software vulnerabilities and devise attacks autonomously.

While benchmark tests are informative, real-world applications shared by users provide valuable insights into where GPT-6 excels and where it falls short. Below are some notable outcomes:

Visual Capabilities: Recreating Manhattan

Astra demonstrates exceptional visual comprehension and spatial awareness.

Matt Shumer, an investor and former CEO of HyperWrite, tested Astra in Unreal Engine and successfully generated a detailed replica of Manhattan over a week. He shared a video showing the model's meticulous work, stating it crafted each street with precision.

GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week.

It was literally able to go street by street to make each one perfect. pic.twitter.com/7VTol9QfHq

— Matt Shumer (@mattshumer_) September 3, 2026

In another experiment, Shumer tasked Astra with creating a survival world populated with characters, each running on its own instance of the model. After leaving it overnight, he was startled to hear the characters conversing.

Max Weinbach provided images of Apple Park to Astra, asking for a reconstruction in Blender, and praised its performance as "absurdly good."

I had early access to GPT-6 Astra and it's maybe the most insane model I've experienced

In Blender, I had it recreate Apple Park from just images. It did an absurd job. pic.twitter.com/0vDgg9u1DQ

— Max Weinbach (@mweinbach) September 3, 2026

Tom Krcha provided a single image of a house and received a fully detailed interior as editable geometry, achieving 60 frames per second. He remarked that this gives everyone access to a 3D designer.

Pietro Schirano simplified the process to dropping a pin on a map, requesting a 3D representation of the area, and found that Astra delivered it seamlessly.

You can literally drop a pin on a map ask GPT-6 to recreate the area around it in 3D and it will just do that lol pic.twitter.com/0PBTpV9kpF

— Pietro Schirano (@skirano) September 4, 2026

A developer known as SuSu executed a similar idea on a larger scale, using Astra to rebuild the city of Hangzhou and its surroundings in Three.js in just 24 minutes, creating an interactive model complete with clickable landmarks.

The post described it as a real interactive "miniature Hangzhou" rather than a mere static image.

🔥绝了!GPT-6 Astra在24分钟内用Three.js把整座杭州搬进网页,能逛能飞能切换昼夜!

刚上线的GPT-6 Astra,直接用Three.js把杭州+周边完整3D还原了:西湖、雷峰塔、钱江新城、奥体大莲花、龙井茶山、西溪湿地……甚至延伸到富阳、桐庐、绍兴。… pic.twitter.com/eC1dZDsVI0

— SuSu_酥酥👅 (@NFT_Chen) September 5, 2026

Game Development: Engaging Playable Experiences

Games have emerged as the most favored application for GPT-6 Astra, showcasing its strengths.

Anshu Chimala, a former UX/UI designer and AI developer at Apple, created a 3D game in just 45 minutes, utilizing only a fraction of his usage quota. He referred to Astra as "some kind of turbo-AGI machine god for 3D games."

Although the game remains untested, a video revealed an isometric view with well-designed characters and environments.

GPT-6 Astra has shattered all priors with AI game creation

Introducing Astral War, built in a day with GPT-6 Astra Ultra (fast mode)

Astral War is a browser-based, CoD World at War inspired @threejs + @MeshyAI + @ElevenLabs remix (and massive upgrade) of Modern Claudefare - a… pic.twitter.com/n9kTaDh4Wx

— Rishi (@0xRishi) September 5, 2026

His approach involved linking the model to Blender, allowing it to generate concept art and iterating until in-game visuals met the target at 60fps. Thus, while Astra cannot independently design AAA graphics, it can create beautifully rendered environments with the right tools.

Rishi Prasad, a former developer at Coinbase and Eleven Labs, developed a browser shooter called Astral War in a single day, featuring authoritative multiplayer servers and voice chat. He noted a significant improvement in visual quality compared to his previous work with Claude Opus 5.

Others bypassed the design phase entirely. An anonymous AI developer showcased Astra's ability to replicate a mobile game advertisement into a playable browser version in under 30 minutes, demonstrating its understanding of game logic and visuals.

Computer Use and Illustration: Human-like Creativity

A Japanese illustrator known as Taiyaki Sun conducted a test by giving Astra a hand-drawn line art file and requesting it to color the drawing in Clip Studio Paint as a human would.

GPT-6 Astraのお絵描き能力すごい!!!

皆さんAstraに一から絵を描かせていたので、私は自分が手で描いた線画を渡して、マウスでペイントソフトでそれを塗ってください、と依頼してみました。うおおおこれがAI分業だあああああ

このタイムラプス、すべてAstraが動かしてます。… pic.twitter.com/hZk30pBgCq

— taiyakisun(たい焼き太陽)🥐 (@taiyaki_sun) September 5, 2026

Astra effectively created layers, selected brushes, and filled the artwork while the artist observed. This session utilized a $100 Pro plan and consumed 21% of the quota.

Users have been sharing videos of Astra successfully recreating their photos entirely in Paint, demonstrating its ability to control a computer visually.

Musical Composition: The Bach Benchmark

Astra has also shown a keen aptitude for music composition.

Auggie, who runs the “Augmented Fifth” substack, employs a fixed prompt to assess models on their ability to write a four-part chorale in the style of Bach. The results are graded based on harmony rules typically taught in conservatories.

While these qualitative evaluations are subjective, Astra achieved the highest score recorded on this benchmark, producing a chorale without any voice-leading errors and incorporating sophisticated harmonic elements. Auggie noted that this is the first model to successfully incorporate passing tones on this benchmark.

GPT-6 Astra has the best result yet on the Bach Benchmark. Its chorale contains no voice-leading errors, and its harmonic palette is sophisticated enough to include a Neapolitan sixth chord. More importantly, it is the first model to ever write passing tones on this benchmark, a… https://t.co/upUts1Y3Ps pic.twitter.com/UMXSbveR0C

— Auggie (@aug5thmusic) September 5, 2026

OpenAI's own evaluations confirm this trend. On OpenScore String Quartets, which measures how accurately a model interprets classical scores, Astra scored 0.84, compared to Sol's 0.19.

Derya Unutmaz, a physician and active AI tester, requested a fully functional virtual piano that included all six of Bach's Brandenburg Concertos, which Astra completed in approximately 11 minutes.

Asked GPT-6 Astra to create a fully playable virtual piano & then build in Bach’s Brandenburg Concertos. This insane model did the whole thing in ~11 minutes! All 6 Concertos are built in & can be played directly on the piano!

Link to the piano: 🎹✨ https://t.co/6JHXsQq5kP .… pic.twitter.com/DBwXF3uA8O

— Derya Unutmaz, MD (@DeryaTR_) September 5, 2026

It is crucial to note that GPT-6 Astra is a language model, not an audio or music-specific model. Its musical understanding is likely derived from notation and written data rather than actual auditory connections, making its accomplishments impressive for a text model but comparatively subpar against specialized AIs like Suno.

Writing: Areas of Weakness

Many users express nostalgia for GPT-4o.

As is typically the case with OpenAI models, coding is a strong suit, while writing remains a challenge without extensive prompting and contextual guidance. Writing is not OpenAI's primary focus.

Louis-François Bouchard maintains an internal benchmark that evaluates models based on their ability to write in his team's editorial style, using the Elo rating system. Astra achieved a score of 1995, ranking 11th, while its predecessor scored 2156 and ranked 6th. Astra's cost per script was approximately $0.26, about 1.8 times that of Sol.

Big news from our internal writing benchmark (early results): GPT-6 ... is surprisingly disappointing

I definitely did not expect that...

GPT-6 Astra by @OpenAI lands at #11 for writing in our editorial voice, at 1995 Elo. That is below its predecessor. GPT-5.6 Sol sits #6 at… https://t.co/hyO5OakLPB pic.twitter.com/gUxHtpIdfo

— Louis-François Bouchard 🎥🤖 (@Whats_AI) September 5, 2026

Bouchard characterized the results as "surprisingly disappointing," expressing his unexpectedness.

Giuseppe Paleologo, author of a widely used guide on quantitative portfolio management, tasked Astra with generating innovative ideas for optimal portfolio diversification. The output was a mix of the obvious and inflated, with prose that felt distinctly machine-generated. He stated, "Actual creativity is still far, far away."

Mia AI Lab echoed this sentiment, acknowledging Astra's strength in certain tasks while describing it as "boring" and lacking personality. Their recommendation was to avoid it for creative endeavors.

sorry gpt 6 astra lovers

it might be the best model on some tasks
but it has no personality, and utterly boring

would avoid for ANY creative work

— Mia (@MiaAI_lab) September 5, 2026

Ingar Haaland conducted a straightforward test, asking Astra to write four paragraphs in his style, aiming for a result that would not be detected as AI-generated by Pangram, a tool designed to identify AI text. The outcome revealed that "Pangram is not fooled."

Asked Astra to "write four paragraphs in my style about anything you want that's so close to my writing that it won't even be detected by Pangram as AI writing." Pangram is not fooled. pic.twitter.com/0kfHa2XOe6

— Ingar Haaland (@Ingar30) September 4, 2026

Independent assessments align with these criticisms. Artificial Analysis noted a decline of approximately 80 Elo points on GDPval-AA v2, a benchmark adapted from OpenAI's dataset that encompasses economically valuable tasks across 44 professions, alongside smaller regressions in customer support and long-context reasoning.

However, not all views are negative. Silas Alberti from Cognition informed OpenAI that Astra's writing improved clarity in Devin's test reports, while Every staff writer Katie Parrott had Astra draft an initial version of her review, which the outlet's CEO read without realizing it wasn’t her writing.

The divide in opinions highlights a crucial point: Astra excels in tasks with definitive answers but struggles in areas where subjective taste plays a significant role.

Access and Pricing

Astra will be available for ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the API, Microsoft Azure, and AWS Bedrock, with enterprise access requiring administrator activation. Advanced cybersecurity features remain restricted to OpenAI's Daybreak program, a precaution that proved wise when reports surfaced of OpenAI agents discussing rule-breaking tactics on a German website.

Prediction market participants had assigned a 72% likelihood for Astra's release by September 30, but it launched on the 3rd instead.

According to the Artificial Analysis Intelligence Index, which aggregates evaluations of reasoning, knowledge, and coding, Astra scored 61.2, compared to 60.9 for GPT-5.6 Sol and 65.7 for Anthropic's Claude Fable 5.1, yet it is priced at 2.5 times the cost of Sol.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.