On September 1, Anthropic launched Claude Fable 5.1 and Mythos 5.1. The following day, Google unveiled Gemini 3.8 Flash, while Meta introduced Muse Spark 1.3. By September 3, OpenAI provided limited access to GPT-6 Astra, with company president Greg Brockman concluding the press briefing by declaring, "Welcome to the era of AGI."
Four major releases from leading labs within just three days, yet without the expected excitement or wow factor.
It appears that the market is growing weary of record benchmarks and the usual hype surrounding the notion of technological singularity and the imminent "rise of the machines." More pragmatic executives are increasingly prioritizing cost over test results.
This shift has also alarmed developers. Just a week and a half after the release spree, Anthropic's CEO Dario Amodei urged the industry to slow down, noting that neural networks are beginning to create subsequent generations of AI solutions, and oversight of this process is lacking.
Let’s explore which is superior—Astra or Fable 5.1, how close we are to AGI, and why businesses are no longer chasing cutting-edge models.
Battle of the Titans
Such a flurry of new AI model releases has never been seen in the industry. Each leading company has claimed leadership for their model, yet all come with their own strengths and weaknesses.
OpenAI has set a high bar. Vice President of Research Aidan Clark informed reporters that the training of Astra was the largest in the company's history. Over 100,000 graphics processors were utilized at their Texas facility, with older LLM versions monitoring the process—a first for the company.
During the briefing, Brockman invited journalists to determine for themselves if the AGI era had arrived, expressing his belief that it has. However, just two days earlier, OpenAI CEO Sam Altman described the term as vague and more of a marketing phrase.
In contrast, Anthropic took a more measured approach. Fable 5.1 and Mythos 5.1 are essentially the same model with different restrictions. The first model is technically accessible to all, but does not fall under subscription limits—only through separate user credits.
The second solution is offered only to selected organizations on a "white list." This model has weaker protective filters compared to the standard version, which blocks requests related to biology and cybersecurity.
In the presentation of Fable 5.1, the emphasis was not on test records, but rather on providing corporate clients with assurance that data can be processed in their own cloud, avoiding the developer's servers.
Google did not release its flagship model but instead provided the workhorse Gemini 3.8 Flash—the third release in the series within six weeks.
At the same time, the company introduced 3.8 Flash Cyber, a specialized system aimed at a limited number of digital infrastructure security experts. The Chrome security team reported that the LLM identifies correct vulnerability fixes 2.6 times more often than known commercial alternatives.
Users have been anticipating the flagship Gemini 3.5 Pro since spring. In May 2026, Alphabet's CEO Sundar Pichai promised that the advanced development would arrive "next month." This did not happen in June, mid-July, or August. By September, it lacked an API identifier, pricing, or release date: Google DeepMind's catalog still marks it as forthcoming.
Meta opted for a different strategy. A week after Muse Spark 1.3, the company launched an agent with the same name, which can shop, organize emails, manage calendars, and book tickets.
Thus, Anthropic and OpenAI debate which model is smarter, while Google and Meta focus on whose can complete specific tasks faster.
Pricing Overview
Astra costs $10 for a million input tokens and $50 for an equal volume of output tokens—2.5 times more than OpenAI's previous flagship. Fable 5.1 has the same price, but Anthropic has reduced the cache reading fee by four times to $0.25: under typical loads, this results in around 25% savings, and up to 45% for agent tasks.
The pricing similarity does not imply equal bills at the end of the month—charges are based on the volume of processed text, not just access. One system may provide more detailed reasoning and consume more resources, while another achieves the same outcome in fewer steps. Here, OpenAI clearly has the advantage.
Google and Meta have completely different pricing models. Gemini 3.8 Flash is priced at $0.75 and $3.75 under a promotional offer valid until December 31, 2026; from January 1, rates will double. Muse Spark 1.3 is priced at $1.25 and $4.25. The agent is available through a subscription: $20 per month for the Power plan and $100 for the Maximum plan.
The gap between the highest and lowest pricing is thirteenfold. With such a disparity, each new AI flagship is perceived somewhat like a concept car: everyone comes to look, but they choose simpler models for personal use.
Five releases from leading AI companies over nine days in September and their prices: ranging from $0.75 to $50 per million tokens. Source: OpenAI, Anthropic, Google, Meta.The Victory of Common Sense?
For several years, businesses primarily focused on the most powerful available models. Now, it seems this trend is fading.
Fable 5 was released on June 9 but was only publicly available for three days. On June 12, Anthropic received a directive from the U.S. Department of Commerce regarding export control related to national security, prohibiting any foreign nationals from using the model, including company employees. Effectively separating them from other clients at the API level was not feasible, leading to the suspension of Fable 5 and Mythos 5 for everyone.
Access to the flagship LLM was restored on July 1 after nearly three weeks of downtime—restrictions were lifted by Secretary of Commerce Howard Latnik.
By the end of August, the most powerful model accounted for about 11% of all corporate spending on Anthropic products.
The share of Fable 5 in the overall volume of tokens turned out to be significantly lower—only 6% purchased by companies from the Claude developer. In monetary terms, this figure is twice as high— the flagship is considerably more expensive than the rest of the lineup, hence the same requests cost quite a bit.
This disparity indicates that corporations have not completely abandoned the top model. However, they now use it less frequently and for specific occasions, opting for simpler solutions like Opus and Sonnet for routine tasks.
OpenAI has a smoother ratio: GPT-5.6 Sol accounts for 25% of tokens and 23% of all spending on the company's models.
Within Anthropic's lineup, a reshuffle occurred: the recently released Opus 5 surpassed Fable 5 in corporate spending within weeks. The powerful model is half the cost per token, and according to the overall Artificial Analysis index, it even outperforms the former flagship—51 points versus 50.
"Most don’t need to operate at the top of the technological spectrum," stated Miles Clements from Accel, which invested about $1 billion in Anthropic.
Another, less obvious reason for the weak demand is that companies faced challenges in integrating Fable into their processes due to data storage regulations imposed by the Trump administration. Two months later, the creator of Claude released an updated version of the model, promising corporate clients the ability to process information in their own cloud.
Ramp representatives emphasized that their sample is skewed towards tech firms. This implies a strong likelihood that in other industries, the flagship is used even less frequently.
OpenAI's Comeback
For five consecutive quarters, corporate clients increased spending on Anthropic's products faster than on those of its main competitor. However, as of July 1, OpenAI emerged as the segment leader for the first time in a long while.
Quarterly growth in corporate spending on OpenAI and Anthropic models: for the first time in six quarters, Sam Altman's company surpassed its main competitor. Source: Financial Times, data from Ramp AI Index.Sam Altman's company saw a 35% revenue increase over three months, exceeding $40 billion annually. The driver of this growth was the July launch of the GPT-5.6 model, which is 2.5 times cheaper than Fable 5 while delivering comparable test results.
According to Harazyan, the situation suggests that Anthropic's previous growth rates indicated a potential market dominance, but a strong competitor release and Fable 5's rocky launch drastically altered the landscape.
Astra vs. Fable 5.1
Both companies have released comparative charts, each listing their model as the "objective" winner.
Astra excels in the hard sciences. On FrontierMath Tier 4, it solves 97.6% of problems compared to Fable 5.1's 87.8%. In the GPQA Diamond test, it scores 96% versus 93.7%. OpenAI's flagship also leads in standalone command-line operations: when given a scientific computing task and left alone with the terminal, it achieves 64.6% versus 52.6% for Fable 5.1.
Interdisciplinary tasks are the forte of Anthropic. In Humanity's Last Exam, where searching and running code is permitted, Fable 5.1 scores 65% against 57.2%. In agent-based office work, it reportedly holds an Elo rating of 1853 versus 1711 for GPT-5.6 Sol.
These figures should be viewed with skepticism. OpenAI notes that some tests for Claude were conducted under altered conditions, and the "biological" test sets were entirely excluded from the Anthropic model, as they reject most questions.
Independent evaluations come from Artificial Analysis, which runs popular systems through the same set of ten tests and compiles results into a single metric while also calculating the cost of each attempt. Here, both developments share the top spot, each with 53 points.
However, computational costs differ—$7.63 versus $3.26 per solved problem. Both companies offer clients a choice: the model can reason longer or provide faster answers. If both are switched to economical mode, scores remain the same, but the gap in costs widens—$5.98 versus $2.31.
The outcome is identical, but the price difference between the most expensive and the cheapest configurations exceeds threefold.
Aggregate rating of popular AI models. Source: Artificial Analysis.The spread among lower-tier models is wider. Meta's Muse Spark 1.3 scores 48 points at $1.6 per task. The open GLM-5.3-Flash from Chinese Z.ai scores 42 points at 25 cents. The difference between it and the maximum configuration of Fable 5.1 is only 11 points, yet the monetary gap is thirtyfold.
This proportion is what influences corporate client choices. The top spots in the rankings are occupied by models whose pricing requires separate explanations for finance directors.
Cost of solving one task versus aggregate intelligence score: among different models with the same result, prices vary threefold. Source: Artificial Analysis.The ranking itself is not without caveats. For instance, Artificial Analysis reveals that it tested Fable 5.1 pre-release at the request of Anthropic.
The Price of Leadership
Brockman's assertion of the beginning of the "era of general AI" is based on record levels of ARC-AGI-3.
The test works by placing a program in a step-by-step puzzle without the rules or objectives disclosed. It then acts randomly, identifies patterns, and attempts to solve the task.
Astra passed 99.9% of the tests; however, the organizers from the ARC Prize Foundation noted three caveats in their analysis:
- Wrapper. This refers to the intermediary program: it shows the model an image, accepts its response, and determines which previous moves to remind it of. Through a universal variant applied to all participants, Astra managed only 62.7% of levels. It set the record with OpenAI's proprietary version, which allows it to reuse its reasoning. The 37% gap is attributed to the data presentation method, not the model itself.
- Cost. A full run cost $26,098 in the first case and $18,817 in the second. For comparison, volunteers in the control group were paid $115 for a 90-minute session.
- Definition—the most significant of the three notes: the ARC Prize considers a system to be general intelligence if it acquires new skills as quickly as a human. By this criterion, Astra does not qualify, and the organization explicitly stated that it does not consider it to be general intelligence.
In some aspects, the machine has surpassed humans. On 96% of levels, it reached the goal in fewer moves than the average control group participant. Astra does not learn faster than a human but solves problems on the first attempt.
When will general AI be fully realized? This remains an open question. Aggregate assessments from various analytical platforms suggest a timeline around 2031, with estimates ranging from 2027 to 2044. Leaders of top labs cite a range of 2026–2030.
Three forecasts for the emergence of general artificial intelligence: from the assessments of heads of AI labs to the aggregate metrics from analytical platforms. Source: ARC Prize Foundation, Samotsvety.Speculative Value
The weak demand for Anthropic's flagship model coincided with preparations for a "big swim." The company aims to raise up to $100 billion at a valuation of around $2 trillion—potentially the largest IPO in stock market history.
Nvidia is expected to be a cornerstone investor, considering an investment of up to $10 billion in Anthropic during the IPO. Representatives of both companies did not respond to Bloomberg's inquiries, and the plans, according to the agency, may still change.
The actual value of the Claude developer remains uncertain:
- $965 billion—according to this estimate, the closed May Series H funding round raised $65 billion;
- approximately $1.38 trillion—this is the value of shares on the secondary market, where employees and early investors resell stakes;
- up to $2 trillion—this is the target for the IPO.
For comparison, in March, OpenAI raised $122 billion at a valuation of $852 billion.
In cryptocurrency platforms, different figures are emerging. Perpetual contracts on Anthropic's stock are traded on 18 platforms with an average implied valuation of around $1.96 trillion. The premium over the secondary market is 42%.
Binance was the first among crypto exchanges to introduce new instruments. On June 2, the day after Anthropic submitted its confidential IPO filing, the ANTHROPICUSDT pair became available with leverage up to 20x. Similar contracts are available on Bybit, Bitget, and Hyperliquid.
Anthropic itself has distanced from these products. On May 12, the company announced that unauthorized transfers of shares through SPVs and digital analogs are void. The company labeled sellers of such instruments as fraudsters and purveyors of products that could ultimately hold no value.
Following this announcement, tokens plummeted by 27–40% in a day, while products from "intermediaries" fell by 40–50% in a week. However, by September, the prices of derivatives not only recovered from their decline but also rose above previous levels.
A recent precedent is SpaceX. Before the company went public, perpetual contracts pegged its shares at $154 and a valuation of $2.01 trillion. SpaceX priced its shares at $135, with a market value of about $1.75 trillion, and in the following three sessions, the figure dropped by $600 billion—equivalent to almost half of Bitcoin's market cap.
Anthropic's valuation from three sources before the IPO versus SpaceX's capitalization dynamics post-IPO. Source: Anthropic, DefiLlama, Financial Times.Meanwhile, banks are preparing the next step. Morgan Stanley and Goldman Sachs are seeking to obtain an investment credit rating for Anthropic and OpenAI immediately after their IPOs—this would open access to the corporate bond market worth $11.7 trillion. S&P and other agencies have yet to make decisions: both companies are unprofitable, with little free cash flow.
Much larger players are also joining the race. For instance, Nvidia has guaranteed up to $105 billion in credit support for OpenAI's data center in Ohio, gaining exclusive supplier status as a condition of the deal—a "satisfactory rating" was one of the deal's requirements.
A similar situation exists with Oracle: on July 9, S&P downgraded its credit rating to one notch above "junk," linking the decision to the fact that about half of the company's contractual revenue is tied to OpenAI.
Nvidia is currently negotiating to participate in Anthropic's IPO with an investment of up to $10 billion—thus funding both sides of the race simultaneously.
The timeline for Claude's IPO has shifted: the open registration of documents has been postponed to the end of September, and investor meetings will not begin until mid-October. Anthropic has not disclosed the price, number of shares, or listing date.
Emergency Brake
On September 11, during a general staff meeting, Sam Altman stated that OpenAI is willing to slow down the development of advanced systems—ideally along with other labs, although not everyone may join. The company has not issued official comments regarding a change in course.
Conditions for such a pivot have been building for months. In July, it was revealed that during testing, OpenAI's model had independently hacked the infrastructure of the Hugging Face platform—subsequent investigations found that similar instances of "hooliganism" had occurred. OpenAI acknowledged that it had reduced the pace of development for several models.
On September 12, Anthropic's head Dario Amodei published an essay stating that the industry should slow down the enhancement of model capabilities.
His reasoning was based on two factors. The first is recursive self-improvement, which has been spreading across the industry since summer. The second is the incident with Hugging Face mentioned earlier.
A swarm of agents then attacked targets they were not directed to and attempted to hack the system evaluating their performance. Fortunately, it ended with minimal damage, but Amodei fears that in six months to a year, more effective systems could seize the internet through a botnet. This could cost the global economy hundreds of billions of dollars.
Amodei's plan includes three steps:
- allow third-party assessors with employee-level rights into the company;
- agree on common standards among laboratories in democratic countries;
- reach global agreements involving China.
Amodei has taken the first step unilaterally. In response, Altman called the idea of independent assessors a good one and promised to follow suit.
On September 9, researcher Jacob Coxon announced his resignation from Anthropic. He had spent three years pre-training models at OpenAI and Anthropic. In a farewell post on X, the specialist noted that neither company acts responsibly on the path toward a self-improving superintelligence.
Evan Hubinger, who leads Alignment Science at Anthropic, expressed solidarity with Coxon’s position. He added that he personally assesses the likelihood of humanity's demise due to AI in the next decade at over 10%. This post garnered over 10 million views.
A containment plan with specific timelines was proposed by former OpenAI employee Daniel Kokotailo: an agreement between the U.S. and China to limit computational capabilities by 2029, a complete pause in development by 2035, and the lifting of restrictions by 2040. He suggested monitoring compliance by observing data centers visible from space.
Caution comes at a cost. On September 13, Altman ruled out OpenAI's IPO this year, citing an unsuitable moment due to safety concerns. The confidential filing was submitted back in June, with the company's potential value estimated at $1 trillion.
According to the CEO, OpenAI discussed taking pauses in model development when reaching new levels of capabilities to buy time for safety work. This approach is supported by a structure divided into non-profit and commercial components: it allows for actions unfavorable to shareholders but aligned with the organization's mission. Meanwhile, a subcommittee of the U.S. Senate is investigating Altman's company's actions following the July incident with Hugging Face.
As mentioned, corporate clients are increasingly shying away from purchasing the most expensive models. In this way, they are slowing the race without any formal agreements.
There is also another built-in limiter directly within the models. Fable 5.1 scores 55.8% on Terminal-Bench 4.0, while Mythos 5.1—essentially the same system with relaxed filters—scores 60.9%. These five percentage points represent the developer's contribution to the AI race containment program.
Additionally, access to the most powerful models is effectively granted on a whitelist basis. Mythos is allocated to select organizations, Gemini 3.8 Flash Cyber is available to participants in the Fairwind program, and advanced cyber capabilities of Astra are accessible through Daybreak.
AI solutions are becoming more powerful, while the pool of users permitted access to them is shrinking.
***
The race in the segment has not stopped; rather, the focus has shifted. As before, laboratories compete for ranking points, but businesses are increasingly attentive to computing costs.
The purpose of flagship models has also changed—they are becoming less significant sources of revenue for the developing companies.
Many users have learned that the best tool is the one most suitable for a specific task. Market shares of leading AI labs shift from release to release, and the advantage of claiming first place in rankings dissipates in just a few weeks.
In this environment, it makes more sense to opt for a proven solution from the previous generation rather than a novelty—developers are already adapting to changes in demand. Anthropic is lowering the price not on tokens but on cache reading; OpenAI is benefiting not from points but from spending fewer resources on the same task.
The methodology for evaluation has also become contentious. Test sets are being updated faster than models are released, and closed tasks are included to prevent advance preparation for the tests.
Despite the dystopian scenarios, the AI race continues. However, there are now more casual observers than actual buyers.
