Overview

  • xAI launched Grok 4.7 on Monday, following five delays since late July.
  • This model ranked second to Claude Fable 5.1 on both GDPval and AA-Briefcase, and it also placed second to GPT-6 Astra on the EEBench electrical engineering benchmark.
  • Grok 4.7 is now available in the Grok app, Cursor, Grok Build, and through the xAI API, featuring 2.1 trillion parameters with additional training on SpaceX engineering data.

Elon Musk’s xAI released its most advanced model yet, Grok 4.7, on Monday afternoon, claiming it represents "a significant enhancement over Grok 4.6 at the same speed and price."

This latest release follows multiple delays, with Musk adjusting the timeline at least five times since late July, initially stating it was "four weeks out," then changing it to "a few weeks," and subsequently to "3 to 4 weeks," before finally saying on September 1 it needed "just a few more days to cook" on September 11.

Myriad: How much will Elon Musk be worth? Click to make your prediction.

xAI indicates that Grok 4.7 dedicates more time to complex problems and verifies its answers more frequently compared to Grok 4.6, alongside implementing what they describe as the best safety features to date.

Musk followed up on X, praising Grok 4.7 as "a strong combination of intelligence, speed, and affordability." Unlike previous launches, there’s no waitlist; it is available now via the Grok app, Cursor, Grok Build, and xAI API.

Grok 4.7 is here.

It's a notable improvement over Grok 4.6 at the same price and speed. pic.twitter.com/H3OTBbXyvO

— SpaceXAI (@SpaceXAI) September 21, 2026

The new model boasts 2.1 trillion parameters, representing a 40% increase from the 1.5 trillion in Grok 4.6, which was an upgrade from Grok 4.5. The pricing is set at $2 per million input tokens and $6 per million output tokens. Parameters serve as the internal settings a model adjusts during training, with a higher number typically allowing for better pattern recognition, while tokens are the basic units of data an AI can either process or produce.

Additionally, xAI incorporated supplementary training data from SpaceX, including Starlink satellite telemetry, manufacturing logs, and engineering failure records. This aims to enhance the model’s reasoning capabilities regarding hardware and physical systems compared to those trained solely on internet text.

However, benchmark results reveal a familiar trend. GDPval evaluates models based on their performance in producing valuable economic outputs—such as legal documents and spreadsheets—scoring them using an Elo rating system akin to that used in chess.

Grok 4.7 achieved a score of 1695 on GDPval, while Claude Fable 5.1 led with a score of 1735.

Similarly, AA-Briefcase, which assesses multi-hour office tasks involving research and document creation, also uses the Elo system. Grok 4.7 scored 1657, compared to Fable 5.1's 1678, showing a consistent pattern across different assessments.

CursorBench 4.0, which measures coding tasks within Cursor's editor, compares accuracy with task costs and token consumption. Grok 4.7 is positioned in the middle, being more expensive per task than GPT-5.6 Sol's successor, GPT-6 Astra, and Claude Sonnet 5, but still trailing behind Fable 5.1, which outperforms at every price point.

This pattern is not new for xAI. The previous model, Grok 4.5, launched in July, utilized the largest training cluster in the industry but still ranked third behind Claude and OpenAI’s models. Earlier, Grok 4.20 prioritized speed and personality over reliability, while Grok 4.6 also fell short in coding autonomy compared to competitors.

Despite these rankings, Grok 4.7 remains a valuable tool for the millions who interact with it via X, the standalone app, or Tesla's dashboard. However, it indicates that the model most likely to assist users is operating on a level that its own developer's benchmarks suggest is not at the forefront of technology.

This performance gap underscores the importance of pricing over leaderboard standing for many users. xAI has consistently offered lower costs per token than competitors like Anthropic and OpenAI, even as it lags in raw performance, banking on the notion that "good enough, affordable, and accessible" is preferable to "the best, but more expensive" for everyday applications.

In anticipation of the launch, Musk had tempered expectations, stating Grok 4.7 would likely be "roughly on par with" Anthropic's Claude Opus 5.0, rather than the newer Opus 5.1, and that multimodal capabilities still required enhancements.

He also outlined the future models: Grok 4.8 is expected to be a significant upgrade, Grok 4.9 is projected to be in the "Astra/Fable class," and Grok 5 could potentially lead the frontier. Release dates for these forthcoming models remain unannounced. "We shall see," he commented.

Daily Debrief Newsletter

Stay informed with the latest news stories, original features, podcasts, videos, and more.