On September 21, SpaceXAI unveiled Grok 4.7, a new model designed for programming, agent tasks, and knowledge management.
— SpaceXAI (@SpaceXAI) September 21, 2026
This latest version features a larger foundational model compared to its predecessor. Developers have also extended the reinforcement learning duration and incorporated more complex tasks that can take several hours to complete.
According to SpaceXAI, Grok 4.7 offers improved verification of its own conclusions and handles longer contexts more effectively. The model has also been trained to operate within the Grok Bot agent environment.
Testing Results
In the internal CursorBench 4.0 tests, Grok 4.7 achieved a score of 46.3%, surpassing Grok 4.6's 40.4%. For context, GPT-5.6 Sol scored 41.7%, while Claude Fable 5.1 achieved 51.8%.
Performance of Grok 4.7 and other AI models in CursorBench 4.0 based on average task execution cost. Source: SpaceXAI.In the AA Briefcase test focused on multi-hour tasks, the new model recorded an Elo score of 1657, compared to 1546 for its predecessor.
Results in EEBench improved from 53% to 64%, and in the Harvey Legal Agent Benchmark, the score increased from 15.8% to 19.6%, according to SpaceXAI's data.
Independent analysis from Artificial Analysis rated Grok 4.7 at 46 points on the Intelligence Index, which is two points higher than Grok 4.6. When paired with Grok Build, the model scored 56 points in the Coding Agent Index, up from 47 for the earlier version.
Among models operating in proprietary agent environments, it ranked fourth, following Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.
However, the new version consumes significantly more tokens for complex tasks. In tests by Artificial Analysis, Grok 4.7 in xhigh mode averaged about 81,000 output tokens, compared to 36,000 for Grok 4.6.
Compare Grok 4.7 (first) and 4.6 (second) building an open world city game. pic.twitter.com/uxo3ays7Bx
— SpaceXAI (@SpaceXAI) September 21, 2026
Pricing Remains Unchanged
For requests up to 200,000 tokens, Grok 4.7 is priced at $2 per million input tokens, $0.50 for cached inputs, and $6 for every million output tokens. For longer requests, the rates double.
The context window supports up to 500,000 tokens, and the model can process both text and images, generating text as output.
It offers four levels of reasoning depth: low, medium, high, and xhigh.
Additionally, SpaceXAI has introduced Grok 4.7 Fast, which doubles the generation speed at twice the cost. This version is exclusively available in Cursor and Grok Build, not through the public API.
SpaceXAI also announced a complete overhaul of the model's protection system. In its HackerBench v0.3, it allowed 3.3% of dangerous queries to pass, while in the LatchBio biosecurity test, it scored 62.4%.
It is worth noting that SpaceXAI released Grok 4.6 in August, which also featured a 500,000 token context window and a base rate of $2 for input tokens and $6 for output tokens.
