On August 13, Google introduced Gemini 3.7 Flash, while OpenAI launched the Ultrafast mode for GPT-5.6 Sol. Both innovations aim to enhance the rapid execution of tasks involving coding and AI agents.
What's New in Gemini 3.7 Flash
Google promotes Gemini 3.7 Flash as a model designed for programming and automating business processes. The company announced improvements in multi-step planning, instruction following, and code generation. Notably, the cost of usage has been halved compared to the initial pricing of its predecessor.
Source: Google DeepMind.The Flash line is tailored for scenarios where both response quality and speed are crucial, such as coding, AI agent operations, and tasks requiring multiple sequential interactions with the model.
In July, Google released Gemini 3.6 Flash, priced at $1.5 per million input tokens and $7.5 per million output tokens. According to the company, this model utilized 17% fewer output tokens than the previous 3.5 Flash.
Meanwhile, the anticipated Gemini 3.5 Pro has yet to be released. Developers are currently testing the flagship model with partners, although the release timeline remains uncertain.
How OpenAI Accelerated GPT-5.6 Sol
OpenAI has launched Ultrafast, a new mode for GPT-5.6 Sol. The company claims it can operate up to 14 times faster than the standard version, generating up to 750 output tokens per second. This acceleration is achieved through the Cerebras infrastructure.
Initially, Ultrafast is available via API to a limited number of clients. As OpenAI's computing capabilities expand, the company plans to onboard more users.
OpenAI had previously mentioned the 750 tokens per second figure during the announcement of GPT-5.6 in June. At that time, developers discussed plans to run Sol on Cerebras but cautioned that not everyone would have immediate access to the accelerated version.
Additionally, in August, the SpaceXAI team unveiled Grok 4.6, a new flagship model focused on long-term agent tasks, programming, and the creation of interactive applications.
