Highlights
- Google has introduced Gemini 3.7 Flash, a cost-effective model for coding and agents, now widely available.
- OpenAI has launched a preview of GPT-5.6 Sol Ultrafast, a high-speed version of its advanced model, operating up to 14 times faster.
- The focus has shifted from pure intelligence to the speed of real-time AI agents, but only Google's offering is accessible to all developers at present.
Both Google and OpenAI have delivered a clear message: AI technology has reached a level of speed that is impressive for users of AI-driven agents, with each company unveiling their respective high-speed models.
The two models differ in design, and only one is available for immediate use.
Myriad: When will OpenAI release GPT-6? Make your prediction.Google has launched Gemini 3.7 Flash, optimized for software development and autonomous business processes. Meanwhile, OpenAI is offering a limited preview of GPT-5.6 Sol Ultrafast, a new tier that enhances its most powerful model's speed to as much as 750 output tokens per second.
Today we unveil Gemini 3.7 Flash, our most advanced model yet for coding and agents.
This model significantly improves performance across software engineering, web development, and complex knowledge tasks.
Gemini 3.7 Flash is available now through the end of the year… pic.twitter.com/RSCBDipjKn
— Google (@Google) August 13, 2026
Gemini 3.7 Flash is now generally available. It can process up to a million input tokens (approximately 750,000 words) and generate 64,000 tokens, handling various data types including text, images, video, audio, and PDFs. It is designed to serve as a budget-friendly brain for automated systems that manage tasks and execute complex jobs with less human intervention.
This model does not compromise on quality for speed; it is both more capable and efficient, completing a test coding task in just 2 minutes and 13 seconds compared to the previous Flash model's over 5 minutes. The enhancement in quality is also evident.
OpenAI's Ultrafast is not a new model but an accelerated version of GPT-5.6 Sol, which was fortified against prompt-injection vulnerabilities with the help of an AI red team before its launch. It operates approximately 14 times quicker than the standard speed of GPT-5.6 Sol.
Previewing Ultrafast mode: GPT-5.6 Sol running at up to 14x the speed.
Initially available in the OpenAI API to a select group of clients, more businesses will gain access as capacity increases. pic.twitter.com/a5dleofiDJ
— OpenAI (@OpenAI) August 13, 2026
The Cerebras chips enable the generation of up to 750 tokens per second, roughly 560 words, providing rapid processing that allows a voice agent to think during calls.
Key Metrics
According to Google's benchmark data, Gemini 3.7 Flash outperforms Claude Sonnet 5, GPT-5.6 Terra, and others in 11 out of 18 tested categories, achieving a top score of 1,588 Elo in Code Arena web development and 30.4% on AutomationBench for enterprise tasks. These results rely on Google's own evaluation methods, so they should be considered claims.
Gemini 3.7 Flash is also cost-effective, priced at 75 cents per million input tokens and $3.75 per million output tokens until the end of the year. This is half the initial rate of Gemini 3.6 Flash. After December 31, prices will increase to $1.50 and $7.50, which remains affordable for a Google model.
OpenAI has not released comparative scores for Ultrafast beyond customer testimonials. John Crepezzi, an AI engineer at Jane Street, noted that Cerebras’ speed "enables different ways of using the models." Courtland Lykins, product lead at Podium, described it as "invaluable in our voice stack," asserting that the speed "completely transforms the call experience." Access to this tier is currently by invitation only.
The emphasis on speed comes as both companies shift their focus from who is the smartest to who can deliver the fastest solutions for AI agents. Google's timing is significant, as its flagship Gemini 3.5 Pro is still pending with no release date announced, following a leadership change at DeepMind that saw Demis Hassabis replaced by deputy Koray Kavukcuoglu. OpenAI, on the other hand, is leveraging Cerebras' speed rather than waiting for its own technology stack to be ready.
Gemini 3.7 Flash is currently available in over 160 countries, while GPT-5.6 Sol Ultrafast remains invite-only.
