Summary

  • Google introduced the Gemini 3.6 Flash and 3.5 Flash-Lite models today, showcasing improvements in efficiency and cost compared to 3.5 Flash; however, the anticipated Gemini 3.5 Pro remains in testing due to internal performance issues.
  • The 3.6 Flash model utilizes 17% fewer output tokens than its predecessor, reducing the cost from $9 to $7.50 per million tokens, thereby lowering the operational costs for AI agents.
  • Google also announced the initiation of pre-training for Gemini 4, which it describes as its most ambitious effort to date.

Today, Google unveiled three new AI models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, which was not what many anticipated.

Following the introduction of Gemini 3.5 Flash at Google I/O 2026 in May and the promise of a Pro version within a month, Google failed to meet its own timeline. The delay of the Gemini 3.5 Pro is attributed to it not achieving internal coding targets, as reported by Bloomberg. An attempt to rectify this with updated training data in late June yielded unsatisfactory results. Following this news, Alphabet’s stock dropped approximately 4.4%, leading to a loss of about $200 billion in market capitalization within a single day.

The last Pro model released by Google was Gemini 3.1 Pro, which came out in February.

The Flash series represents Google's fast, cost-efficient models designed for AI agents that can perform tasks like document management and data processing autonomously. In contrast, Pro models are more powerful but also slower and more expensive, aimed at complex reasoning tasks.

Overview of Each AI Model and Their Use Cases

The primary release, Gemini 3.6 Flash, utilizes 17% fewer output tokens—where tokens are the basic units processed by AI, roughly equating to three-quarters of a word—compared to 3.5 Flash, according to the Artificial Analysis Index. It is also more affordable, costing $1.50 per million input tokens and $7.50 per million output tokens, a drop from $9 for the output tokens in 3.5 Flash. This price reduction is significant for businesses operating AI agents at scale.

In benchmark tests that measure AI performance, 3.6 Flash scored 49% on DeepSWE v1.1, which assesses long-term software engineering tasks, compared to 37% for 3.5 Flash. On the MLE-Bench, it achieved a score of 63.9%, up from 49.7%. It excelled in OSWorld-Verified, where it achieved 83.0%, surpassing Claude Sonnet 5 at 81.2% and GPT-5.6 Luna at 72.6%.

However, competitors still outperform in certain areas: GPT-5.6 Luna scored 67% on DeepSWE and 84.7% on Terminal-Bench 2.1, which evaluates coding capabilities, while Claude Sonnet 5 leads in knowledge work on the GDPval-AA v2 benchmark with a score of 1607, compared to 3.6 Flash's 1421.

In our testing of the model for coding tasks, the results were disappointing. Our simple coding test produced an unusable file with improperly formatted HTML and incorrectly rendered elements. Efforts to correct the issues through further coding attempts were unsuccessful.

We enlisted Deepseek to troubleshoot the initial model and fix the bugs. It identified 11 bugs and implemented 8 critical fixes, eventually resulting in a playable game.

The adjustments made by Deepseek improved the game, indicating that while Gemini's foundational reasoning was sound, the inaccuracies rendered the initial output ineffective. Users should be prepared for lengthy coding sessions with this model if they intend to use it for such tasks.

The second model, Gemini 3.5 Flash-Lite, is designed for high-volume tasks, processing 350 output tokens per second at a cost of $0.30 per million input tokens and $2.50 per million output tokens. This model is tailored for high-throughput applications, such as massive document processing or agentic search systems, and it outperforms the older 3 Flash in key coding benchmarks like Terminal-Bench 2.1 (54% versus 31%), despite its lower cost.

It may also serve as an effective session compactor, analyzing long sessions and extracting key elements to prevent agents from becoming overwhelmed by excessive information, particularly for users of Hermes and Openclaw.

The third model, Gemini 3.5 Flash Cyber, will not be available to the public. Google is limiting access to this model to governments and trusted partners needing to identify and address software vulnerabilities, a capability that the company is hesitant to release widely.

In the meantime, Google’s DeepMind team has already begun work on future developments. The company confirmed in the official announcement that they have initiated "our most ambitious pre-training run yet, for Gemini 4," generating excitement about the project’s progress.

Pre-training is a crucial stage where the model learns from extensive datasets before moving on to specific task fine-tuning, indicating that Gemini 4 is actively in development rather than just in the planning phase.

We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress : )

— Logan Kilpatrick (@OfficialLoganK) July 21, 2026

Both the 3.6 Flash and 3.5 Flash-Lite models are currently operational within the Gemini app, Google AI Studio, and accessible via the API. Google stated that the Gemini 3.5 Pro will be released "as soon as it's ready," although a specific timeline remains unclear.

Daily Debrief Newsletter

Stay updated with the latest top news stories along with original features, podcasts, videos, and more.