Overview
- Gemini 3.7 Flash successfully created a playable browser game from a single prompt in just 2 minutes and 13 seconds, a feat that Gemini 3.6 Flash failed to accomplish three weeks prior.
- It struggled with a bridge logic puzzle, providing the same incorrect answer as Claude Fable 5, and did not complete the math problem correctly.
- The cost of using the model is set at 75 cents per million input tokens until December 31, which is half the rate of 3.6 Flash, after which it will increase to $1.50.
On August 13, Google launched Gemini 3.7 Flash, making it available in over 160 countries from day one. The model can handle up to a million input tokens, process 64,000 outputs, and interpret images, videos, audio, and PDFs, as well as interact with tools and control a computer.
Traditionally, Flash is not the go-to model for complex problems; rather, it serves well for organizing text, managing agent sessions, and summarizing documents without incurring the cost of a premium model.
In terms of performance on simpler tasks, 3.7 Flash is indeed a significant improvement. However, compared to more advanced software, it still falls short in capabilities.
Google's own testing metrics indicate that 3.7 Flash outperforms Claude Sonnet 5 and GPT-5.6 Terra in 11 out of 18 assessment categories. Notably, it scored 1,588 Elo on Code Arena’s web development leaderboard and 30.4% on AutomationBench. These results stem from Google's proprietary methodology, so they should be considered as the company's claims rather than absolute facts.
We conducted tests to determine if the model meets Google's assertions. Here are our findings.
Coding: Can it create a functional product on the first attempt?
This evaluation assesses zero-shot code generation—whether the model can transform a single instruction into operational software without prior examples or the ability to self-correct. We provided one prompt for a browser game and accepted whatever output was generated, including any bugs. There were no follow-ups or error corrections allowed.
Gemini 3.7 Flash succeeded within 2 minutes and 13 seconds. The game was playable on the first attempt, with clean syntax and functioning collision and scoring mechanisms, exceeding the expectations set by its pricing tier.
In this context, the relevant comparison is not with flagship models but with Gemini 3.6 Flash, which was unable to produce a working file. Its HTML was flawed, elements failed to render, and requests for corrections went unanswered.
We ultimately had to send that output to DeepSeek, which identified 11 bugs and provided 8 fixes to make it functional. In contrast, just three weeks later, 3.7 Flash needed no assistance, producing results comparable to GPT-5.6 Sol from our previous review.
Gemini 3.7 Flash clearly outperforms its predecessor, making it the primary reason to consider a switch. However, it relies on executing specifications rather than generating creative ideas, so a vague prompt will yield a vague game.
You can try the game made with Gemini 3.7 Flash here.
Creative Writing: Can it maintain a paradox and craft a coherent sentence?
This test evaluates literary quality and the model's ability to adhere to structural constraints across an extended narrative. The prompt sends Jose Lanz from the year 2150 back to 1000 AD, requiring a closed causal loop—his actions must create the future he aims to prevent.
The critical rule is that he cannot comprehend his actions until he returns home.
Gemini 3.7 Flash produced a reasonably good result. Jose inadvertently creates an obelisk that enslaves 22nd-century Iberia by firing an entropic cannon into a fissure, ultimately realizing the loop while still in the past: "It was the base of the Cinder Spire."
The plot mechanics are quite solid. The narrative includes a falling star witnessed by ancient monks, which turns out to be the flash of Jose's arrival, and the weapon he brings to eliminate the anomaly is what ultimately creates it. The ending—"It had simply been waiting for him to complete it"—captures the determinism requested in the prompt.
However, for those familiar with AI-generated content, the writing style is immediately recognizable. Most nouns are accompanied by multiple adjectives: "hyper-luminescent towers," "damp, moss-choked earth," and "thick, obsidian hair." This indicates a model selecting the most probable next word rather than making distinct choices, leading to phrases like a monolith "humming with a low-frequency hum."
In comparison, we evaluated it against Qwopus3.5-27B-v3, a community-tuned version of Qwen3.5-27B that embodies Claude Opus-style reasoning and operates on a consumer GPU for free. Qwopus adhered to the rule that Gemini failed to follow.
In its version, Jose kills a monk at the historically significant San Millán de la Cogolla monastery, only comprehending his actions after returning to 2150 and discovering his DNA in a sealed codex.
Although Qwopus is not without flaws—it included its planning notes with typos above the story and its ending broke the closed loop by allowing Jose to correct his actions—it ultimately performed better. Gemini produced a more polished presentation and a more structured conclusion, but it failed the one critical instruction of the prompt, while a free model on a gaming GPU delivered a superior narrative.
Associative Thinking: Can a metaphor support an argument?
This test assesses associative reasoning—whether a model can make connections between unrelated ideas without needing to explain them. The prompt requests a description of a twig, which is then used to discuss worker exploitation and the veneration of wealth, concluding with a description of a lettuce.
Failing to signpost the metaphor is crucial; naming it undermines its effectiveness.
Gemini names the metaphor in the opening line of its second paragraph: "This is the precise mechanics of the modern proletariat," which detracts from its impact. The preceding imagery, however, is quite evocative. The worker is described as receiving "just enough bark to stay rigid for another week of output," and fallen twigs are conditioned to believe that with enough rigidity, any one of them might become a trunk. This paragraph also references a worker "bound to a vast, top-heavy corporate hierarchy."
Overall, the logic and structure are commendable.
However, the transition is where it falters. Gemini narrates the shift instead of executing it—the hierarchies "crumble, dissolving into the quiet, humble reality of the organic world underneath"—and then a lettuce appears with no connection to prior descriptions.
GPT-5.6 Sol effectively integrates the twig into soil and grows the lettuce from it: "Rain enters the grain. Fibers loosen, darken." The argument is embedded within the object itself, framing wealth as a language of virtue where "The mansion signifies intelligence."
GPT-5.6 Sol outperformed Gemini significantly. While Gemini produced some strong individual lines, it explained its metaphor and failed to execute the transition that the prompt specifically required.
Logic: Does it comprehend the prompt or recognize the puzzle?
This evaluation focuses on non-mathematical reasoning, specifically assessing whether a model reads the question accurately or merely applies memorized patterns. Our bridge prompt presents four individuals with one torch and crossing times of 1, 2, 5, and 10 minutes, asking how quickly they can all cross.
The key detail omitted from the prompt is that only two people can be on the bridge at once, meaning the correct answer is 10 minutes—all can walk over together at the pace of the slowest individual.
Gemini responded with 17 minutes, following a memorized sequence from the textbook version of the puzzle. It asserted the constraint as fact without verifying whether it was included in the prompt.
Its reasoning is less impressive than the answer it provided. The explanation suggests that sending the two slowest individuals across together would be inefficient because someone would need to return the torch—and yet it ultimately suggests they cross together anyway. This inconsistency occurs within a single response, which it presents with unwarranted confidence.
Claude Fable 5 provided the same incorrect answer back in July, beginning by stating its assumption: "assuming the classic constraint that the bridge holds only two people at a time," which differentiates a misleading answer that can be caught from one that cannot.
In conclusion, neither model excelled. Fable wins on transparency alone, while Gemini's overconfidence in its incorrect answers is evident in a puzzle that can be manually checked.
Math: Does it complete the task?
This evaluation assesses symbolic mathematics beyond typical consumer applications, as well as basic task completion—whether the model fulfills its request. The prompt requires a degree-19 odd monic polynomial with real coefficients and a linear coefficient of -19, whose curve splits into at least three irreducible components, and then asks for p(19).
Both models identified the same polynomial. Gemini and Qwen 3.7 Max Preview recognized the Dickson polynomial, solved for its parameters, and derived the closed form accurately.
However, Gemini stopped short, presenting p(19) as an unevaluated expression involving the 19th power of a square root, failing to provide the number or demonstrate the required component count. It delivered all of this in a styled HTML format with CSS and a drop shadow that was unnecessary.
Qwen completed the task, providing the full factorization into 10 components—one linear and nine quadratic—running the recurrence to 1,876,572,071,974,094,803,391,179, and verifying the result modularly. We confirmed this figure independently using SymPy, and it holds true.
Qwen wins based on the only criterion that truly mattered. Gemini started strong but inexplicably skipped the arithmetic, which is an odd place to halt.
Final Thoughts
Gemini 3.7 Flash is a worthwhile upgrade for those already within Google's ecosystem. It shows significant improvements in coding compared to its predecessor, operates quickly enough for agent tasks, and is economical enough to make high-volume usage negligible in cost.
Its strengths lie in execution and structure: provide it with detailed specifications, and it will produce the required output, maintain narrative coherence, and uphold causal logic over extended texts.
However, it falls short in creativity and reasoning. The writing is often predictable enough to be recognized as machine-generated, and the model confidently presents incorrect answers without acknowledging the assumptions that led to those errors.
Its pricing is its most compelling feature. At 75 cents per million input tokens and $3.75 for output, it undercuts GPT-5.6 Sol's $5 input rate by 85% and costs half of what 3.6 Flash did at launch.
On the flip side, there is a free 27B model on a gaming GPU that produced a better story without any cost. The introductory pricing from Google will expire on December 31, after which input costs will rise to $1.50 and output to $7.50.
