Google introduces Gemini 3.7 flash with huge upgrades

Gemini 3.7 Flash can process text, images, audio and video.

Google has officially introduced Gemini 3.7 Flash on August 13, 2026. The new multimodal model is designed as a fast and efficient “workhorse” for coding, Web development, knowledge-based tasks and agent workflows. The release came only three weeks after Gemini 3.6 Flash. Google did not train the new model from scratch. Instead, developers used algorithmic improvements and user feedback to completely replace the previous version.

At the same time, Google has still not released its flagship Gemini 3.5 Pro model, which the company had originally promised for June 2026. Industry analysts believe the quick release of several Flash updates could be linked to an internal loss of AI talent. Reports also suggest Google is trying to catch up with rival AI companies after Gemini’s earlier coding performance fell behind competitors.

Gemini 3.7 Flash can process text, images, audio and video. It can generate text with a large 1 million-token context window, while its maximum output is 65,536 tokens. The model also allows users to adjust its reasoning level. Users can choose LOW, MEDIUM or HIGH thinking levels, with MEDIUM set as the default. However, the model does not support a MINIMAL thinking level and will return an API validation error if users request it.

The model is designed to handle complex, multistep tasks. It can respond to problems during a task, adjust its approach and ask for clarification when needed. This allows it to complete complicated plans more effectively and reduces the need for users to retry tasks manually.

Gemini 3.7 Flash also brings major improvements in software engineering. Developers can expect more accurate code on the first attempt and faster debugging. For Web development, the model can create applications with complete features while requiring fewer prompts.

However, independent analysts warn users not to rely completely on benchmarks reported by vendors. High benchmark scores do not always mean the model will be reliable in real-world tasks such as code reviews or permission management. Engineering teams should therefore continue testing the model on workloads that match their specific needs.