Google Rolls Out Gemini 3.8 Flash Model with Faster Coding Performance and Same Intro Pricing

Google launched Gemini 3.8 Flash, its newest AI model aimed at coding and complex reasoning, just weeks after the 3.7 Flash release. The model is priced at $0.75 per million input tokens and $3.75 per million output tokens, but may consume more tokens due to deeper reasoning steps. Early reports suggest the rollout is part of a rapid update cycle to compete with OpenAI and Anthropic, and developers can still use the prior 3.7 Flash to control costs.

Google unveiled Gemini 3.8 Flash on September 2, 2026, positioning the model as its strongest offering for coding and reasoning tasks to date. The company says the new version "works harder" by performing additional reasoning steps and calling tools iteratively, while retaining the introductory pricing of $0.75 per million input tokens and $3.75 per million output tokens. However, Google warned that the model could end up costing users more because it may use additional tokens to maximize performance, especially at higher effort levels, and developers can revert to Gemini 3.7 Flash to limit token consumption .

The rollout follows a rapid update cadence, arriving just three weeks after the previous Flash release and marking the third Flash update in three months. This fast-track deployment underscores Google's intent to stay ahead in the AI coding race against rivals such as OpenAI and Anthropic, which have recently launched their own advanced models .

Google published specific benchmark gains rather than leaving the performance story vague: on the DeepSWE v1.1 long-horizon software engineering benchmark, Gemini 3.8 Flash outperforms most larger frontier models, and on Terminal-Bench 2.1, a measure of end-to-end command-line and coding task completion, it scores well ahead of the prior Gemini 3.7 Flash version. Google is also making a security-focused variant, Gemini 3.8 Flash Cyber, available to trusted testers including government and critical-infrastructure operators through its Fairwind Program, a narrower rollout than the general-purpose coding model .

Analysts will monitor adoption rates and actual token usage to gauge whether the higher computational effort translates into higher costs for users. The ability to switch back to the prior model may help cost-sensitive developers, but the perceived performance gains could drive broader uptake, influencing the competitive dynamics of the AI coding market.

Related Stocks

Powered by SentiSense - Intelligent Market Analysis