Google has released two new versions of its Gemini Flash model series in a single announcement, doubling down on both agentic AI performance and cybersecurity specialisation simultaneously. Gemini 3.8 Flash is positioned as a hardworking workhorse for software development, multi-step reasoning and agentic tasks, while its companion Gemini 3.8 Flash Cyber is Google's most capable cybersecurity model to date — achieving frontier-level performance on vulnerability discovery and patching across more than 20 programming languages.
Key Points
- Google released two variants of Gemini 3.8 Flash — a standard agentic model and Flash Cyber, a dedicated cybersecurity model
- Flash Cyber achieved 86.2% on the CyberGym cybersecurity benchmark and 47.2% on CWE-Bench, which evaluates AI patching capabilities
- In an internal Google benchmark, Flash Cyber achieved more than 70% success rate discovering vulnerabilities across 20 programming languages
- 3.8 Flash outperformed many large frontier models on the DeepSWE coding benchmark at significantly lower cost
- 3.8 Flash ranked No. 14 in Agent Arena — above DeepSeek-V4-Pro — and No. 7 in Text Arena, ahead of Claude Opus 5
- 3.8 Flash is priced at $0.75 per million input tokens and $3.75 per million output tokens — the same as Gemini 3.7 Flash
- This is Google's third Flash release in six weeks
The Standard Model — Built for Agents
Gemini 3.8 Flash is designed around the demands of agentic workloads — tasks that require a model to plan, reason across multiple steps, use tools and execute complex sequences of actions with minimal human intervention. Google CEO Sundar Pichai described the model as delivering significant leaps from 3.7 Flash across software engineering, agentic tasks and multi-step reasoning.
The model is available now in Gemini Enterprise and through the Gemini API via Google AI Studio, Google Antigravity, Android Studio and Stitch. Pricing matches Gemini 3.7 Flash — $0.75 per million input tokens and $3.75 per million output tokens — with users able to adjust effort levels based on their priorities around quality, cost and latency.
Google senior product director Tulsee Doshi and Gemini security lead Raluca Ada Popa described 3.8 Flash as a model that "works harder" with "greater diligence" on complex tasks — executing additional reasoning steps where needed, though at times using more tokens to maximise performance. The model features a one million token input window and a 64,000 token output limit, and can process text, images, audio, video and PDF files.
The practical capability demonstrations Google shared illustrate the model's range. Using a simple prompt on the Antigravity platform, 3.8 Flash built a functioning 3D game — a wizard navigating a castle — using looping techniques, procedural storytelling and textures from Nano Banana. In other examples, the model created a fully functional DOS version of Google Maps with interactive locations, directions and street views; a 3D visualiser that decomposed devices into inspectable layers; and a topographic map of famous geographical sites based on real U.S. Geological Survey datasets, with real-time cross-sections, 2D projections and scientific explanations.
On Arena.ai's independent evaluation, 3.8 Flash landed at No. 14 in Agent Arena — above DeepSeek-V4-Pro and significantly ahead of Gemini 3.7 Flash, which sits at No. 32. It debuted at No. 7 in Text Arena, placing ahead of Anthropic's Claude Opus 5 and Gemini 3.7 Flash. Improvements over its predecessor were recorded across multi-turn requests, writing, literature, language, longer queries, hard prompts, coding, instruction following and business and financial operations.
The model also performed strongly on specialist knowledge benchmarks. It outperformed predecessors and other frontier models on Vals Finance Agent V2 for financial analysis and Harvey's Legal Agent Benchmark for legal reasoning. On Humanity's Last Exam Verified — one of the most demanding AI capability tests, covering multi-step reasoning across mathematics, science and humanities — 3.8 Flash scored 54.9%.
Flash Cyber — A Dedicated Vulnerability Hunter
Gemini 3.8 Flash Cyber is a purpose-built cybersecurity model that Google describes as its most capable to date for vulnerability detection and mitigation. Pichai said it matches frontier-level performance when it comes to discovering vulnerabilities and patching them at scale — a claim backed by specific benchmark results.
Flash Cyber achieved 86.2% on CyberGym — the standardised cybersecurity benchmark that has become an industry reference point for comparing AI cybersecurity models — and 47.2% on CWE-Bench, which specifically evaluates AI systems' ability to patch software weaknesses. In an internal Google benchmark assessing vulnerability discovery across 20 programming languages, the model achieved a success rate exceeding 70%.
The breadth of the language coverage is particularly significant. Cybersecurity vulnerabilities do not confine themselves to popular languages — legacy systems running in critical infrastructure often use older or less common languages that receive less security research attention. A model that can hunt vulnerabilities across 20 languages is covering a substantially larger portion of the actual software ecosystem than one that performs well only on the most common languages.
The CyberGym score of 86.2% places Flash Cyber in competitive territory with other leading cybersecurity AI models. Microsoft's MAI-Cyber-1-Flash and OpenAI's GPT-5.6 Sol have both claimed benchmark leadership in the same space, while Anthropic's Mythos 5 — accessible through Project Glasswing — has set a high bar on CyberGym. Google's entry into the dedicated cybersecurity model category with benchmark numbers of this calibre signals that the company is treating AI-powered security as a strategic priority rather than a feature of its general-purpose models.
Three Flash Releases in Six Weeks
The pace of Google's Flash releases is itself a signal about the competitive dynamics in the AI model market. Gemini 3.8 Flash is the third Flash release in six weeks — a cadence that reflects both the technical feasibility of rapid iteration and the commercial pressure to maintain visibility and relevance in a market where competitors including OpenAI, Anthropic and DeepSeek are releasing capable models at high frequency.
Google's decision to maintain the same pricing as 3.7 Flash despite the performance improvements is consistent with a strategy of competing on value — offering more capability at the same price point rather than charging a premium for advances. That approach is particularly effective against the cost efficiency narrative that DeepSeek has used to gain traction in the developer market, and against Microsoft's Teams and Azure enterprise relationships that create switching cost advantages for competing products.
For users who prioritise efficiency over maximum performance, Google noted that 3.7 Flash remains fully supported for efficiency-first workloads — giving the product line a clear tiering rather than requiring users to move to the newest model regardless of their specific needs.
What This Means for the Cybersecurity AI Race
The release of Flash Cyber as a dedicated model — rather than simply noting that 3.8 Flash performs well on security tasks — reflects a broader industry pattern. OpenAI has GPT-5.6 Sol and the forthcoming Astra for cybersecurity. Microsoft has MAI-Cyber-1-Flash. Anthropic has Mythos 5 through Glasswing. Google now has Flash Cyber. The market for purpose-built AI cybersecurity models is consolidating around a small number of major players, each with specific benchmark claims and access structures.
The practical competition for enterprise cybersecurity buyers will increasingly come down to deployment flexibility, integration with existing security infrastructure, access structure and cost — in addition to raw benchmark performance. Flash Cyber's availability through Google's existing API and developer platform infrastructure gives it an immediate deployment path that models available only through restricted access programmes cannot match for most enterprise buyers.
Sources
Google blog post on Gemini 3.8 Flash and 3.8 Flash Cyber, by Tulsee Doshi and Raluca Ada Popa, Wednesday September 3, 2026. Sundar Pichai post on X regarding Gemini 3.8 Flash capabilities, September 3, 2026. Arena.ai Agent Arena and Text Arena rankings for Gemini 3.8 Flash, September 3, 2026. CyberGym benchmark score 86.2% and CWE-Bench score 47.2%, per Google announcement. Vals Finance Agent V2 and Harvey's Legal Agent Benchmark performance references, Google blog post, September 3, 2026.