
Just now, Google is back!
Early this morning, Google officially released Gemini 3.8 Flash, as well as Gemini 3.8 Flash Cyber, which focuses on network security.
This is the third Flash version released by Google in just six weeks.
The previous two updates seemed like just a warm-up, because this time, Google has pushed cost-effectiveness to the forefront—a single task cost of $0.58! Input cost of only $0.75 per million tokens! This incredibly affordable model has managed to reach the levels of the world's most powerful flagship processors, Opus 5, GPT-5.6 Sol, and Grok 4.6, in multiple core benchmark tests. Google DeepMind's Yao Shunyu commented: "For the model, this is just a small step; but for RSI, it's a huge leap." On X, the developers are in an uproar. After testing it, someone asked the existential question: "My God, did Google send the wrong product? Isn't this just the Gemini 3.5 Pro that they haven't released yet? It's charging the price of Flash but doing the work of a Pro!" "Someone on the DeepMind team probably hasn't slept since July." Everyone is still digesting Fable On May 1st, Google once again turned the tables. Google, the "King of Price/Performance Ratio," is back: the king of value for money. After a long period of silence, Google, in an extremely competitive manner, announced to the world: the big boss is back. To understand the shock brought by Gemini 3.8 Flash, we must first look at the current landscape of the AI market. Originally, the Artificial Analysis test already had extremely strict barriers. To achieve top-tier intelligence (Intelligence Index around 60 points), one had to pay a high price in tokens. However, Gemini 3.8 Flash acted like an assassin, completely disregarding ethical standards. In the high-reasoning mode of Artificial Analysis, Gemini 3.8 Flash's Intelligence Index soared to 59 points! What does this mean? It's just one step away from the top-tier GPT-5.6 Sol and Grok 4.6, and even surpasses Claude Opus 5 in some dimensions! And what's the cost of all this? Its cost per task is only $0.58. Specifically, the input cost remains at $0.75 per million tokens, and the output cost is $3.75 per million tokens. Google created a chart: on the Pareto frontier of intelligence versus cost, the blue dot representing "most efficient" for Gemini 3.8 Flash is located in the upper right corner of the software engineering benchmark DeepSWE, leaving the second-place device far behind. Gemini 3.8 Flash is not only intelligent, but also ridiculously fast. Its generation speed reaches an astonishing 305 tokens/s, while the next-fastest model barely manages 154 tokens/s. "Fast, smart, and so cheap it's practically free—this is the kind of model humans can truly use every day." Especially on the long-cycle software engineering benchmark (DeepSWE v1.1), 3.8 Flash achieved an astonishing 73.7% success rate. In this test, which required AI to autonomously solve complex engineering problems and write end-to-end code, it not only surpassed most expensive cutting-edge large models, but also cost only a fraction of the latter. It's important to note that this is only the "Flash" version, not "Pro," much less "Ultra." This time, Google has once again allowed a smaller model to steal the spotlight from its flagship model. Someone personally tested the Gemini 3.8 Flash and Opus 5 performing the same task.


In addition, Google also poached Barret Zoph—co-founder of Thinking Machines Lab, former head of post-training at OpenAI, now serving as VP of Research, overseeing reinforcement learning and post-training. This move clearly shows they're going all out.
Last month, co-founder and Nobel laureate Demis Hassabis stepped down as CEO, and his successor, Koray Kavukcuoglu, immediately declared: Speed up, speed up, and speed up again.
Is the benchmark inflated, or is it simply superior performance? However, amidst the praise, Google is generally criticized for prematurely celebrating. The most frustrating thing is Meta—Meta's Muse Spark 1.3 and Gemini 3.8 Flash went head-to-head, with the former scoring 62 points on the Artificial Analysis Intelligence Index, matching Claude Fable 5 and placing it among the top performers. Muse's input cost is 8 times lower than Fable 5's, and its output cost is nearly 12 times lower, making it extremely cost-effective. Meta's Chief AI Officer, Alexandr Wang, bluntly stated: "In the Artificial Analysis intelligence index, Gemini is nothing; it can only rely on the exhaust fumes of other models."


Muse Spark 1.3 scores higher than GPT-5.6 Sol and Grok 4.6 and Gemini 3.8 Flash are available, but currently only in limited preview. Some have pointed out that while the benchmark charts for 3.8 Flash look good, their performance in real-world testing may be different. Indeed, Gemini 3.8 Flash outperformed GPT-5.6 Sol and Opus 5 in tests like Terminal-Bench 2.1 and HLE, but the huge score difference between TBench 2.1 and TBench 4 suggests possible "ranking manipulation." In other words, when faced with unfamiliar real-world Zero-shot challenges not yet included in the training set, 3.8 Flash will likely still be inferior to a parameter behemoth like Opus 5. But Google's solution is incredibly clever—hard work makes up for shortcomings. As Google stated in its official blog: "These performance improvements stem from core design choices: 3.8 Flash works harder." When faced with complex tasks, 3.8 Flash is designed to perform additional inference steps and iteratively call tools. When encountering problems it doesn't understand, it doesn't simply give up; instead, it consumes more tokens to perform rapid self-verification and reflection internally. Even though it consumes more tokens through multiple iterations, its base unit price is so cheap ($0.75/million tokens) that, overall, it's still much cheaper than directly calling GPT-5.6 once! This strategy directly breaks the past single-dimensional competition of "large parameters = strong capabilities". It proves that in a perfect agent cycle, speed and extremely low inference cost are themselves a powerful form of intelligence. Why release Flash? The "rejected" Pro version. Did Google really not make a mistake this time? Gemini 3.5 Pro is nowhere to be seen. The next-generation ace, Gemini 4, has impressive pre-training data, but the post-training phase is not yet complete.

In fact, Google's delay in releasing the Pro version has a painful history behind it.
Spicy boasted in May of this year that a more powerful Pro series would be launched "next month," but it has been delayed ever since.
Spicy had boasted that a more powerful Pro series would be launched "next month" this year, but it has been delayed ever since.
Actually, Google originally had several candidate models for 3.5 Pro, but they were all ruthlessly rejected—because they weren't significantly better than the current Flash series! Meanwhile, the highly anticipated next-generation flagship Gemini 4, while performing well in pre-training evaluations, is currently stuck in the crucial post-training stage. In short, everything has inadvertently led to the frenzied iterations of the Flash series.
A vulnerability that would take months to discover was uncovered in 2 hours: The cybersecurity version topped the charts
Is Google just letting 3.8 Flash drive prices for routine tasks this time?
Far more than that.
The real killer weapon of this launch event is its twin brother—Gemini 3.8 Flash Cyber.
Even developers are exclaiming: "What kind of security task is worth Google releasing a Cyber-specific variant for Flash?"
Google's answer is: to combat future AI hackers.
On CyberGym, a benchmark test for vulnerability discovery, Flash Cyber 3.8 achieved a score of 86.2%, dominating the leaderboard and leaving previous large models far behind. Moreover, real-world code goes beyond C/C++. In comprehensive testing within Google's internal database of complex codebases spanning 20 programming languages, this model achieved a success rate of over 70% in discovering a wide range of vulnerabilities. Even more impressive is its real-world performance. The Chrome security team tested it and found that it generated 2.6 times more correct patches than the previous best-performing, much larger commercial model. Wiz Security's tests found that it improved recall rates in penetration testing by 7.5-9.7%, while reducing costs by 2.3 to 5.2 times. Most surprisingly, the Google Cloud vulnerability research team used it to discover a critical fundamental vulnerability in just 2 hours—something that previously took top human experts months to find! In the CWE-Bench (Patching Capability Test), Flash Cyber 3.8 achieved a first-pass rate of 47.2%, almost on par with the most advanced frontier models (47.8%), but at a fraction of the price. Because of Flash Cyber's formidable network attack and defense capabilities, Google did not choose to fully open-source it, but instead opened it only to trusted organizations through the new Fairwind program. OpenAI, Anthropic, and Google: The Three Kingdoms Battle of Large Models Reaches a Watershed Moment. The release of Gemini 3.8 Flash reveals that the three giants of AI have embarked on three distinct evolutionary paths. Anthropic (Claude-affiliated) is focused on the reliability and stability of knowledge. In tests that prioritize knowledge coverage and factual accuracy (such as AA-Omniscience), Claude remains the clear winner. The top seven are all from the Claude group, and they are now clearly focusing on improving their error-free capabilities in enterprise-level tasks, striving for consistent output. OpenAI (GPT group) is betting on "deep reasoning and insight." In the CritPt test, which leans towards physics and mathematical derivation, GPT-5.6 Sol leads by a significant margin. OpenAI seems to be pursuing a pure intellectual emergence, requiring models to "think" out answers like scientists without being given them them. Google (Gemini-affiliated) is creating an "endless cycle of ultra-high-speed laborers." Google has offered a third answer: Perhaps I can't solve the most difficult physics problem on my first try, but my reaction time is extremely fast and cost-effective. In the era of Agents, I could think 100 times per second, write code, run it, report errors, and modify it, using extremely low-cost brute-force iteration to force the correct result. The problems Agents encounter in reality require repeated trial and error, calling tools, and dynamic adjustments. And Gemini 3.8 Flash is practically tailor-made for this kind of "long-term workflow." In Google Antigravity, you can use just a hint and a loop command to have Gemini 3.8 Flash automatically generate a complete game with puzzles, environment maps, and 3D levels. You can also use it to generate a DOS version of Google Maps with Street View and navigation with a single click. Clearly, Gemini 3.8 Flash's ease of use and cost have reached a new level. The arrival of Gemini 3.8 Flash seems to be Google's declaration to the entire industry: large-scale computing is moving away from being "only affordable for giants" and is rapidly becoming more accessible to everyone. Those tasks that used to require tens of thousands of dollars and days of running for intelligent agents can now be completed in minutes for the price of a few cups of coffee. The era of the super-individual, belonging to ordinary people, has truly arrived. This is the awakening moment for small models.