Gemini 3.8 Flash Is Cheaper on Paper, Costlier per Task
Gemini 3.8 Flash keeps Google's low introductory token price but uses more reasoning and output tokens. Independent testing shows why its real cost per completed task can be significantly higher.
On this page
Google launched Gemini 3.8 Flash on September 2 with the same introductory token price as its predecessor, but independent testing points to a catch: getting the new model to complete harder work can cost substantially more. Artificial Analysis found that Gemini 3.8 Flash improved its intelligence score while using more output tokens and taking more agentic turns, pushing its average cost per task about 40% above Gemini 3.7 Flash.
Gemini 3.8 Flash keeps the price, but changes how it spends tokens
Google describes Gemini 3.8 Flash as its most capable Flash model for long-running software engineering, autonomous agents and complex enterprise workflows. It accepts text, images, video, speech and other supported inputs, and has a 1-million-token context window, meaning it can work with an unusually large amount of information in a single conversation. The model also has three thinking levels β low, medium and high β allowing developers to trade reasoning depth against speed and token consumption.
The API price initially looks familiar: $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. That is the same introductory rate Google lists for Gemini 3.7 Flash. But token price is only part of the bill when a model is allowed to reason repeatedly, call tools and revise its work. Google explicitly says 3.8 Flash may use more tokens on difficult tasks, particularly at higher reasoning levels.
The independent numbers show why the distinction matters
Artificial Analysis independently evaluated Gemini 3.8 Flash and gave its high-reasoning configuration an Intelligence Index score of 59, three points above Gemini 3.7 Flash at the same setting. The improvement was driven largely by agentic evaluations, which test whether a model can use tools and complete multi-step work rather than simply answer a question. On its testing, Gemini 3.8 Flash also reached roughly 300 output tokens per second, so the extra work does not mean the model has suddenly become slow at producing text.
The more interesting result is cost per completed task. Artificial Analysis measured Gemini 3.8 Flash at about $0.58 per task at high reasoning, compared with roughly $0.40 for Gemini 3.7 Flash. That is an increase of around 40% even though Google's listed price per token had not changed. The reason was straightforward: the newer model generated about 30% more output tokens on the evaluated tasks and used more turns in agentic workflows.
That changes how developers should read a price comparison. If an application simply sends short prompts and receives short answers, the difference may be small. If an agent spends several minutes inspecting files, calling tools, writing code and checking its own work, the amount of computation required to finish the job becomes much more important than the headline token rate.
Google is buying better results with more reasoning
Gemini 3.8 Flash is designed around a simple trade-off: spend additional computation when a problem is difficult instead of treating every request as equally cheap. Google says the model can execute additional reasoning steps and iteratively call tools, which is particularly useful for software engineering and other tasks where the first answer is not necessarily the final answer.
That approach helps explain why the model can compete with larger systems on selected benchmarks without being priced like a large frontier model. Google's own evaluations show strong results on long-horizon software engineering and specialized agent tasks, while Artificial Analysis independently found meaningful gains over 3.7 Flash. Neither set of results means the model is universally better, however. Benchmarks measure particular abilities under particular conditions, and real applications can behave very differently.
The three reasoning levels give developers a useful escape hatch
Developers do not have to run every request at maximum reasoning. Gemini 3.8 Flash supports low, medium and high thinking effort, with medium as the default. Artificial Analysis measured an Intelligence Index score of 57 at medium reasoning and 52 at low reasoning, while its cost-per-task measurements fell as the reasoning level decreased.
That makes the model more interesting as an engineering component than as a single benchmark score. A developer could use low reasoning for routine classification or straightforward generation, medium for normal coding work, and high only when an agent has to solve a difficult problem autonomously. The right setting depends on whether an extra few points of benchmark performance are worth additional latency and token consumption.
The bigger price increase arrives in 2027
There is another cost change hiding behind the introductory rate. Google's current pricing documentation says Gemini 3.8 Flash will move to $1.50 per million input tokens and $7.50 per million output tokens starting January 1, 2027. Both rates are exactly double the introductory prices available through the end of 2026.
For developers building production systems, that date matters more than the launch-day price. A workload that already costs more per task because 3.8 Flash reasons longer will face a higher underlying token rate once the introductory period ends. Teams comparing 3.7 and 3.8 should therefore measure the complete cost of successfully finishing their own tasks rather than multiplying the advertised input and output prices by a theoretical workload.
Gemini 3.8 Flash Cyber pushes the same idea into security
Google also introduced Gemini 3.8 Flash Cyber, a version aimed specifically at cybersecurity work. It is not a generally available model: access is being provided to trusted defenders through Google's Fairwind Program. The model focuses on finding software vulnerabilities and producing fixes, rather than primarily optimizing for offensive exploitation.
Google reports that Flash Cyber scored 86.2% pass@1 on CyberGym for autonomous vulnerability discovery and 47.2% on the external CWE-Bench automated-patching benchmark. Google also says its Chrome security team found that the model produced 2.6 times more correct vulnerability patches than the best commercial models it compared against. Those company-reported results are useful signals, but they should not be treated as independent confirmation of performance across the broader cybersecurity market.
The real upgrade is not simply a smarter Flash model
Gemini 3.8 Flash shows where AI model competition is moving: the useful unit is increasingly the completed job rather than the response. A model that thinks longer, calls tools and checks its own work can produce a better result while also consuming more compute, so a lower token price does not automatically translate into a lower operating cost.
For users, that distinction is mostly invisible. For developers running thousands or millions of agent tasks, it can determine whether an upgrade is economical. Gemini 3.8 Flash is therefore worth testing against real workloads, especially coding and multi-step tasks, but the sensible comparison is not β3.8 versus 3.7 per million tokens.β It is how much each model costs to get the job done correctly β and whether the extra reasoning saves enough human work to justify the bill.
Written by


