AGI Soon As Possible · Deep reads on AI & tech
Article

Gemini 3.7 Flash Arrived Three Weeks After 3.6 Flash: Same Price Sheet, and AutomationBench Went From 17.0% to 30.4%

2026-08-15 · 8 min read

Google released Gemini 3.7 Flash on August 13, 2026 and called it "our most intelligent workhorse model yet for coding and agents." All five benchmarks in the announcement are ahead of 3.6 Flash: FrontierCode 1.1 Main rises from 34.4% to 43.6%, DeepSWE v1.1 from 49.0% to 65.3%, and AutomationBench, which measures enterprise workflow automation, from 17.0% to 30.4%. The pricing document, however, lists exactly the same rates as gemini-3.6-flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, and those rates are valid only through December 31, 2026 before doubling on January 1, 2027. ASAP works only from figures verifiable in Google's own announcement and API pricing documentation to separate what this release changed from what it deferred by four months.

The biggest gains are in document comprehension and workflow automation, not coding

Gemini 3.7 Flash is ahead of 3.6 Flash on all five benchmarks in Google's announcement, and the average gain across the four percentage-scored items is 12.7 percentage points. FrontierCode 1.1 Main rises 9.2 points from 34.4% to 43.6%, DeepSWE v1.1 rises 16.3 points from 49.0% to 65.3%, the GDP.pdf document comprehension benchmark rises 12.0 points from 22.0% to 34.0%, and AutomationBench rises 13.4 points from 17.0% to 30.4%. WebDev Arena, scored in Elo, moves from 1538 to 1588.

Ranking the same table by multiple rather than by point gain reverses the order. AutomationBench leads at 1.79x and GDP.pdf follows at 1.55x, while the coding items trail at 1.33x for DeepSWE v1.1 and 1.27x for FrontierCode 1.1 Main. Google introduced this model for coding and agents, yet the two items that moved most measure reading documents and completing clerical procedures rather than writing code.

The 50-point Elo gain deserves its own translation. In an Elo system a 50-point gap means the higher-rated side wins roughly 57% of head-to-head matchups. Unlike a 9.2-point or 16.3-point move on a percentage metric, 50 Elo means that a human comparing two outputs side by side picks the new model about six times out of ten. Numbers in a benchmark table do not share a unit, and that has to be re-checked every time the table is quoted.

The pricing document applies identical rates to 3.7 Flash and 3.6 Flash

Google's Gemini API pricing documentation is explicit that gemini-3.7-flash and gemini-3.6-flash carry the same rates of $0.75 per 1M input tokens and $3.75 per 1M output tokens. Output pricing includes thinking tokens. Context cache reads cost $0.075 per 1M tokens and cache storage costs $0.50 per 1M tokens per hour. Batch and Flex modes run at half price, which puts input at $0.375 and output at $1.875. On the free tier, input, output, and caching are all free.

Google Search grounding includes 5,000 free requests per month, and that free allowance is shared across all Gemini 3.x models. Beyond it, the rate is $14 per 1,000 requests, and Google Maps queries carry the same rate. For teams running agents with search attached, this per-request charge often shows up on the invoice before token costs do.

The verifiable fact here is simple. Google raised capability in this release and left the price sheet untouched. The same work costs the same money at a higher completion rate, so cost per unit of successful output falls by exactly the margin shown in the table. On AutomationBench, the share of tasks completed for the same 1M tokens goes from 17.0% to 30.4%.

The doubling scheduled for January 1, 2027 is the real variable in any adoption model

Current rates for Gemini 3.7 Flash are valid only through December 31, 2026, and Google's API pricing documentation states that input becomes $1.50, output becomes $7.50, cache reads become $0.15, and cache storage becomes $1.00 per hour on January 1, 2027. All four items double exactly. The relationship between introductory and standard pricing is a uniform multiple rather than a per-item adjustment, and that uniformity defines the character of this price sheet.

A uniform increase leaves no room for optimization. Cache reads stay at one tenth of standard input after the change, and batch stays at half. Raising cache-hit share or shifting traffic to batch does not offset the increase, because every denominator moves with every numerator. Under a price revision with item-by-item differences, teams can restructure to absorb the shock; under a flat 2x, the only available responses are to accept double the invoice or to cut call volume outright.

Converted into money, the judgment sharpens. A service consuming 10M input tokens and 2M output tokens per day pays $15 a day and $450 a month at current rates. The same traffic costs $900 a month starting January 1, 2027. Teams running pilots today and drafting budgets from them need approval against the figure four months out, and any break-even computed on introductory pricing has to be re-tested before year end. The stated expiry date is the single most operational fact in this release.

The item with the largest multiple still sits at 30.4%

AutomationBench rose 1.79x from 17.0% to 30.4%, and its absolute level is still three attempts in ten. In the same table GDP.pdf sits at 34.0% and FrontierCode 1.1 Main at 43.6%, and the only item above half is DeepSWE v1.1 at 65.3%. Across all five items, the larger the relative gain, the lower the starting point.

This distribution splits deployment decisions in two. In software engineering work, where the absolute score reached 65.3%, the model can carry real tasks under human review. In enterprise workflow automation, sitting in the 30.4% band, the model still fails to finish most runs, so attaching it as a drafting step with human confirmation costs less than wiring it for autonomous execution. "Workhorse" describes where the model sits in a lineup; it does not claim substitution for a person on every task.

The three-week cadence belongs in the same reading. Google described this release as a direct result of developer feedback and algorithmic innovations. An improvement that lifts five benchmarks in three weeks points to post-training and execution changes rather than a fresh pretraining run. When builds are swapped on that cadence, an architecture frozen around one build's scores becomes a re-validation task at the next one.

No competing model appears anywhere in the comparison

Every benchmark in Google's announcement is compared against exactly one baseline, the company's own 3.6 Flash, and no other vendor's scores appear. Measurement conditions, run counts, and agent framework settings are not stated either. What the table establishes is how far one build moved past its predecessor inside a single company, not where it stands against other models at the same price point.

Several specifications are also absent. Context window size, latency figures, tokens per second, and knowledge cutoff date are all missing from the announcement. Given that "workhorse" normally implies throughput and speed, a launch that lists intelligence metrics with no speed metric at all is worth flagging. On safety, Google says it updated safeguards against misuse in chemical, biological, radiological, and nuclear domains and in cyber offense, without publishing evaluation numbers.

What this release settles is a five-item improvement over the previous generation plus two price sheets by date; what it leaves open is the model's position against competitors and its speed profile. For organizations where the second question decides adoption, the step of measuring it on their own tasks remains exactly where it was.

Three access paths exist for teams in Korea

Google describes three delivery paths, and they are split by audience: developers work through Google Antigravity for agent-first workflows or build in the Gemini API via Google AI Studio and Android Studio; enterprises reach 3.7 Flash through the Gemini Enterprise Agent Platform and the Gemini Enterprise app; individuals get it through Spark, the 24/7 personal agent in the Gemini app, for Google AI Pro and Ultra subscribers.

Regional conditions attach only to the consumer path. Google limits Spark to subscribers in supported countries and explicitly excludes the EEA, Nigeria, Switzerland, and the UK. Korea is not on that exclusion list. The announcement does not enumerate supported countries individually, so final confirmation comes from what the app shows.

Developer and enterprise paths carry no such regional condition. The fastest way for a Korean team to evaluate this model today is to call the Gemini API with gemini-3.7-flash, replay an existing 3.6 Flash workload unchanged, and compare the two builds against their own task set. Since the rates are identical, the marginal cost of that comparison is one extra run, and finishing it inside 2026 means the post-increase budget can be built from in-house numbers.

Source: Google's official blog announcement of Gemini 3.7 Flash (August 13, 2026), and Google's official Gemini API pricing and model documentation, compiled by ASAP

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts