Case TS-B77A46FE4 Sept 2026releaseCompound claim

AI

“Google has launched Gemini 3.8 Flash, its latest AI model focused on reasoning, coding and autonomous agent tasks, which improves on Gemini 3.7 Flash while keeping the same pricing of $0.75 per million input tokens and $3.75 per million output tokens, and performs better across software engineering, finance, legal, and multi-step…”

Plain restatementGoogle released a model named Gemini 3.8 Flash. Google states it improves on Gemini 3.7 Flash, is priced at the same per-token rate ($0.75 input / $3.75 output per million tokens), and scores higher on software engineering, finance, legal, and multi-step reasoning evaluations.

Mostly accurateConfidence High
What this verdict means →

Distortion codes this site does not recognise yet: cost_compute_omission, benchmark_cherry_picking, unreleased_as_released. Not collectible until the field guide has an entry.

This post is mostly accurate. Google did launch Gemini 3.8 Flash on September 2, 2026, and the prices quoted, $0.75 per million input tokens and $3.75 per million output tokens, match Google's own pricing pages and are indeed the same rate as the previous model. Two things the post leaves out matter. First, that price is an introductory promotion that Google says ends on December 31, 2026, after which it doubles to $1.50 and $7.50. Second, the same price per token does not mean the same bill: Google's own docs say the model deliberately uses more tokens on hard tasks, and the independent evaluator Artificial Analysis measured cost per task roughly 40% higher than the previous version. The benchmark table in the post is Google's own, with Google's scores run by Google, and on that same table the model loses clearly to Claude Opus 5 on several agentic and computer-use tests. Google also released a second, restricted-access security model the same day that the post does not mention. The caption does correctly say "Google says" for the performance claims, which is a point in its favour.

The drift / as claimed vs as evidenced

Google [drifted from the evidence:] has launched Gemini 3.8 Flash, [drifted from the evidence:] its latest AI model focused on reasoning, coding and autonomous agent tasks, which improves on Gemini 3.7 Flash [drifted from the evidence:] while keeping the same [drifted from the evidence:] pricing of $0.75 [drifted from the evidence:] per million input [drifted from the evidence:] tokens and $3.75 per million [drifted from the evidence:] output tokens, and [drifted from the evidence:] performs better across software engineering, finance, legal, and multi-step reasoning [drifted from the evidence:] benchmarks.


Google [added by the neutral restatement:] released a model named Gemini 3.8 Flash. [added by the neutral restatement:] Google states it improves on Gemini 3.7 Flash, [added by the neutral restatement:] is priced at the same [added by the neutral restatement:] per-token rate ($0.75 input [added by the neutral restatement:] / $3.75 [added by the neutral restatement:] output per million tokens), and [added by the neutral restatement:] scores higher on software engineering, finance, legal, and multi-step reasoning [added by the neutral restatement:] evaluations.

Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.

The trace / claim to source

Where it appeared
⌿ Omitted qualifier
A load-bearing condition from the source quietly disappears from the claim.
cost_compute_omission
$ Marketing as evidence
Promotional material dressed up as independent proof.
benchmark_cherry_picking
unreleased_as_released
Secondary sourcenamed-outlet journalism
Fortune, "Google shipped four Gemini Flash models in 106 days" (Sept 3, 2026)
Secondary sourcetech press
The Register, 9to5Google, The Decoder launch coverage
Primary sourcevendor of record
Google official announcement, "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber," blog.google
Primary sourcevendor documentation of record
"What's new in Gemini 3.8 Flash," Gemini API docs, ai.google.dev
Primary sourcevendor pricing of record
Gemini Enterprise Agent Platform pricing page, Google Cloud
Primary sourcevendor card of record
Gemini 3.8 Flash Model Card, Google DeepMind
Primary sourcevendor documentation
Gemini 3.8 Flash model page, Gemini API docs and Google Cloud docs
Primary sourceindependent evaluator
Artificial Analysis independent evaluation of Gemini 3.8 Flash
● Primary source found
What is true
  • Google launched Gemini 3.8 Flash on September 2, 2026. This is confirmed on Google's own blog, developer docs, Cloud docs, and DeepMind model card.
  • The positioning is accurately reported: Google describes the model as aimed at software engineering, agentic and autonomous tasks, and multi-step reasoning.
  • The pricing figures are correct. $0.75 per million input tokens and $3.75 per million output tokens is the rate Google published, and it is the same rate as Gemini 3.7 Flash.
  • Google does claim improvements over 3.7 Flash across software engineering, agentic tasks, finance and legal domain benchmarks, and multi-step reasoning. On Google's published table, the 3.8 versus 3.7 comparisons move in Google's stated direction.
  • The caption attributes the performance and pricing statements to Google ("The company says," "Google says," "Source: Google"), which is correct sourcing practice.
  • An independent evaluator, Artificial Analysis, separately measured a 3 point index improvement over 3.7 Flash, so the improvement claim is not vendor-only.
What is misleading
  • Omitted qualifier: the post says the model keeps "the same pricing." Google's docs and Cloud pricing page state the $0.75/$3.75 rate is introductory and expires December 31, 2026, after which it doubles to $1.50/$7.50. A reader takes "same pricing" as the standing price of the model. It is a promotion with a published end date seventeen weeks away.
  • Cost compute omission: the same per-token rate does not mean the same bill. Google's own documentation says the model uses more tokens by design on complex tasks, and Artificial Analysis measured cost per task about 40% higher than 3.7 Flash from roughly 30% more output tokens. "Improves while keeping the same pricing" reads as a free upgrade. Measured real-world cost went up.
  • Marketing as evidence: the comparison table in the image is Google's launch table. Gemini's rows are Google's own runs and the competitor rows are figures reported by those competitors, not a common harness. The caption's "Google says" framing covers this, but the claim text as submitted for verification drops the attribution and states the performance results as fact.
  • Benchmark cherry picking: the image highlights the Claude Opus 5 and GPT-5.6 Sol columns for comparison, and the caption's framing is uniformly favourable. On the very same Google table the model trails Opus 5 heavily on Terminal-bench 4.0 (19.1% versus 51.8%), OSWorld-2.0 (59.0% versus 75.4%), and GDPVal-AA Elo (1545 versus 1824). Several of the apparent wins are also sub-point differences that are ties rather than leads.
  • Omission of a co-released product: Google launched Gemini 3.8 Flash Cyber the same day, and it is not generally available. The post's "its latest AI model," singular, does not reflect that two models shipped and one is gated.
What is uncertain
  • Which effort level (low, medium, high) produced each row of Google's comparison table. The official materials I could reach do not state this per row, and the post does not either.
  • Whether every reasoning row improved. One write-up reports Google's developer docs showing a broader Humanity's Last Exam figure essentially flat at 45.4% for 3.8 Flash versus 45.7% for 3.7 Flash, while other write-ups report the HLE-Verified row rising from 53.6% to 54.9%. I could not retrieve the underlying evaluation PDF in full to reconcile the two rows, so the "multi-step reasoning" part of the claim is supported on the verified split and unresolved on the broader one.
  • The exact DeepSWE v1.1 figure. Several outlets circulated 71.0% while at least one report says Google's evaluation PDF states 73.7%. I did not retrieve the PDF, so I cannot settle which figure Google published.
  • I retrieved Google's official pages as search-index excerpts rather than full page loads. The quoted pricing and improvement language is verbatim from those excerpts, but I did not read the complete pages.
Evidence summary

The release is real and confirmed on Google's own channels. Google's announcement states that, building on 3.7 Flash from three weeks earlier and marking its third Flash release in six weeks, it is introducing Gemini 3.8 in two variants: Gemini 3.8 Flash, described as its most intelligent workhorse model with improvements from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. On pricing, the announcement says directly: "It is available at the same introductory price 1 as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens." The docs attach a condition the post does not carry. Google's developer documentation states that the $0.75/$3.75 rate is introductory through December 31, 2026, and that standard pricing of $1.50 per million input tokens and $7.50 per million output tokens takes effect January 1, 2027. The same page notes that Gemini 3.8 Flash can use more tokens on longer running and complex tasks by design, taking smaller reasoning steps, calling tools iteratively, and verifying its work. Google Cloud's pricing page confirms the same introductory rate and end date applies to 3.8 Flash, 3.7 Flash and 3.6 Flash alike, with standard pricing of $1.5/$7.5 from January 1, 2027. Independent evaluation partially corroborates the improvement. Artificial Analysis reports that with high reasoning, Gemini 3.8 Flash scores 59 on its Intelligence Index, up 3 points from Gemini 3.7 Flash, matching Gemini 3.7 Flash's discounted pricing until the end of the year, and sitting on the intelligence versus cost-per-task frontier at $0.58 per task, which is about 40% higher than its predecessor, driven by a 30% increase in average output tokens per task. Artificial Analysis attributes the index gain primarily to stronger performance on agentic evaluations including tool use, Terminal-Bench v2.1, and GDPval-AA v2. On the benchmark table itself, the numbers are Google's. One analysis notes that on Google's announcement table the reported figures are DeepSWE v1.1 73.7% versus 65.3%, Terminal-bench 2.1 89.4% versus 85.8%, OSWorld-2.0 59.0% versus 50.6%, Vals Finance Agent v2 61.4% versus 59.0%, and HLE-Verified 54.9% versus 53.6%, with every published row moving up, and that all numbers are Google's own runs. The same table shows clear losses against the circled competitors: 3.8 Flash trails Claude Opus 5 on Terminal-bench 4.0 at 19.1% versus 51.8%, on OSWorld-2.0 at 59.0% versus 75.4%, and on GDPVal-AA knowledge-work Elo at 1545 versus 1824. Google also launched a second model the same day that the post does not mention. Gemini 3.8 Flash Cyber is a security-focused variant that is not publicly available, distributed through the Fairwind Program to government agencies and critical infrastructure operators.

Complete reasoning
Every checkable element resolves in the claim's favour on Google's own channels as of 2026-09-04: the model exists, it is generally available, the two price figures are exactly right, and Google does assert improvement over 3.7 Flash in the named domains, with independent corroboration from Artificial Analysis. I considered and rejected "Accurate" because "keeping the same pricing" omits that the rate is a promotion expiring December 31, 2026 and that measured cost per task rose about 40%, which is material to the reader the post is aimed at. I considered and rejected "Source exists but framing is misleading" because the caption itself attributes each performance statement to Google rather than asserting it independently, and the post reproduced the full comparison table including the rows where the model loses, so the framing does not invert the meaning. The gaps are omissions around a correct core, which is exactly what "Mostly accurate" covers. Confidence is High because the deciding artifacts are vendor official channels and this is a release and pricing claim, the class where the vendor channel is primary and decisive.
Use this case

The reply is formatted for pasting into the thread where the claim is circulating.

Compact share page: ai.trueseeker.com/s/b77a46fe9fce/xEYEk7PT9KJS6nIwTXL7Fuj_I6e

Ask this case

Answers come only from the case file above; nothing is added.

Is the $0.75/$3.75 pricing really the same as before?

Yes, that rate matches Gemini 3.7 Flash, but Google's own documentation says it is an introductory price that ends December 31, 2026. After that, standard pricing doubles to $1.50 per million input tokens and $3.75 per million output tokens rising to $7.50 per million output tokens.

Does the same price per token mean the same overall cost to use the model?

No. Google's documentation says Gemini 3.8 Flash is designed to use more tokens on complex tasks, and the independent evaluator Artificial Analysis measured real-world cost per task about 40% higher than Gemini 3.7 Flash, driven by roughly 30% more output tokens.

Are the benchmark improvements confirmed by anyone besides Google?

Partially. Artificial Analysis independently measured a 3 point gain on its Intelligence Index over Gemini 3.7 Flash, mainly from stronger agentic task performance. However, the detailed comparison table in the post comes from Google's own launch materials, with Google running its own model's tests.

How does Gemini 3.8 Flash compare to competitor models like Claude Opus 5?

On Google's own published table, Gemini 3.8 Flash loses clearly to Claude Opus 5 on several tests, including Terminal-bench 4.0, OSWorld-2.0, and GDPVal-AA knowledge-work Elo. The post highlights favorable comparisons but does not mention these losses.

Did Google launch any other models alongside Gemini 3.8 Flash?

Yes. Google also released Gemini 3.8 Flash Cyber, a security-focused model, on the same day. It is not publicly available and is distributed only to government agencies and critical infrastructure operators through a restricted program, a fact the post does not mention.

Similar cases on record