AI News

Gemini 3.8 Flash Review: Benchmarks, Specs, Pricing & Verdict

Verdict: meaningful capability gains with a workload-cost premium

Gemini 3.8 Flash is a meaningful upgrade for long-horizon coding, tool-using agents, finance and legal workflows, and scientific research tasks. Teams buying accepted results rather than the fewest generated tokens should test it now. It keeps Gemini 3.7 Flash’s one-million-token context window, multimodal input, broad tool support and temporary $0.75/$3.75 token prices, while improving Google’s side-by-side scores across every reported 3.7 benchmark row.

It is not an automatic migration for everyone. Google says 3.8 “works harder”: at high effort it can reason longer, call tools more often and consume more output tokens. Independent Artificial Analysis testing found a higher intelligence score but roughly 40% higher cost per benchmark task than 3.7. Use 3.8 for difficult work where better completion rates can repay that overhead. Stay on 3.7 for latency-sensitive, high-concurrency or tightly budgeted pipelines that already pass their acceptance tests.

Testing disclosure, updated September 2, 2026: Kingy ran six low-thinking API calls against gemini-3.8-flash, covering four small deterministic tasks and two targeted retries. The metered-cost estimate was $0.002466. This was a smoke check, not a benchmark; no Kingy-tested claim about published benchmark scores, multimodal work, tools, long context or safety appears below.

Evidence labels: Google-reported means Google documentation, model-card or launch-chart evidence. Independent means a benchmark operator or third-party lab ran the model. Kingy-tested means the six-call smoke check described below. Kingy analysis means arithmetic or a purchasing interpretation derived from labelled evidence.

Gemini 3.8 Flash at a glance

Specification Verified detail Evidence
Status and exact ID Stable / GA; gemini-3.8-flash Google-reported
Direct predecessor Gemini 3.7 Flash; Google says 3.8 is based on it Google-reported
Context window 1,048,576 input tokens Google-reported
Maximum output 65,536 tokens Google-reported
Input / output modalities Text, image, video, audio and PDF in; text out Google-reported
Thinking levels Low, medium and high; medium is the default. minimal returns an error Google-reported
Supported tools Caching, code execution, computer use (preview), file search, function calling, Google Maps and Search grounding, structured output and URL context Google-reported
Not supported Native audio generation, image generation and Live API Google-reported
Consumption modes Standard, Batch, Flex and Priority Google-reported
Availability Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, Gemini app, AI Mode and Antigravity; Google’s launch also names Android Studio, Stitch and Gemini in Sheets Google-reported
Knowledge cutoff March 2026, with some domains potentially limited to January 2025 Google-reported
Undisclosed Parameter count, architecture size/topology, active parameters, layer count, training-token count and training compute Google has not disclosed these specifications

The product envelope barely changes from the Gemini 3.7 Flash review. That is useful: this is primarily a reasoning and agent-behaviour delta, not a new context, modality or storage architecture.

Gemini 3.8 Flash pricing

Google’s paid Gemini Developer API rates below are US dollars per one million tokens. Output prices include thinking tokens. The introductory rates expire after December 31, 2026.

Paid route Through Dec. 31, 2026: input / cached input / output From Jan. 1, 2027: input / cached input / output
Standard $0.75 / $0.075 / $3.75 $1.50 / $0.15 / $7.50
Batch $0.375 / $0.0375 / $1.875 $0.75 / $0.075 / $3.75
Flex $0.375 / $0.0375 / $1.875 $0.75 / $0.075 / $3.75
Priority $1.35 / $0.135 / $6.75 $2.70 / $0.27 / $13.50
Cache storage, per 1M tokens per hour $0.50 $1.00

Google Search and Maps grounding have separate allowances and per-query charges; the pricing page currently lists 5,000 free requests shared across eligible Gemini 3.x usage, then $14 per 1,000 queries. Taxes, enterprise discounts, tool infrastructure and retry costs are not in the token table. Kingy’s AI API Price Index keeps standard, Batch, Flex, Priority and cache identities separate for the same reason.

Gemini 3.8 vs 3.7: benchmark deltas

The following is one evidence set: Google DeepMind’s September 2026 evaluation chart. It is not mixed with independent leaderboard runs. Metric and configuration labels are retained; deltas are Kingy arithmetic.

Workload Benchmark and source configuration 3.8 Flash 3.7 Flash Delta
Coding DeepSWE v1.1, long-horizon software engineering 73.7% 65.3% +8.4 pp
Terminal work Terminal-Bench 2.1, agentic terminal coding 89.4% 85.8% +3.6 pp
Agents Terminal-Bench 4.0, general agent capabilities 19.1% 11.2% +7.9 pp
Professional knowledge GDPVal-AA v2, Elo 1,545 1,482 +63 Elo
Finance agent Vals Finance Agent v2, chart score 61.4% 59.0% +2.4 pp
Legal agent Harvey Legal Agent Benchmark, all-pass rate 10.0% 8.8% +1.2 pp
PDFs GDP.PDF, all-pass rate 35.0% 34.0% +1.0 pp
Multimodal reasoning CharXiv Reasoning, no tools 86.2% 84.5% +1.7 pp
Long video LVBench 87.8% agentic / 87.1% static 85.4%, condition not shown Not cleanly comparable
Expert reasoning HLE-Verified 54.9% 53.6% +1.3 pp
Computer use OSWorld 2.0, partial score, batch tool enabled 59.0% 50.6% +8.4 pp
Science BioMysteryBench, human-solvable split 88.8% 87.1% +1.7 pp
Science BioMysteryBench, human-difficult split 56.5% 43.5% +13.0 pp
Science LABBench2 86.2% 82.1% +4.1 pp

Configuration warning: Google’s separate Gemini Enterprise developer guide reports Terminal-Bench 2.1 at 90.8% versus 81.6%, not 89.4% versus 85.8%. Those pairs belong to different published surfaces or configurations. Kingy does not average or reconcile them. The same caution applies to the 3.7 baselines: Kingy’s August review recorded 1,525 Elo on GDPVal-AA v2 and 47.9% on OSWorld from the then-current 3.7 launch package; Google’s new side-by-side chart uses 1,482 and 50.6%. This review uses the new paired chart for its deltas without rewriting the historical record.

What improved, and what still loses

Google-reported wins: The strongest direct delta is the 13-point jump on BioMysteryBench’s human-difficult split, followed by DeepSWE and OSWorld at +8.4 points. Those are substantial enough to justify workload tests. LABBench2 and Terminal-Bench 2.1 improve more moderately. Finance, legal, PDF, CharXiv and HLE-Verified move by only one to 2.4 points; small chart differences should not decide a purchase by themselves.

The absolute results matter as much as the deltas. Terminal-Bench 4.0 rises sharply from 11.2% to 19.1%, yet Claude Opus 5 reaches 51.8% in Google’s chart. GDP.PDF is 35% against GPT-5.6 Sol at 40%. OSWorld reaches 59%, still behind Opus 5 at 75.4%. On BioMysteryBench’s human-solvable split, 3.8 scores 88.8% versus Opus 5 at 90.1%. These are losses against peers even where 3.8 beats 3.7.

Independent DeepSWE: The public DeepSWE v1.1 leaderboard runs every model with mini-swe-agent. At high effort, it reports 3.8 at 74% ±1%, versus 65% ±2% for 3.7. The 3.8 run averages 143,000 output tokens and 166 agent steps, compared with 107,000 and 125 for 3.7. This independently supports the coding gain and the “works harder” mechanism; it does not validate every Google benchmark.

Independent Artificial Analysis: At high reasoning, 3.8 scores 59 on Artificial Analysis Intelligence Index v4.1.1, up from 56 for 3.7. It generates about 305 output tokens per second, but averages 48,000 output tokens per index task. Artificial Analysis says the 30% token increase and extra agent turns raise cost per task from $0.40 to $0.58 and time per task from 2.2 to 2.5 minutes.

“Works harder” changes the economics

Equal token prices do not mean equal workload cost. A model can spend more billable reasoning tokens, call retrieval or execution tools repeatedly, recover from a failed path and retry a subtask. Those behaviours may increase success, but they also increase token spend, tool fees, wall-clock time and the number of failure points before an answer is accepted.

DeepSWE offers a useful example. Average run cost rises from $2.18 on 3.7 to $2.36 on 3.8, while success rises from roughly 65% to 74%. Dividing average run cost by pass rate gives a crude expected cost-per-pass proxy of about $3.35 for 3.7 and $3.19 for 3.8. That suggests harder work can pay for itself on this benchmark even though each run costs more. It is a Kingy calculation, not a leaderboard metric, and real retry policies will differ.

The opposite appears in Artificial Analysis’s broader index: intelligence improves three points while per-task cost rises about 40%. Buyers should log reasoning tokens, visible output, tool calls, retries, accepted-result rate and p95 latency. Start routine extraction, classification and formatting at low or medium effort; reserve high for tasks where the acceptance-rate gain justifies the overhead.

Kingy-tested: a six-call output-cap smoke check

On September 2, Kingy called the Gemini Interactions API with the stable ID, low thinking and deterministic prompts. Four tasks returned the expected answer. Structured extraction produced correct JSON with zero thought tokens, while an unsupported future claim produced a refusal. Arithmetic returned 415 twice, although both calls remained incomplete after using 89 of 96 and then 185 of 192 available tokens in thought. A code answer was initially cut to 3; at 192 tokens it returned the correct 385 with completed status.

Across six calls, the API reported 218 input tokens, 31 visible output tokens and 583 thought tokens. Latency ranged from 0.968 to 1.842 seconds. At Google’s introductory standard rates, estimated spend was $0.002466, about one-quarter of one cent. The narrow result supports a practical warning: even low thinking can exhaust a tight generation ceiling before the visible answer finishes. It does not validate Google’s benchmarks or establish production reliability.

Safety, hallucinations and operational risk

Google’s model card is unusually direct about the trade-offs. It warns that Gemini 3.8 Flash can hallucinate, may be occasionally slow or time out, and can use more tokens at higher effort. Search grounding can improve freshness, but it does not make an answer correct or fully sourced.

On Google’s automated development evaluations versus 3.7, text-to-text safety improves by 0.4 percentage points and image-to-text safety is unchanged. Tone improves 0.2 points. Two results move the wrong way: unjustified refusals worsen 1.1 points, and multilingual safety worsens 5.4 points where lower is better. Google calls the multilingual regression slight and says manual review found losses overwhelmingly false positives or non-egregious. Buyers operating in non-English languages should still treat it as a deployment-specific test requirement, not dismiss it.

Google says specialist red teams found similar or improved content-safety performance overall, child-safety launch thresholds were met, and no egregious concerns emerged. Its Frontier Safety conclusion is inherited partly from 3.7: Google found no meaningful new capability increase in tracked high-risk domains and considers 3.8 unlikely to reach a Tracked or Critical Capability Level. That is vendor safety evidence, not an independent audit.

Use-case recommendations

Use 3.8 Flash for repository-scale coding agents, multi-step terminal automation, PDF and mixed-media research, finance or legal assistants with human review, browser/computer-use pilots, and scientific workflows where a higher accepted-result rate matters more than minimum token use. It is also the better shortlist candidate for new Gemini deployments because the API ID is stable and the headline price is unchanged.

Remain on 3.7 Flash when the workload is already validated, high-volume and latency-sensitive; when output-token or tool-call budgets are hard limits; when lower effort on 3.8 does not reproduce your 3.7 quality; or when multilingual safety behaviour is material and has not been requalified. The Gemini 3.7 migration guide remains useful because 3.8 keeps the same Gemini 3.x thinking and request-contract constraints. Do not change the model ID and traffic allocation in one step. Shadow, compare and roll forward by task class.

For broader model discovery, use Kingy’s AI Model Intelligence Hub. Context size, benchmark score and token rate are inputs to a decision, not substitutes for a production acceptance suite.

Kingy scorecard

Category Score Why
Capability 9.2/10 Material gains in hard coding, computer use and difficult science
Price-performance 9.0/10 Excellent temporary rates, tempered by higher tokens and cost per task
Multimodal breadth 9.2/10 Text, image, audio, video and PDF input under one stable ID
Agents and tooling 9.1/10 Broad tools and stronger loops; computer use remains preview
Operational efficiency 8.1/10 Fast decoding, but more reasoning, occasional slowness and timeouts
Safety and transparency 7.9/10 Useful model-card disclosure; multilingual safety regressed
Evidence quality 8.6/10 Independent AA and DeepSWE support key claims; Kingy added a narrow API smoke check, while configuration variance remains
Overall 8.9/10 Recommended for controlled evaluation, not a blanket 3.7 replacement

Methodology and limitations

Kingy checked Google’s launch article, Gemini API model page, paid pricing table, Enterprise developer guide and September 2026 model card on September 2, 2026. We transcribed Google’s evaluation chart, preserved its metric labels and recomputed the direct 3.8-versus-3.7 deltas. We separately checked Artificial Analysis and the DeepSWE public leaderboard. Kingy also ran six calls through the Gemini Interactions API and calculated cost from the returned input, visible-output and thought-token counts. Vals Finance Agent and Harvey Legal pages were reviewed for benchmark definitions and harness context, but Google’s launch-chart figures remain labelled Google-reported because Kingy did not reproduce those runs.

This is launch-day evidence. Scores can change with model snapshots, agent harnesses, effort levels, tool permissions, task revisions, graders and service routes. LVBench’s 3.7 condition is not shown in Google’s chart, so no clean delta is claimed. No latency service-level objective, parameter count, architecture size or training-compute figure has been disclosed. Kingy’s six-call smoke check did not test benchmark replication, multimodal inputs, tools, long-context recall, multilingual safety or reliability at scale.

FAQ

Is Gemini 3.8 Flash generally available?

Yes. Google lists gemini-3.8-flash as stable, and its Enterprise guide labels the launch stage GA. Computer use remains preview.

What is the Gemini 3.8 Flash context window?

The model accepts up to 1,048,576 input tokens and can produce up to 65,536 output tokens. Capacity is not proof of perfect recall across a million-token prompt.

How much does Gemini 3.8 Flash cost?

Standard paid pricing is $0.75 per million input tokens, $0.075 cached input and $3.75 output through December 31, 2026. On January 1, 2027, those rates double to $1.50, $0.15 and $7.50. Thinking tokens are billed as output.

Is Gemini 3.8 better than 3.7?

Yes for the difficult coding, agent, computer-use and science workloads covered by the strongest evidence. Not necessarily for cost- or latency-constrained production jobs: independent testing shows 3.8 using more tokens and turns.

Does Gemini 3.8 Flash support PDF, image, audio and video input?

Yes. It accepts all four plus text and returns text. It does not natively generate images or audio, and it does not support the Live API.

Was this a hands-on Kingy review?

Partly. The benchmark analysis remains research-led, but Kingy later added six low-thinking API calls across four deterministic tasks. Treat those calls as an operational smoke check, not an independent benchmark.

Primary sources and update log

Update log, September 2, 2026: Initial publication. Verified stable/GA status, exact model ID, context and output limits, modalities, thinking levels, tools, availability, cutoff, all paid pricing routes and scheduled January 2027 increase. Added Google-reported benchmark deltas, independent Artificial Analysis and DeepSWE evidence, model-card safety results and the initial no-testing disclosure.

Update log, September 2, 2026 at 18:40 UTC: Added Kingy’s six-call low-thinking smoke check, exact token and latency totals, the $0.002466 metered-cost estimate, the output-cap finding and revised testing disclosures. No published benchmark was rerun.