Trending on Kingy
Keep reading with the stories getting the most attention now.
Verdict: meaningful capability gains with a workload-cost premium
Gemini 3.8 Flash is a meaningful upgrade for long-horizon coding, tool-using agents, finance and legal workflows, and scientific research tasks. Teams buying accepted results rather than the fewest generated tokens should test it now. It keeps Gemini 3.7 Flash’s one-million-token context window, multimodal input, broad tool support and temporary $0.75/$3.75 token prices, while improving Google’s side-by-side scores across every reported 3.7 benchmark row.
It is not an automatic migration for everyone. Google says 3.8 “works harder”: at high effort it can reason longer, call tools more often and consume more output tokens. Independent Artificial Analysis testing found a higher intelligence score but roughly 40% higher cost per benchmark task than 3.7. Use 3.8 for difficult work where better completion rates can repay that overhead. Stay on 3.7 for latency-sensitive, high-concurrency or tightly budgeted pipelines that already pass their acceptance tests.
Testing disclosure, updated September 2, 2026: Kingy ran six low-thinking API calls against
gemini-3.8-flash, covering four small deterministic tasks and two targeted retries. The metered-cost estimate was $0.002466. This was a smoke check, not a benchmark; no Kingy-tested claim about published benchmark scores, multimodal work, tools, long context or safety appears below.
Evidence labels: Google-reported means Google documentation, model-card or launch-chart evidence. Independent means a benchmark operator or third-party lab ran the model. Kingy-tested means the six-call smoke check described below. Kingy analysis means arithmetic or a purchasing interpretation derived from labelled evidence.
Gemini 3.8 Flash at a glance
| Specification | Verified detail | Evidence |
|---|---|---|
| Status and exact ID | Stable / GA; gemini-3.8-flash |
Google-reported |
| Direct predecessor | Gemini 3.7 Flash; Google says 3.8 is based on it | Google-reported |
| Context window | 1,048,576 input tokens | Google-reported |
| Maximum output | 65,536 tokens | Google-reported |
| Input / output modalities | Text, image, video, audio and PDF in; text out | Google-reported |
| Thinking levels | Low, medium and high; medium is the default. minimal returns an error |
Google-reported |
| Supported tools | Caching, code execution, computer use (preview), file search, function calling, Google Maps and Search grounding, structured output and URL context | Google-reported |
| Not supported | Native audio generation, image generation and Live API | Google-reported |
| Consumption modes | Standard, Batch, Flex and Priority | Google-reported |
| Availability | Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, Gemini app, AI Mode and Antigravity; Google’s launch also names Android Studio, Stitch and Gemini in Sheets | Google-reported |
| Knowledge cutoff | March 2026, with some domains potentially limited to January 2025 | Google-reported |
| Undisclosed | Parameter count, architecture size/topology, active parameters, layer count, training-token count and training compute | Google has not disclosed these specifications |
The product envelope barely changes from the Gemini 3.7 Flash review. That is useful: this is primarily a reasoning and agent-behaviour delta, not a new context, modality or storage architecture.
Gemini 3.8 Flash pricing
Google’s paid Gemini Developer API rates below are US dollars per one million tokens. Output prices include thinking tokens. The introductory rates expire after December 31, 2026.
| Paid route | Through Dec. 31, 2026: input / cached input / output | From Jan. 1, 2027: input / cached input / output |
|---|---|---|
| Standard | $0.75 / $0.075 / $3.75 | $1.50 / $0.15 / $7.50 |
| Batch | $0.375 / $0.0375 / $1.875 | $0.75 / $0.075 / $3.75 |
| Flex | $0.375 / $0.0375 / $1.875 | $0.75 / $0.075 / $3.75 |
| Priority | $1.35 / $0.135 / $6.75 | $2.70 / $0.27 / $13.50 |
| Cache storage, per 1M tokens per hour | $0.50 | $1.00 |
Google Search and Maps grounding have separate allowances and per-query charges; the pricing page currently lists 5,000 free requests shared across eligible Gemini 3.x usage, then $14 per 1,000 queries. Taxes, enterprise discounts, tool infrastructure and retry costs are not in the token table. Kingy’s AI API Price Index keeps standard, Batch, Flex, Priority and cache identities separate for the same reason.
Gemini 3.8 vs 3.7: benchmark deltas
The following is one evidence set: Google DeepMind’s September 2026 evaluation chart. It is not mixed with independent leaderboard runs. Metric and configuration labels are retained; deltas are Kingy arithmetic.
| Workload | Benchmark and source configuration | 3.8 Flash | 3.7 Flash | Delta |
|---|---|---|---|---|
| Coding | DeepSWE v1.1, long-horizon software engineering | 73.7% | 65.3% | +8.4 pp |
| Terminal work | Terminal-Bench 2.1, agentic terminal coding | 89.4% | 85.8% | +3.6 pp |
| Agents | Terminal-Bench 4.0, general agent capabilities | 19.1% | 11.2% | +7.9 pp |
| Professional knowledge | GDPVal-AA v2, Elo | 1,545 | 1,482 | +63 Elo |
| Finance agent | Vals Finance Agent v2, chart score | 61.4% | 59.0% | +2.4 pp |
| Legal agent | Harvey Legal Agent Benchmark, all-pass rate | 10.0% | 8.8% | +1.2 pp |
| PDFs | GDP.PDF, all-pass rate | 35.0% | 34.0% | +1.0 pp |
| Multimodal reasoning | CharXiv Reasoning, no tools | 86.2% | 84.5% | +1.7 pp |
| Long video | LVBench | 87.8% agentic / 87.1% static | 85.4%, condition not shown | Not cleanly comparable |
| Expert reasoning | HLE-Verified | 54.9% | 53.6% | +1.3 pp |
| Computer use | OSWorld 2.0, partial score, batch tool enabled | 59.0% | 50.6% | +8.4 pp |
| Science | BioMysteryBench, human-solvable split | 88.8% | 87.1% | +1.7 pp |
| Science | BioMysteryBench, human-difficult split | 56.5% | 43.5% | +13.0 pp |
| Science | LABBench2 | 86.2% | 82.1% | +4.1 pp |
Configuration warning: Google’s separate Gemini Enterprise developer guide reports Terminal-Bench 2.1 at 90.8% versus 81.6%, not 89.4% versus 85.8%. Those pairs belong to different published surfaces or configurations. Kingy does not average or reconcile them. The same caution applies to the 3.7 baselines: Kingy’s August review recorded 1,525 Elo on GDPVal-AA v2 and 47.9% on OSWorld from the then-current 3.7 launch package; Google’s new side-by-side chart uses 1,482 and 50.6%. This review uses the new paired chart for its deltas without rewriting the historical record.
What improved, and what still loses
Google-reported wins: The strongest direct delta is the 13-point jump on BioMysteryBench’s human-difficult split, followed by DeepSWE and OSWorld at +8.4 points. Those are substantial enough to justify workload tests. LABBench2 and Terminal-Bench 2.1 improve more moderately. Finance, legal, PDF, CharXiv and HLE-Verified move by only one to 2.4 points; small chart differences should not decide a purchase by themselves.
The absolute results matter as much as the deltas. Terminal-Bench 4.0 rises sharply from 11.2% to 19.1%, yet Claude Opus 5 reaches 51.8% in Google’s chart. GDP.PDF is 35% against GPT-5.6 Sol at 40%. OSWorld reaches 59%, still behind Opus 5 at 75.4%. On BioMysteryBench’s human-solvable split, 3.8 scores 88.8% versus Opus 5 at 90.1%. These are losses against peers even where 3.8 beats 3.7.
Independent DeepSWE: The public DeepSWE v1.1 leaderboard runs every model with mini-swe-agent. At high effort, it reports 3.8 at 74% ±1%, versus 65% ±2% for 3.7. The 3.8 run averages 143,000 output tokens and 166 agent steps, compared with 107,000 and 125 for 3.7. This independently supports the coding gain and the “works harder” mechanism; it does not validate every Google benchmark.
Independent Artificial Analysis: At high reasoning, 3.8 scores 59 on Artificial Analysis Intelligence Index v4.1.1, up from 56 for 3.7. It generates about 305 output tokens per second, but averages 48,000 output tokens per index task. Artificial Analysis says the 30% token increase and extra agent turns raise cost per task from $0.40 to $0.58 and time per task from 2.2 to 2.5 minutes.
“Works harder” changes the economics
Equal token prices do not mean equal workload cost. A model can spend more billable reasoning tokens, call retrieval or execution tools repeatedly, recover from a failed path and retry a subtask. Those behaviours may increase success, but they also increase token spend, tool fees, wall-clock time and the number of failure points before an answer is accepted.
DeepSWE offers a useful example. Average run cost rises from $2.18 on 3.7 to $2.36 on 3.8, while success rises from roughly 65% to 74%. Dividing average run cost by pass rate gives a crude expected cost-per-pass proxy of about $3.35 for 3.7 and $3.19 for 3.8. That suggests harder work can pay for itself on this benchmark even though each run costs more. It is a Kingy calculation, not a leaderboard metric, and real retry policies will differ.
The opposite appears in Artificial Analysis’s broader index: intelligence improves three points while per-task cost rises about 40%. Buyers should log reasoning tokens, visible output, tool calls, retries, accepted-result rate and p95 latency. Start routine extraction, classification and formatting at low or medium effort; reserve high for tasks where the acceptance-rate gain justifies the overhead.
Kingy-tested: a six-call output-cap smoke check
On September 2, Kingy called the Gemini Interactions API with the stable ID, low thinking and deterministic prompts. Four tasks returned the expected answer. Structured extraction produced correct JSON with zero thought tokens, while an unsupported future claim produced a refusal. Arithmetic returned 415 twice, although both calls remained incomplete after using 89 of 96 and then 185 of 192 available tokens in thought. A code answer was initially cut to 3; at 192 tokens it returned the correct 385 with completed status.
Across six calls, the API reported 218 input tokens, 31 visible output tokens and 583 thought tokens. Latency ranged from 0.968 to 1.842 seconds. At Google’s introductory standard rates, estimated spend was $0.002466, about one-quarter of one cent. The narrow result supports a practical warning: even low thinking can exhaust a tight generation ceiling before the visible answer finishes. It does not validate Google’s benchmarks or establish production reliability.
Safety, hallucinations and operational risk
Google’s model card is unusually direct about the trade-offs. It warns that Gemini 3.8 Flash can hallucinate, may be occasionally slow or time out, and can use more tokens at higher effort. Search grounding can improve freshness, but it does not make an answer correct or fully sourced.
On Google’s automated development evaluations versus 3.7, text-to-text safety improves by 0.4 percentage points and image-to-text safety is unchanged. Tone improves 0.2 points. Two results move the wrong way: unjustified refusals worsen 1.1 points, and multilingual safety worsens 5.4 points where lower is better. Google calls the multilingual regression slight and says manual review found losses overwhelmingly false positives or non-egregious. Buyers operating in non-English languages should still treat it as a deployment-specific test requirement, not dismiss it.
Google says specialist red teams found similar or improved content-safety performance overall, child-safety launch thresholds were met, and no egregious concerns emerged. Its Frontier Safety conclusion is inherited partly from 3.7: Google found no meaningful new capability increase in tracked high-risk domains and considers 3.8 unlikely to reach a Tracked or Critical Capability Level. That is vendor safety evidence, not an independent audit.
Use-case recommendations
Use 3.8 Flash for repository-scale coding agents, multi-step terminal automation, PDF and mixed-media research, finance or legal assistants with human review, browser/computer-use pilots, and scientific workflows where a higher accepted-result rate matters more than minimum token use. It is also the better shortlist candidate for new Gemini deployments because the API ID is stable and the headline price is unchanged.
Remain on 3.7 Flash when the workload is already validated, high-volume and latency-sensitive; when output-token or tool-call budgets are hard limits; when lower effort on 3.8 does not reproduce your 3.7 quality; or when multilingual safety behaviour is material and has not been requalified. The Gemini 3.7 migration guide remains useful because 3.8 keeps the same Gemini 3.x thinking and request-contract constraints. Do not change the model ID and traffic allocation in one step. Shadow, compare and roll forward by task class.
For broader model discovery, use Kingy’s AI Model Intelligence Hub. Context size, benchmark score and token rate are inputs to a decision, not substitutes for a production acceptance suite.
Kingy scorecard
| Category | Score | Why |
|---|---|---|
| Capability | 9.2/10 | Material gains in hard coding, computer use and difficult science |
| Price-performance | 9.0/10 | Excellent temporary rates, tempered by higher tokens and cost per task |
| Multimodal breadth | 9.2/10 | Text, image, audio, video and PDF input under one stable ID |
| Agents and tooling | 9.1/10 | Broad tools and stronger loops; computer use remains preview |
| Operational efficiency | 8.1/10 | Fast decoding, but more reasoning, occasional slowness and timeouts |
| Safety and transparency | 7.9/10 | Useful model-card disclosure; multilingual safety regressed |
| Evidence quality | 8.6/10 | Independent AA and DeepSWE support key claims; Kingy added a narrow API smoke check, while configuration variance remains |
| Overall | 8.9/10 | Recommended for controlled evaluation, not a blanket 3.7 replacement |
Methodology and limitations
Kingy checked Google’s launch article, Gemini API model page, paid pricing table, Enterprise developer guide and September 2026 model card on September 2, 2026. We transcribed Google’s evaluation chart, preserved its metric labels and recomputed the direct 3.8-versus-3.7 deltas. We separately checked Artificial Analysis and the DeepSWE public leaderboard. Kingy also ran six calls through the Gemini Interactions API and calculated cost from the returned input, visible-output and thought-token counts. Vals Finance Agent and Harvey Legal pages were reviewed for benchmark definitions and harness context, but Google’s launch-chart figures remain labelled Google-reported because Kingy did not reproduce those runs.
This is launch-day evidence. Scores can change with model snapshots, agent harnesses, effort levels, tool permissions, task revisions, graders and service routes. LVBench’s 3.7 condition is not shown in Google’s chart, so no clean delta is claimed. No latency service-level objective, parameter count, architecture size or training-compute figure has been disclosed. Kingy’s six-call smoke check did not test benchmark replication, multimodal inputs, tools, long-context recall, multilingual safety or reliability at scale.
FAQ
Is Gemini 3.8 Flash generally available?
Yes. Google lists gemini-3.8-flash as stable, and its Enterprise guide labels the launch stage GA. Computer use remains preview.
What is the Gemini 3.8 Flash context window?
The model accepts up to 1,048,576 input tokens and can produce up to 65,536 output tokens. Capacity is not proof of perfect recall across a million-token prompt.
How much does Gemini 3.8 Flash cost?
Standard paid pricing is $0.75 per million input tokens, $0.075 cached input and $3.75 output through December 31, 2026. On January 1, 2027, those rates double to $1.50, $0.15 and $7.50. Thinking tokens are billed as output.
Is Gemini 3.8 better than 3.7?
Yes for the difficult coding, agent, computer-use and science workloads covered by the strongest evidence. Not necessarily for cost- or latency-constrained production jobs: independent testing shows 3.8 using more tokens and turns.
Does Gemini 3.8 Flash support PDF, image, audio and video input?
Yes. It accepts all four plus text and returns text. It does not natively generate images or audio, and it does not support the Live API.
Was this a hands-on Kingy review?
Partly. The benchmark analysis remains research-led, but Kingy later added six low-thinking API calls across four deterministic tasks. Treat those calls as an operational smoke check, not an independent benchmark.
Primary sources and update log
- Google, Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, accessed September 2, 2026.
- Google AI for Developers, Gemini 3.8 Flash model documentation, accessed September 2, 2026.
- Google AI for Developers, Gemini Developer API pricing, accessed September 2, 2026.
- Google DeepMind, Gemini 3.8 Flash model card, accessed September 2, 2026.
- Google Cloud, Developer’s guide to Gemini 3.8 Flash, accessed September 2, 2026.
- Artificial Analysis, Gemini 3.8 Flash independent analysis, accessed September 2, 2026.
- DataCurve, DeepSWE v1.1 leaderboard, accessed September 2, 2026.
- Vals AI, Finance Agent v2 and Harvey’s Legal Agent Benchmark, accessed September 2, 2026.
Update log, September 2, 2026: Initial publication. Verified stable/GA status, exact model ID, context and output limits, modalities, thinking levels, tools, availability, cutoff, all paid pricing routes and scheduled January 2027 increase. Added Google-reported benchmark deltas, independent Artificial Analysis and DeepSWE evidence, model-card safety results and the initial no-testing disclosure.
Update log, September 2, 2026 at 18:40 UTC: Added Kingy’s six-call low-thinking smoke check, exact token and latency totals, the $0.002466 metered-cost estimate, the output-cap finding and revised testing disclosures. No published benchmark was rerun.
