AI Tool Profile

Gemini 3.5 Flash Computer Use: How the Agent Loop and Safety Controls Work

A Gemini API computer-use tool that reasons over screenshots and proposes browser, mobile, or desktop interface actions for a controlled client to validate and execute.

QA engineer supervises a computer-use workflow across desktop, laptop, and phone

Verification & Sources

Evidence state
Recheck due
Source links
1
Freshness
Needs recheck: checked July 23, 2026
Last updated
July 24, 2026
What this evidence state means
Definition
The claim was previously checked, but its review window expired or a material change may have invalidated it.
Required provenance
The prior evidence and check date are retained, together with the expiry or change signal that triggered recheck.
Owner
Kingy freshness queue owner and assigned editorial reviewer
Freshness rule
This is already outside its freshness rule. It must not be presented as current until reviewed against current evidence.
Disputes and corrections
Use “Suggest a correction” on the record. Kingy editorial reviews the cited evidence, records material corrections, and changes or removes the state when it is not supported.

Key source checks

Suggest a correction

Form submissions, correction notes, score details, URLs, and analytics events may be stored for editorial review, spam prevention, product improvement, and follow-up. Do not submit secrets, unreleased financials, private customer data, or regulated personal data through these forms.

What It Does

A Gemini API computer-use tool that reasons over screenshots and proposes browser, mobile, or desktop interface actions for a controlled client to validate and execute.

Full Guide

Last updated: 2026-07-23

Last verified: 2026-07-23

TL;DR: Gemini 3.5 Flash Computer Use adds a built-in tool for agents that inspect screenshots, decide what to do next, and request interface actions across browser, mobile, and desktop environments. It can support testing and operations, but the developer—not the model—must execute actions, enforce confirmations, and decide what the agent may touch.

What Google launched

Google announced the capability on June 24, 2026. The company describes it as a computer-use model built on Gemini 3.5 Flash and exposed through the Gemini API and Google Cloud’s enterprise agent platform. The model can reason over a visual interface and propose actions such as clicking, typing, scrolling, or navigating.

The important distinction is that the model does not directly control a machine. An application sends a screenshot and task context, receives a structured action request, executes an allowed action in its own environment, captures the new screen, and returns that state for the next step. That loop gives the application a place to validate coordinates, restrict domains, redact data, and interrupt the run.

How the computer-use loop works

The Gemini API documentation lays out a repeated observe, reason, act cycle. A client supplies the current screenshot and recent history. Gemini can return an action, a safety decision, or a request for confirmation. The client then decides whether the proposed action is valid and whether a person must approve it.

  1. Open a controlled browser, mobile emulator, or desktop session.
  2. Send the current visual state and a narrowly written task.
  3. Validate the requested function and its arguments against an allowlist.
  4. Pause for human confirmation before consequential or ambiguous actions.
  5. Execute the approved action, capture the result, and repeat until a stop condition is reached.

Where it may be useful

Computer use is most credible for tasks where interfaces are the only practical integration point. Examples include reproducing a web bug, checking a checkout flow in a test environment, moving through a legacy admin console, or comparing the same interaction across desktop and mobile layouts.

Teams tracking related AI News should separate a successful demo from a reliable production process. Interfaces change, pop-ups appear, coordinates drift, and an agent can mistake an advertisement or destructive button for the next step. A useful pilot measures recovery and supervision, not only task completion.

Safety controls to require

Google’s documentation distinguishes actions that can proceed, actions that require confirmation, and actions that should be blocked. An implementation should add its own policy on top. Domain allowlists, disposable accounts, limited credentials, action budgets, recorded screenshots, and an immediate stop control reduce the impact of a bad decision.

Do not let an early evaluation send messages, publish content, approve payments, change access controls, or delete records. Begin with read-only work and synthetic data. If the agent later receives write access, require a person to review the exact proposed change at the point of action.

Pricing and availability

Google’s current Gemini API pricing page lists a free tier for Gemini 3.5 Flash and paid rates of $1.50 per million input tokens and $9 per million output tokens, with lower batch rates. Computer-use workloads also incur the cost of screenshots, repeated turns, browser infrastructure, logging, and human review.

Documentation now identifies Gemini 3.5 Flash as a previous stable model and points developers toward a newer Flash version. That makes model selection part of the test plan: record the exact model identifier, rerun the same task set after an upgrade, and avoid assuming identical behavior between versions.

Evaluation checklist

  • Use at least one success case, one blocked action, and one deliberately confusing interface.
  • Measure completion rate, wrong-action rate, confirmations, latency, and total API usage.
  • Verify that the agent stops when the page leaves an approved domain.
  • Inspect every screenshot and action log for exposed credentials or private data.
  • Test whether a human can interrupt the run before an external change occurs.

Kingy AI verdict

Gemini computer use is worth a controlled evaluation for visual testing and bounded legacy workflows. Its value comes from the application’s execution and safety layer as much as the model. Keep the first pilot read-only, make every consequential step reviewable, and promote it only after repeated tests show that the controls work when the model is wrong.

FAQ

Does Gemini directly control the computer?

No. It proposes structured actions. The client application validates and executes those actions, then returns a new screenshot.

Can it work outside a browser?

Google describes browser, mobile, and desktop use. The available actions and reliability depend on the environment built by the developer.

Should it be given production credentials?

Not during an initial evaluation. Use limited test accounts and require confirmation before any action that changes external state.

Official links

Related Kingy AI links

Launch History

AI Developer Tools

Gemini 3.5 Flash Computer Use

Google launched public preview support for the Computer Use tool in Gemini 3.5 Flash on June 24, 2026.

Recheck due Free: Yes API: Yes Open: Unknown
Clear use caseVideo demoTraction signal
Launch readiness
7.6 / 10
Demo evidence
Not scored yet
Creator-story fit
Not scored yet
Score definitions and rubric

These are launch-record readiness heuristics, not product ratings.

Launch readiness

How complete and reviewable the launch record is, not the quality of the product.

Inputs and weights: Launch date 15%; qualifying source 10%; what launched 10%; demo 15%; category 10%; audience 10%; editorial assessment 10%; traction evidence 10%; creator or audience fit 10%.

Evidence inputs: Reviewed launch metadata, public source links, demo links, taxonomy, audience, editorial notes, and recorded traction signals.

Demo evidence

Whether the record contains useful, reviewable demonstration evidence; it is not a rating of product output quality.

Inputs and weights: Working demo URL 45%; video walkthrough 25%; clear description of what launched 10%; audience 10%; editorial assessment 10%.

Evidence inputs: Demo and video URLs plus the reviewed launch description, audience, and editorial notes.

Creator-story fit

Whether a launch has enough demonstrable evidence and audience relevance for a useful creator story; it does not predict views or guarantee coverage.

Inputs and weights: Demo evidence 25%; visual creator category 15%; audience 15%; editorial assessment 15%; traction evidence 10%; pricing clarity 10%; API or open-weight evidence 10%.

Evidence inputs: Reviewed demo, category, audience, editorial, traction, pricing, API, and open-weight fields.

Scale
0.0–10.0. A present qualifying input receives its published weight; a missing input receives zero. Scores are rounded to one decimal.
Assigned by
Suggested by the deterministic field-completeness helper and assigned or approved by a Kingy editorial reviewer.
Rubric and check date
Rubric version P0-2026-08-10. The record’s “Last verified” date is the score check date. Checked: 2026-06-25.
Confidence and missing data
Confidence depends on source completeness. “Not scored yet” means no reviewed value; “Needs review” means the value or score set failed validation.
Freshness
Recalculate after a material launch, source, demo, pricing, audience, or traction change and during the record freshness review.
Disputes
Use “Suggest a correction” on the record and cite the relevant evidence. Commercial relationships cannot buy or alter a score.

Google shipped public-preview Computer Use support in Gemini 3.5 Flash on June 24, 2026, with dedicated API docs and a live browser demo (ai.google.dev).…