AI Tool Profile
Gemini 3.5 Flash Computer Use: How the Agent Loop and Safety Controls Work
A Gemini API computer-use tool that reasons over screenshots and proposes browser, mobile, or desktop interface actions for a controlled client to validate and execute.

Verification & Sources
- Evidence state
- Recheck due
- Source links
- 1
- Freshness
- Needs recheck: checked July 23, 2026
- Last updated
- July 24, 2026
What this evidence state means
- Definition
- The claim was previously checked, but its review window expired or a material change may have invalidated it.
- Required provenance
- The prior evidence and check date are retained, together with the expiry or change signal that triggered recheck.
- Owner
- Kingy freshness queue owner and assigned editorial reviewer
- Freshness rule
- This is already outside its freshness rule. It must not be presented as current until reviewed against current evidence.
- Disputes and corrections
- Use “Suggest a correction” on the record. Kingy editorial reviews the cited evidence, records material corrections, and changes or removes the state when it is not supported.
Key source checks
Suggest a correction
What It Does
A Gemini API computer-use tool that reasons over screenshots and proposes browser, mobile, or desktop interface actions for a controlled client to validate and execute.
Full Guide
Last updated: 2026-07-23
Last verified: 2026-07-23
TL;DR: Gemini 3.5 Flash Computer Use adds a built-in tool for agents that inspect screenshots, decide what to do next, and request interface actions across browser, mobile, and desktop environments. It can support testing and operations, but the developer—not the model—must execute actions, enforce confirmations, and decide what the agent may touch.
What Google launched
Google announced the capability on June 24, 2026. The company describes it as a computer-use model built on Gemini 3.5 Flash and exposed through the Gemini API and Google Cloud’s enterprise agent platform. The model can reason over a visual interface and propose actions such as clicking, typing, scrolling, or navigating.
The important distinction is that the model does not directly control a machine. An application sends a screenshot and task context, receives a structured action request, executes an allowed action in its own environment, captures the new screen, and returns that state for the next step. That loop gives the application a place to validate coordinates, restrict domains, redact data, and interrupt the run.
How the computer-use loop works
The Gemini API documentation lays out a repeated observe, reason, act cycle. A client supplies the current screenshot and recent history. Gemini can return an action, a safety decision, or a request for confirmation. The client then decides whether the proposed action is valid and whether a person must approve it.
- Open a controlled browser, mobile emulator, or desktop session.
- Send the current visual state and a narrowly written task.
- Validate the requested function and its arguments against an allowlist.
- Pause for human confirmation before consequential or ambiguous actions.
- Execute the approved action, capture the result, and repeat until a stop condition is reached.
Where it may be useful
Computer use is most credible for tasks where interfaces are the only practical integration point. Examples include reproducing a web bug, checking a checkout flow in a test environment, moving through a legacy admin console, or comparing the same interaction across desktop and mobile layouts.
Teams tracking related AI News should separate a successful demo from a reliable production process. Interfaces change, pop-ups appear, coordinates drift, and an agent can mistake an advertisement or destructive button for the next step. A useful pilot measures recovery and supervision, not only task completion.
Safety controls to require
Google’s documentation distinguishes actions that can proceed, actions that require confirmation, and actions that should be blocked. An implementation should add its own policy on top. Domain allowlists, disposable accounts, limited credentials, action budgets, recorded screenshots, and an immediate stop control reduce the impact of a bad decision.
Do not let an early evaluation send messages, publish content, approve payments, change access controls, or delete records. Begin with read-only work and synthetic data. If the agent later receives write access, require a person to review the exact proposed change at the point of action.
Pricing and availability
Google’s current Gemini API pricing page lists a free tier for Gemini 3.5 Flash and paid rates of $1.50 per million input tokens and $9 per million output tokens, with lower batch rates. Computer-use workloads also incur the cost of screenshots, repeated turns, browser infrastructure, logging, and human review.
Documentation now identifies Gemini 3.5 Flash as a previous stable model and points developers toward a newer Flash version. That makes model selection part of the test plan: record the exact model identifier, rerun the same task set after an upgrade, and avoid assuming identical behavior between versions.
Evaluation checklist
- Use at least one success case, one blocked action, and one deliberately confusing interface.
- Measure completion rate, wrong-action rate, confirmations, latency, and total API usage.
- Verify that the agent stops when the page leaves an approved domain.
- Inspect every screenshot and action log for exposed credentials or private data.
- Test whether a human can interrupt the run before an external change occurs.
Kingy AI verdict
Gemini computer use is worth a controlled evaluation for visual testing and bounded legacy workflows. Its value comes from the application’s execution and safety layer as much as the model. Keep the first pilot read-only, make every consequential step reviewable, and promote it only after repeated tests show that the controls work when the model is wrong.
FAQ
Does Gemini directly control the computer?
No. It proposes structured actions. The client application validates and executes those actions, then returns a new screenshot.
Can it work outside a browser?
Google describes browser, mobile, and desktop use. The available actions and reliability depend on the environment built by the developer.
Should it be given production credentials?
Not during an initial evaluation. Use limited test accounts and require confirmation before any action that changes external state.
Official links
Related Kingy AI links
The Kingy Brief
Follow The Kingy Brief.
One consequential launch, one pricing, limit, or shutdown change, one hands-on test, one exact prompt or Test Pack, and one try / watch / skip verdict.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
Tool Links
Launch History
Gemini 3.5 Flash Computer Use
Google launched public preview support for the Computer Use tool in Gemini 3.5 Flash on June 24, 2026.
- Launch readiness
- 7.6 / 10
- Demo evidence
- Not scored yet
- Creator-story fit
- Not scored yet
Score definitions and rubric
These are launch-record readiness heuristics, not product ratings.
Launch readiness
How complete and reviewable the launch record is, not the quality of the product.
Inputs and weights: Launch date 15%; qualifying source 10%; what launched 10%; demo 15%; category 10%; audience 10%; editorial assessment 10%; traction evidence 10%; creator or audience fit 10%.
Evidence inputs: Reviewed launch metadata, public source links, demo links, taxonomy, audience, editorial notes, and recorded traction signals.
Demo evidence
Whether the record contains useful, reviewable demonstration evidence; it is not a rating of product output quality.
Inputs and weights: Working demo URL 45%; video walkthrough 25%; clear description of what launched 10%; audience 10%; editorial assessment 10%.
Evidence inputs: Demo and video URLs plus the reviewed launch description, audience, and editorial notes.
Creator-story fit
Whether a launch has enough demonstrable evidence and audience relevance for a useful creator story; it does not predict views or guarantee coverage.
Inputs and weights: Demo evidence 25%; visual creator category 15%; audience 15%; editorial assessment 15%; traction evidence 10%; pricing clarity 10%; API or open-weight evidence 10%.
Evidence inputs: Reviewed demo, category, audience, editorial, traction, pricing, API, and open-weight fields.
- Scale
- 0.0–10.0. A present qualifying input receives its published weight; a missing input receives zero. Scores are rounded to one decimal.
- Assigned by
- Suggested by the deterministic field-completeness helper and assigned or approved by a Kingy editorial reviewer.
- Rubric and check date
- Rubric version P0-2026-08-10. The record’s “Last verified” date is the score check date. Checked: 2026-06-25.
- Confidence and missing data
- Confidence depends on source completeness. “Not scored yet” means no reviewed value; “Needs review” means the value or score set failed validation.
- Freshness
- Recalculate after a material launch, source, demo, pricing, audience, or traction change and during the record freshness review.
- Disputes
- Use “Suggest a correction” on the record and cite the relevant evidence. Commercial relationships cannot buy or alter a score.
Google shipped public-preview Computer Use support in Gemini 3.5 Flash on June 24, 2026, with dedicated API docs and a live browser demo (ai.google.dev).…