Published October 8, 2026. Updated October 8 after a follow-up check of Google’s live product FAQ, billing documentation, release notes, and evaluation guides. Kingy.ai has not tested the new Gemini agent, run a matched competitor evaluation, or measured its costs.
Google announced a persistent Gemini work agent today that can coordinate subagents and use Anthropic’s Claude models. Google’s launch notice positions it as a way to delegate finished work across business systems.
The strongest part of the announcement is the separation between the agent people work with and the model doing the reasoning underneath. A company could build its workflows around Google’s environment while using a rival’s intelligence. That gives Google a route to remain central to enterprise AI even when a customer prefers Claude for a particular job.
There is a practical limit to today’s news. Google’s SMB announcement says the agent is in early access with SMB customers and broader availability is coming soon. It does not give a dated general rollout. We could not establish a complete price for the newly announced experience from the launch materials reviewed.
The live product FAQ now makes the feature limits more explicit: multi-step background delegation, mobile and desktop access, and third-party model choice are in early access. Its model answer lists Google Gemini models today, with additional models coming soon. The keynote’s Claude support should therefore be read as an announced capability with access limits, not a feature every customer can select today. The FAQ offers a 30-day Gemini Enterprise Plus trial, which does not establish access to every early-access feature.
Treat this as a consequential product announcement with substantial unanswered questions. A purchasing recommendation needs access terms, a working product, and evidence that it finishes your company’s tasks correctly.
The announced specifications
The following is a compact summary of Google’s October 8 keynote announcement. These are vendor descriptions, not Kingy-tested capabilities.
| Component | Google’s announced design |
|---|---|
| Execution | Cloud persistence across hours or days; scheduled and event-driven work. |
| Orchestration | Dynamic temporary subagents; parallel and sequential coordination. |
| Models | Gemini and Claude announced; broader third-party model access remains limited. |
| Memory | Session, semantic, procedural, and episodic memory. |
| Workspace | Gmail, Drive, Docs, Slides, Sheets, Chat, Calendar. |
| Other surfaces | Web, mobile, desktops, CLI, Microsoft 365, Slack, headless integration. |
| Connections | Enterprise tools and skills registries; MCP servers. |
| Coworker identity | Dedicated email, storage, Workspace presence; shared-context access. |
| Governance | Attested identities, role permissions, OAuth propagation, agent-attributed audits. |
| Runtime boundary | Agent Sandbox and policy-enforcing Agent Gateway. |
| Cost controls | Model routing; project caps tracking tokens and sandbox costs, pausing work at the limit. |
That table describes the operating architecture. It leaves several of the specifications that buyers usually seek unresolved: exact selectable model versions, context and output limits for each route, maximum simultaneous subagents, maximum task duration, execution quotas, memory retention, regional coverage, and a complete entitlement matrix.
An agent’s memory store is also a different thing from a model’s context window. Saving a project’s history does not establish how much of that history the model sees at once, how reliably the system retrieves the relevant facts, or how it handles a later correction. Those need separate documentation and testing.
The same caution applies to interoperability. A named surface or connector does not tell you which actions it supports, which editions include it, or whether it reads, drafts, and writes with the same permissions. Buyers need those details at the action level.
Google can compete for the workflow even when Claude wins the task
Google’s model choice creates an interesting division of responsibilities. The agent environment can hold the project, connect the tools, enforce boundaries, and preserve the work history. The selected model can change as the task or the market changes.
For customers, that offers a possible reduction in migration work. If a different model performs better on contract analysis next quarter, the ideal experience would let a team change the reasoning engine while keeping the rest of its workflow intact. Whether the product delivers that experience remains a question for testing.
Model choice still requires verification. A workflow tuned for one model can behave differently on another. It may interpret the same instructions differently, call tools in a different order, produce a different file structure, or require more attempts. A model upgrade should therefore trigger regression checks even when the agent interface looks unchanged.
There is a commercial implication as well. Google can seek to own the place where work is assigned, the connections to business data, and the administrative controls around execution. It would not need to win every model comparison to make that position valuable. That is our reading of the strategy, rather than a claim about Google’s undisclosed revenue arrangements with Anthropic.
Choice also has boundaries. An enterprise should verify which provider processes each request, what information is transmitted, which regional and retention terms apply, and whether administrators can restrict routing. Supporting multiple model families does not, by itself, establish interchangeable privacy contracts or unlimited portability.
Persistent work changes what a successful assistant looks like
A chat answer can be judged when it appears. An ongoing work assignment needs to survive new information, interruptions, missing permissions, and changes in priority.
Consider a hypothetical launch coordinator. On Monday, it assembles a readiness brief. On Tuesday, a supplier changes a delivery date. On Wednesday, the team adjusts the release scope. A useful coordinator would update the same plan, identify the new dependency, and explain which earlier assumptions are now obsolete.
That example is an illustration of the work pattern, not a demonstration of Gemini. It shows why continuity matters. The system has to distinguish an old decision from the current decision and a draft proposal from an approved commitment.
Persistence also raises the cost of an unnoticed error. An incorrect deadline can become tomorrow’s reminder, next week’s status report, and the basis for another agent’s assignment. A reviewer needs a way to correct the original fact and verify that future work uses the correction.
Subagent coordination adds another dimension. Breaking a job into pieces can help when the pieces are independent. It can create extra work when several agents need the same context, make conflicting edits, or rely on an unverified output from an earlier step.
A reasonable design test is whether delegation improves the final result enough to justify the extra calls and coordination. Counting the number of subagents is a poor substitute for measuring completion quality.
Workspace gives Google a useful starting point
Google was already describing cross-application work before today. Its September 9 Workspace post covers creating Docs, Sheets, and Slides from contextual requests and reviewing an email before sending it from Docs. That earlier rollout has its own plan eligibility; it should not be treated as proof that today’s persistent coworker experience is included.
The practical attraction is the handoff between artifacts. A launch report may need a spreadsheet, a written brief, and a presentation. If those outputs refer to different source dates or definitions, the team still has to reconcile them manually.
For a Workspace-heavy organization, a useful evaluation would test whether a changed figure propagates coherently through the related work. Update a cost in the source sheet, request a revised deck, and check both the calculation and the explanation. Then ask the agent to identify what changed.
The most useful output may be the evidence attached to the work: the source documents, the formulas used, the unresolved questions, and the decisions that require a person. A polished deck with an unsupported number is still a bad deliverable.
Agent identities make responsibility easier to examine
An agent with an identifiable role creates an administrative object that a company can inspect. In a well-designed deployment, administrators should be able to determine who owns it, what it can access, and how to stop it.
Separate attribution matters when several people and automated processes touch the same document. If an incorrect change appears, investigators need to trace the originating task, the input, the tool call, and the permission used. A display name alone would not provide that evidence.
Identity also needs lifecycle management. A temporary research agent and an ongoing team coordinator should not automatically keep the same access forever. Organizations need an owner, an expiry or review policy, and a way to revoke access after the assignment ends.
For background on existing protections, Google’s August Workspace Studio security post describes approval controls, DLP rules, execution audit events, and revocation of OAuth scopes for flows. Those are Studio-specific documented controls; this article does not assume every setting maps unchanged onto the new agent.
A credible test should exercise the boundary. Remove permission to a folder during a task and verify that later steps cannot retrieve it. Put misleading instructions inside a source document and check whether the agent treats them as content rather than authority. Inspect the audit record after both tests.
Logs help explain an incident. They do not prove that the system prevents every incident. Useful governance requires both enforceable controls and evidence that those controls work in the configured deployment.
Spending caps need a precise definition
Google says a triggered project cap pauses the agent and can be resumed in the console. The details of enforcement deserve as much attention as the headline capability.
Its earlier August 27 billing announcement describes seat subscriptions, a consumption edition for select customers, pooled quotas, and spend guardrails. It also discusses commitment discounts and future deferred execution pricing. These are surrounding Gemini Enterprise billing options, not a confirmed complete price list for every feature announced today.
What the current billing documentation confirms
Google’s current quotas and overages documentation supplies concrete billing information for the Gemini Enterprise Pay-as-you-go edition. These figures describe that edition, rather than a flat price for the new persistent coworker.
| Billing item | Documented position |
|---|---|
| Account eligibility | An invoiced Cloud Billing account receiving an active monthly invoice is required. |
| Storage and data indexing | US$5 per GiB per month for the foundational engagement layer. |
| Assistant and no-code agent usage | Charged at the linked Agent Platform prices; the document does not give one universal per-task price. |
| Feature quotas | Pay-as-you-go has no feature quota limits; this does not establish exemption from technical service limits. |
| Seat-based overages | A single-user subscription quota does not restrict pay-as-you-go overages; enabled overages can continue up to a configured project monthly spend limit. |
The page also warns that customers who received a specific August 17 overage-billing notice must use its separate legacy documentation. Billing rules therefore need to be checked against the customer’s account, not inferred from a product headline.
Google’s edition comparison currently lists one minimum seat to get started with Pay-as-you-go and repeats the invoiced-account requirement. Its edition feature table does not establish that all newly announced coworker features are generally available.
There has also been a surrounding billing change: October 7 release notes describe enabling usage above quota for Frontline, EDU, and Emerging Market editions. That expands overage options; it does not supply a general availability date for the new universal agent.
For an ongoing agent, the full cost can include model calls, retries, sandbox execution, and connected services. An administrator needs to know which of those charges the cap covers, how quickly usage is counted, and what happens to work already in progress.
The pause behavior matters. Suppose an agent reaches its budget after drafting a report but before checking the totals. The interface should make the incomplete verification visible. If a user resumes it, the system should preserve the work already done and avoid repeating actions that have external effects.
Here is a small illustrative calculation, using invented inputs rather than Google prices:
| Hypothetical outcome | Total agent cost | Accepted deliverables | Cost per accepted deliverable |
|---|---|---|---|
| Lower-cost route, more failures | $100 | 20 | $5.00 |
| Higher-cost route, more successes | $180 | 60 | $3.00 |
The second route costs more in total and less per useful result. Human review would add another cost. If one route needs thirty minutes of correction per document, a lower token bill may have little business value.
The useful purchasing metric is cost per accepted outcome, measured with the review and retry policy your team will use.
How Gemini compares with other work agents
The competitor set already includes systems designed for delegated and recurring work. The table below compares documented product approaches. It is not a performance ranking, and availability must be checked separately for each deployment.
| Product or system | Documented work pattern | Main evaluation question |
|---|---|---|
| Google Gemini agent | Announced persistent, multi-model coworker design summarized above. | Does continuity hold across real projects within the configured boundaries? |
| ChatGPT Work | Files, plugins, finished artifacts, and scheduled work; cloud and local execution have different dependencies. | Can the selected environment reach the required resources reliably? |
| Claude Cowork | Multi-step work alongside the user across files and applications. | How much review and correction does the deliverable require? |
| Claude Managed Agents | Hosted autonomous workflows, long-running sessions, scoped tool permissions, credential vaults, and audit logs. | Can the team operate and inspect its custom workflow over time? |
| Microsoft Copilot Autopilot | Cloud-hosted teammate with identity, memory, and workspace; recurring work across the Microsoft environment. | Do its tenant context and administrative controls fit the process? |
OpenAI’s ChatGPT Work documentation distinguishes cloud work from tasks dependent on a connected computer. Local-resource steps still need that computer available. This makes runtime selection a meaningful part of a comparison: “continues working” needs to be tested with the actual files and tools the task requires. The documentation also describes recurring updates and plugin connections, so ongoing work is already part of the competitive landscape.
Anthropic’s financial-services agent announcement distinguishes Cowork plugins running alongside an analyst from Managed Agents running autonomously on the Claude Platform. It describes skills, connectors, and subagents in the templates, with human review before outputs are sent, filed, or acted upon. Buyers should compare the relevant execution model rather than treating every Claude product as the same assistant.
Microsoft’s September 25 Copilot announcement is especially close to Google’s direction. Autopilot is described as a persistent cloud teammate with its own identity and context. Microsoft said it was expanding to private preview at the end of September and described usage-based billing for long-running agentic capabilities. That establishes a close architectural competitor, not a verified head-to-head result or universal availability as of October 8.
The choice may turn on the work environment before it turns on the model. A team whose essential processes live in Microsoft 365 will have different integration questions from one centered on Workspace. A custom nightly finance workflow has different operating needs from an analyst asking for a one-off presentation.
What the benchmarks establish, and what they leave open
We did not find a reproducible, agent-specific head-to-head evaluation in the October 8 launch materials reviewed. That is a statement about the evidence checked for this article, not proof that no unpublished evaluation exists.
The follow-up review did find official evaluation tooling. Google’s Gen AI agent evaluation guide separates final-response quality from the trajectory of tool calls. It documents exact-match, in-order, and any-order trajectory checks, along with precision, recall, latency, and invocation-failure metrics for supported agent integrations. This provides a way to measure a workflow; it supplies no benchmark score for the October 8 Gemini coworker.
Google’s separate Managed Agents evaluation guide describes generating multi-turn scenarios, executing them, retaining conversation traces, and scoring final-response quality, safety, and multi-turn task success. That guide labels its offering Pre-GA with restrictions on production use and confidential data. Those restrictions belong to the offering described in that guide; they should not be generalized to every Gemini Enterprise feature. Neither guide establishes that Kingy has access to the announced coworker or has run these tests.
A model score cannot be transferred directly to this product. The agent adds routing, memory, retrieval, tool permissions, execution infrastructure, and a stopping policy. Those can improve a result, introduce failure modes, or change its cost.
Published research gives useful context for designing an evaluation:
| Benchmark | What it tests | Evidence and limitation |
|---|---|---|
| APEX-Agents | Long-horizon work across applications, created by banking, consulting, and legal professionals. | The February 2026 v3 abstract reports 480 tasks and a 24.0% Pass@1 result for Gemini 3 Flash with high thinking. Historical research configuration, not the October agent. |
| WorkArena | Enterprise web interactions on ServiceNow. | The 2024 paper introduces 33 tasks. Useful for enterprise actions; limited as evidence about multi-day coordination. |
| τ-bench | Tool use with simulated users and domain policies. | Tests interaction and policy-constrained outcomes. It does not reproduce a Workspace coworker deployment. |
| Hyper-tau-bench | Building an agent from business records, client requirements, and production APIs. | September 2026 research reports 53 tasks; Claude Opus 5 under Claude Code, its strongest tested configuration, passed 23.9% of evaluation simulations, versus 82.2% for an expert-authored reference. Different task and setup from Gemini coworker work. |
The APEX number should not be read as “Gemini completes only 24% of office work.” It measures the published tasks under the paper’s configuration. Nor should the hyper-tau result be used to rank Claude Cowork against Gemini: the evaluated job is constructing a customer-service agent, with its own harness and constraints.
These studies illustrate why finished-work claims require careful measurement. Generating plausible text, navigating an application, and producing an accepted business artifact are distinct targets.
Research also warns that the scoring mechanism itself needs scrutiny. The Agentic Benchmark Checklist paper identifies task and reward-design problems in existing benchmarks. For a buyer, that means checking what “success” verifies, whether external actions are inspected, and whether repeated attempts are included in the cost.
The evaluation a persistent coworker needs
Kingy’s proposed evaluation would start with a small set of realistic assignments, source files with known answers, and a clear acceptance rubric. The following tests are a proposed methodology. We have not run them on Gemini or its competitors.
| Test | Setup | Evidence to collect |
|---|---|---|
| Finished work | Prepare a brief, spreadsheet, and deck from the same evidence. | Accuracy, citations, formulas, consistency, review time. |
| Continuity | Change a deadline and a source figure after the first delivery. | Correct updates, retained decisions, explicit handling of stale facts. |
| Delegation | Give independent research steps and dependent calculation steps. | Coordination quality, conflicting edits, calls, elapsed time. |
| Access boundary | Revoke a folder permission while the task is paused. | Denied retrieval, safe continuation, attributable audit record. |
| Injection resistance | Place misleading action instructions in an input document. | Whether unauthorized actions are blocked and logged. |
| Budget exhaustion | Set a small cap partway through a multi-step job. | Pause timing, visible incomplete work, covered charges. |
| Recovery | Interrupt execution and later resume the assignment. | Preserved state, duplicate actions, completion and verification. |
| Model routing | Repeat the same assignment with allowed model routes controlled. | Quality, cost, latency, selected model and configuration. |
For a fair comparison, each system should receive equivalent information, equivalent action permissions, a comparable spending limit, and the same acceptance criteria. Product-specific advantages should be disclosed rather than quietly removed or added.
A native Workspace workflow and a browser-driven workflow may solve the same assignment in different ways. That difference is worth measuring, but the report should explain it. Otherwise, a reader cannot tell whether the result reflects the model, the connector, or the execution environment.
Repeated runs matter because an ongoing assignment will encounter variation. Report the fraction of accepted outcomes, the spread of completion times, and the cases where the agent needed intervention. Include failures in the cost calculation. Save enough execution evidence for another reviewer to understand what happened.
An attractive demonstration is useful for identifying a possible workflow. A purchasing decision needs measurements across representative tasks, including the awkward cases.
Availability and pricing remain the next evidence to watch
The product FAQ’s feature-specific early-access statement narrows the availability question, but “soon” leaves procurement teams without a rollout date. Existing Gemini Enterprise subscriptions, the Plus trial, and earlier Workspace features should not be assumed to grant every announced capability. The current billing documentation gives useful edition-level charges and prerequisites, while a complete cost and entitlement matrix for the new coworker remains unresolved.
The next documentation needs to answer several concrete questions:
- Which editions, regions, and administrator settings enable persistent agents and coworker identities?
- Which exact Gemini and Claude models can users select, and which routes can administrators permit or block?
- What are the task, concurrency, storage, retention, and execution limits?
- How are seats, model usage, sandbox time, and third-party actions billed?
- What happens to active work when a cap, permission revocation, or service failure interrupts execution?
- Which actions require human approval, and what can administrators inspect afterward?
Google’s announcement makes a credible strategic case for putting persistent work, multiple models, and enterprise controls in one environment. The product case will depend on whether a team can get access, understand the bill, and demonstrate that its recurring assignments finish correctly. Until those details are established, Kingy.ai’s position is announcement coverage with a defined evaluation agenda.
Update record, October 8, 2026: Added the live FAQ’s feature-specific early-access limits and model-access clarification, verified Pay-as-you-go charges and eligibility, the October 7 overage update, and official evaluation methods. These are newly verified additions to this article; some supporting documentation predates the agent announcement. No new reproducible agent-specific performance result or dated general rollout was established in this follow-up.
Source and testing note: All linked evidence is from vendor announcements, official product documentation, or original research. Competitor descriptions are documentation comparisons. Historical benchmark results are explicitly dated and are not scores for the newly announced Gemini agent. The cost example uses invented values. Kingy.ai performed no hands-on agent trial, paid model calls, or matched performance test for this article.
The Kingy Brief
Get The Kingy Brief.
AI changes, original tests and one practical thing to try. Fridays at 09:00 Vancouver time.
Free · Double opt-in · Unsubscribe anytime
Signup help and newsletter schedule
Signup form provided by Beehiiv. After submitting, check your inbox for "Confirm your subscription to The Kingy Brief" and open its confirmation link. Check Spam or Promotions if you cannot find it.
Fridays at 09:00 Vancouver time: source-checked AI changes, original tests and one practical thing to try. The weekly restart begins October 9, 2026. We skip a week when there is not enough verified material. Free. Unsubscribe anytime.
