OpenAI’s new personal agents can keep working between conversations. Here is what they can do, how access and costs work, and what the launch evaluations establish.
Published September 29, 2026. Updated September 29, 2026 at 2:36 p.m. PDT. Launch-day research based on linked documentation and published evaluations. The example briefs and cost scenarios below are illustrative; Kingy.ai has not run a hands-on dots benchmark.
OpenAI launched dots at DevDay on September 29, 2026. A dot is an individual personal agent inside ChatGPT, powered by GPT-6 Astra. It has its own cloud computer, uses connected tools, retains useful context, and keeps work moving between conversations. OpenAI’s launch announcement
The useful promise is continuity. You can hand over a responsibility, return later with a change, and review the work that accumulated in between. That could mean tracking an API migration, maintaining a research brief, or preparing material for a meeting. The practical questions are whether it finishes your particular work accurately, respects its scope, and saves enough review time to justify the cost.
Some answers are already concrete. Others remain open: permanent post-launch usage terms, measured everyday task costs, and independent end-to-end productivity results. This guide separates those categories so you can make a buying decision without treating a demo as a guarantee.
What is a dot, and how does it work?
The documented system combines Astra, a cloud computer and browser, connected apps, remembered context, and background work. It can research, analyze data, prepare documents, and build software while your own device is off. You can keep talking to it while it works. Meet dots
Think about a software migration. “Find every use of this retiring API” is one task. “Keep the migration moving until all services are off the API” is a responsibility: identify dependencies, coordinate fixes, verify progress, and return with decisions. A useful agent needs to carry the earlier findings forward as new information arrives.
A dot can divide work among parallel background agents, create visible cloud threads, and work through local Work or Codex tasks when a computer is connected. Its context includes relevant ChatGPT memory and its own notes; those notes are not a complete transcript. Proactive research uses restricted read-only tools, with follow-up actions subject to permissions. Tasks and memory
That distinction matters in a business workflow. Finding a stale claim in a proposal is one operation. Editing the shared document or emailing the customer is another. Write the expected deliverable and the allowed actions into the request, especially when a job runs for days.
Availability and setup
At launch, eligible adult Pro users have access outside the European Economic Area, Switzerland, and the UK. Business Premium users can access dots across supported ChatGPT regions. Enterprise access is a beta that administrators must enable; it starts disabled. The rollout may take several days. OpenAI Help Center, Access requirements
The Enterprise beta also covers Edu and Healthcare. Launch announcement
Create your dot in the desktop app or a desktop browser. Name it, review the proposed connections, and give it an initial responsibility. Continue in the mobile app when its supporting update is available; mobile web is unsupported. You can start with a narrow set of connections. Getting started
For a first job, choose something you already know how to check. A weekly change brief is easier to evaluate than a broad instruction to “handle my business.” Supply the source list, the deliverable, the deadline, and the decisions that need you. Give it feedback on the first result before adding more responsibilities.
Product specifications and practical limits
| Item | What is documented |
|---|---|
| Core model | GPT-6 Astra |
| Default execution | Separate cloud computer and browser |
| Work types | Research, analysis, documents, software and coordinated background tasks |
| Continuity | Relevant context carries forward between conversations |
| Local access | Optional; only one personal computer connected at a time |
| Local availability | Computer online with the ChatGPT app open |
| Cloud browser sessions | Separate from personal browser logins |
| Cloud hardware and storage | No specifications established in the checked launch documentation |
| Product context, concurrency and runtime ceilings | No numeric limits established in the checked launch documentation |
The first four rows summarize Meet dots. The connection details come from Computers and apps. Unspecified rows are gaps in the evidence, rather than claims that no limit exists.
You can inspect the cloud computer and take over its mouse and keyboard. Website sign-in uses a private form or takeover flow. A personal browser login does not sign the dot in, and some websites block cloud browsers or require verification. Local skills require a connected computer. Messaging, app access, and computer access are separate connections. Computers and apps
Astra’s API specifications are useful background, but an API context window does not establish how much information your dot retains or retrieves across months of work.
| Astra API specification | Published value |
|---|---|
| Model ID | gpt-6-astra |
| Context window | 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Modalities | Text input/output and image input; no native audio/video |
| Reasoning effort | low, medium, high, xhigh, max |
| Developer features | Streaming, function calling and Structured Outputs; fine-tuning unsupported |
These are Astra model specifications. Dots’ voice interface can combine additional capabilities; the base model’s modalities do not describe the whole product.
Messaging, voice and workplace connections
You reach the same dot through its available contact methods. Channels retain their own messages while relevant context can carry across them. You can start a voice call and type during it; assigned work can continue after the call ends. Adding a dot to Slack alone does not establish a monitoring schedule. Specify what to watch and confirm the recurring task. Messaging documentation
Treat access labels carefully. The enterprise administration guide describes Teams as an invite-only alpha. It also says only the owner can direct a dot through supported Slack interactions; another participant’s message does not itself start work. Dots administration
Texting, where enabled, is a limited beta for US Pro users, unavailable in Business and Enterprise. Message and data rates may apply. A dot cannot initiate calls or have its own standalone email address at launch, although you can connect a personal email account. OpenAI Help Center
OpenAI advertises connections to more than 4,000 apps through its plugin ecosystem. That is ecosystem reach, not a promise that every account has every integration or action enabled. Launch announcement
Examples OpenAI has published
OpenAI’s product page illustrates several ongoing jobs:
| Responsibility | Illustrated output |
|---|---|
| Investor update | Track usage and revenue, flag changes, refresh a presentation |
| Enterprise sales evaluation | Compare customer requirements with product documents and prepare an evaluation plan |
| API retirement | Map dependencies, prepare code and tests, track remaining calls |
| Product improvement | Use customer feedback to scope, build and test a change |
| Content production | Find interview moments and draft show notes and social posts |
These are OpenAI’s product examples, not Kingy.ai test results. The page does not supply a matched task dataset, completion-time distribution, or invoice-backed cost for each workflow.
The strongest demonstration for your own use case would include the original input, the resulting files or changes, a check against the input, and every correction a person had to make. A polished final slide by itself leaves most of that invisible.
Five practical task briefs you can adapt
The following prompts are original, untested examples. They specify what a useful result should look like. Adjust them to the connections and permissions your account supports.
1. Maintain a vendor-change brief
Every weekday at 8:30 a.m. America/Vancouver, review the official changelogs and pricing pages in the attached vendor list through October 9. Keep a dated change log with direct source links. Flag changes affecting our current plan, API behavior, or launch schedule. Put routine updates in this conversation; notify me immediately only for a breaking change or a decision due within two business days. Ask before contacting a vendor or changing an account. Confirm the saved schedule and end date.
The deliverable is a dated record of changes and their effect on your configuration. Check whether each alert identifies an actual change. Include one deliberately unchanged page in the source list to see whether the agent invents news.
2. Turn a bug report into a reviewable fix
Investigate the mobile checkout failure described in this issue. Use the specified repository and test environment. Reproduce the failure, prepare the smallest useful fix, and create a draft pull request with reproduction steps, test evidence and a short explanation. Keep desktop behavior intact. Ask before changing dependencies, committing money, deploying, or merging. Tell me if you cannot reproduce it.
A passing result needs a reproduction or a clear account of why reproduction failed. Run the relevant checks yourself and inspect the changed behavior. “Tests passed” should come with enough evidence to identify which tests ran.
3. Prepare for a customer meeting
Prepare a brief for my meeting with Acme on Thursday. Use the supplied notes, the connected calendar and the specified customer folder. Include the current goal, unresolved questions, commitments we have already made, and three decisions the meeting should settle. Link the source for each commitment. Draft an agenda for my review. Ask before sending messages or moving calendar events.
Check dates, owners, and promised deliverables against the originals. A plausible but invented commitment can cause more work than the preparation saves. A useful brief makes missing or conflicting information visible.
4. Convert an interview into an editorial package
Use this transcript to draft show notes, five clip candidates with timestamps, and three social posts in the attached house style. Distinguish direct quotations from paraphrases. Link each proposed factual claim to the transcript or a supplied source. Save the package for review and ask before publishing. Apply my edits consistently across the package.
For a test, pick a transcript you have already edited. Verify every quotation and timestamp. Count how many proposed clips survive review, rather than counting how many the agent produces.
5. Reconcile a recurring reporting spreadsheet
Update the weekly report from the two supplied exports. Preserve the formulas and column definitions in the template. Reconcile row counts and totals, identify duplicates, and list records that could not be matched. Save a new dated version plus a change log. Ask before overwriting the original or distributing the file. Stop and explain any mismatch you cannot resolve.
Check the arithmetic independently. Include a duplicate record and a missing identifier in the test data. The useful behavior is to flag both and explain their effect on the totals.
How much do dots cost?
There are three separate questions: the subscription needed for access, usage charged when work runs, and the value of the useful tasks that survive review.
| Eligible plan | Current published subscription price, USD |
|---|---|
| Pro 100 | $100/month |
| Pro 200 | $200/month |
| Pro 500 | $500/month |
| Business Premium | $125/user/month, or $100/user/month billed annually |
| Enterprise | Custom pricing |
Pro prices come from the updated Pro tiers FAQ. Business and Enterprise prices come from Business pricing. These are subscription prices, before applicable tax. Dots’ access documentation names all three Pro tiers for adults and Business Premium; Free, Go, Plus and Business Standard are not listed as eligible launch plans.
Your first dot is included in Pro or Business Premium at no extra cost. OpenAI plans additional dots and options to increase speed or monthly work; prices and timing are unannounced. Launch announcement
New Business workspaces require at least two paid seats, which can mix Standard and Premium. Standard costs $25 monthly or $20 per month billed annually. One Premium plus one Standard therefore totals $150/month with monthly billing, or $1,440/year with annual billing. This is our arithmetic for that seat mix, not a standalone dots price. Business signup requirements, Business pricing
The launch Help Center says dots usage will not count toward eligible plan allowances for the next month; later terms remain unannounced. The detailed dots docs separately say conversations do not count toward ChatGPT usage limits, while delegated Work/Codex tasks count normally. They describe expanded launch-month deeper-work allowances. The promotion’s coverage of every delegated task should not be assumed. Launch FAQ, Dots usage and access
No fixed dollars-per-dot-task tariff or permanent post-promotion task allowance was established in the checked sources. Buying a subscription also does not establish how many useful tasks you will finish. Failed runs, corrections, connected-service fees and human review belong in that calculation.
A transparent subscription cost calculation
Suppose you allocate a whole subscription to this workflow. Divide its monthly price by the number of completed tasks you accept:
| Accepted tasks per month | $100 subscription | $200 subscription | $500 subscription |
|---|---|---|---|
| 20 | $5.00/task | $10.00/task | $25.00/task |
| 100 | $1.00/task | $2.00/task | $5.00/task |
| 250 | $0.40/task | $0.80/task | $2.00/task |
These are arithmetic scenarios, not usage entitlements or measured dots results. They exclude additional usage, taxes, other services and review time. If you already pay for ChatGPT, the incremental subscription cost of trying an included feature may be zero; allocating the whole plan price to dots answers a different budgeting question.
Add human time for a more useful estimate. Assume a $200 subscription, 100 accepted tasks, five minutes of review per accepted task, and labor valued at $60/hour. Review costs $5 per task. Subscription allocation adds $2, giving $7 per accepted task before retries, setup and other charges. Those are declared assumptions, not a claim about dots speed.
If the same job takes you 20 minutes manually, its assumed labor cost is $20. That suggests a possible $13 saving. Change the review time to 18 minutes and the saving disappears. This is why a task’s correction burden matters as much as its apparent completion rate.
Astra API costs provide background, not a dots bill
Developers using Astra directly pay for tokens. Standard short-context rates are $10 per million uncached input tokens, $1 per million cached reads, $12.50 per million cache writes and $50 per million output tokens. API pricing
Above 272,000 input tokens, the entire request uses doubled input/cache rates and 1.5 times the output rate. Astra model page
| Hypothetical single Standard API request | Token-only calculation | Charge |
|---|---|---|
| 5,000 uncached input; 2,000 billable output | 0.005 × $10 + 0.002 × $50 | $0.15 |
| 50,000 uncached input; 20,000 billable output | 0.05 × $10 + 0.02 × $50 | $1.50 |
| 300,000 uncached input; 20,000 billable output | 0.3 × $20 + 0.02 × $75 | $7.50 |
Billable output includes reasoning tokens, which may never appear in the final answer. Reasoning documentation
The token counts here are invented inputs for transparent arithmetic. They are not measurements of a research brief, bug fix or dots task. The examples exclude tools, cache writes, multiple calls, retries and external services. An agent can make many model calls while completing one human request, so a single-request price cannot establish an end-to-end task price.
Benchmarks and evaluations
Read the evidence in three groups: Astra capability benchmarks, dots-specific safety tests, and your own workflow results. They measure different things.
Astra’s capability benchmarks
| OpenAI-reported benchmark | GPT-6 Astra | GPT-5.6 Sol | What it tests |
|---|---|---|---|
| Agents’ Last Exam | 59.3% | 53.6% | Professional tasks in real software |
| OSWorld 2.0 | 72.6% | 65.7% | Computer tasks; offline v2026.08.08 subset, partial score |
| ScreenSpot-Pro | 92.7% | 76.9% | Screen grounding without tools |
| AutomationBench | 41.4% | 18.1% | Professional automation |
| BrowseComp | 91.5% | 90.4% | Browsing research |
| Terminal-Bench 4.0 | 57.9% | 37.3% | Complex terminal tasks |
OpenAI reports maximum scores across reasoning efforts in research or API environments. Prompts and tools can differ from production ChatGPT. These are base-model results, not dots completion rates. Astra launch benchmarks
A benchmark score can help identify an underlying model’s strengths. It cannot tell you whether a connected calendar is configured correctly, whether a website will allow the cloud browser through, or whether a generated presentation matches your reporting definitions. Those dependencies are part of the product experience.
Dots-specific safety evaluations
OpenAI added a dots appendix to the Astra system card on September 29. Selected results:
| Test | Reported result and conditions |
|---|---|
| Bulk malicious-email exposure | Zero scored attack successes in 100 rollouts containing 50,000 emails, including 16,600 attacks |
| Iterative email attacks | Zero scored successes in 2,638 valid attempts across 100 attack chains; each candidate used fresh defender context |
| Changing scope and permissions | 91.8% alignment pass rate, 45/49 episodes; all 17 explicit permission-change cases passed |
| Boundaries across chained tasks | Moderate-severity flags rose from 8.6% with five intervening tasks to 19.7% with ten; no severe breach or exfiltration observed in this test |
These are OpenAI-run evaluations under specific conditions. Some tests isolate behavior without the full production safeguards. Zero observed successes does not establish immunity, and the boundary failures deserve attention. These numbers measure safety behavior, not everyday completion or uptime. Astra system card, dots appendix
For a buyer, the actionable implication is to make authorization changes explicit. When a project switches from preparing a draft to making a live change, state the new boundary. Revisit persistent responsibilities when their owners, data sources or intended audience change.
What remains unmeasured
The sources checked for this article do not establish an independent dots benchmark for accepted work per dollar, multi-day completion reliability, average correction time, or p95 task latency. Kingy.ai has not performed those tests. Launch examples and model scores should not fill those gaps by implication.
How to evaluate a dot on your own work
Start with ten historical tasks whose answers or acceptable outputs you already know. Give the dot the original inputs, without your completed solution. Use the same scope and grading rules for a manual baseline or another tool. Keep time spent setting up connections separate from time spent reviewing each result.
| Test case | Acceptance criterion |
|---|---|
| Research brief | Every material claim has a supporting source; stale information is identified |
| Spreadsheet reconciliation | Totals match; planted duplicates and missing records are flagged |
| Meeting preparation | Dates, attendees and commitments match the source records |
| Transcript package | Quotes and timestamps are exact; drafts remain unpublished |
| Software fix | Failure reproduced or inability explained; targeted behavior verified |
| Mid-task change | Revoked permission or revised scope is honored |
| Multi-day follow-up | New information is incorporated without redoing completed work |
| Blocked connection | Agent reports the blocker and required action accurately |
| Quiet monitoring | Unchanged sources do not produce unnecessary alerts |
| Context separation | Private information stays within the intended audience |
Score each result as accepted, accepted after correction, or rejected. Record elapsed time, review minutes, rework minutes, observed usage debit, additional fees and any action outside scope. Preserve the output that was graded. A vague overall satisfaction score makes it hard to diagnose failures.
Calculate cost per accepted task as allocated subscription cost plus metered extras, connected-service costs and human review/rework, divided by accepted tasks. Record rejected-task costs too. Compare that figure with the manual baseline. Run a few tasks again on another day to check consistency before relying on recurring work.
For long jobs, include a change halfway through: revoke a permission, revise a deadline, or correct an earlier assumption. Grade the subsequent actions. A persistent agent needs to respond to changed conditions as reliably as it handles the initial prompt.
Permissions, memory and stopping work
Automatic review checks consequential actions against instructions, permissions, custom rules and safeguards. Drafting does not authorize sending. Rules can require approval or hand a step to you, but do not grant app access or override safeguards. Open Activity to inspect delegated results and Scheduled to inspect recurring work. Controls documentation
Stopping has separate layers. Pause stops the main task; stop delegated tasks through Activity and cancel recurring runs through Scheduled. Completed actions are not undone. OpenAI says proactive research and private notes are not directly used for training; eligible conversation and task information follows ChatGPT data settings. Controls documentation
Disconnecting an app does not erase information already obtained. Resetting deletes the dot, its conversations, saved memories and scheduled tasks. OpenAI Help Center
Administrators have another important check: enterprise model defaults and controls do not apply to dots. The administration guide also restricts local access when the applicable residency policy requires it. Verify the actual workspace configuration before assuming a familiar policy covers the new feature. Dots administration
Specialist dots and what is still coming
OpenAI describes specialist dots for organizational responsibilities, including procurement, invoice processing, customer support and commercial contracting. These are focused enterprise pilots. Microsoft Agent 365 integration is described as a coming capability. Launch announcement
An enterprise pilot does not establish general self-service availability. Ask for the supported workflow, accountable owner, data boundaries, pricing and rollout status for the specific specialist you want.
When dots are worth trying
A dot is a sensible candidate when work recurs, context accumulates, and the output can be checked. Research monitoring, document preparation and a migration with reviewable changes fit that pattern. The likely value comes from reducing repeated coordination as well as producing the artifact.
For a short, isolated question, an ordinary conversation may be enough. For a narrowly specified coding task, directing Codex yourself may give you the control you need. Use a dot when you want an ongoing responsibility coordinated across tasks and tools. These are workflow recommendations, not comparative performance results.
Start with one responsibility and an acceptance standard. Track the useful results, the corrections, and the full cost through the launch period. Recheck OpenAI’s usage terms before budgeting for the following month.
Update log
- September 29, 2026, 2:36 p.m. PDT: Added included-dot and expansion details; clarified Edu/Healthcare beta access from the September 29 launch announcement.
