AI News

OpenAI Dot vs. Grok Bot vs. Meta Muse: Pricing, Specs, Benchmarks and the Evidence

Evidence checked September 29, 2026. Prices are US dollars unless stated otherwise. This is a comparison of published product documentation and evaluation reports, not a hands-on three-way performance test.

OpenAI’s Dot, Grok Bot and Meta’s Muse are personal agents that can continue working between messages. They connect to services, use computers, retain context and ask for permission when an action crosses a boundary. Their usefulness depends on the work they can finish, the access they need and the cost of checking their results.

The strongest initial choice depends on your existing subscriptions and workflows. Dot deserves a first trial if you already use ChatGPT and Codex extensively. Grok Bot deserves one if you want several specialized Bots and already pay for Cursor. Muse offers a free starting point and now has a substantial small-business connector offering. These are judgments about documented fit and entry cost. The public evidence does not establish an overall performance winner.

The names also need care. OpenAI calls the product dots, and an individual agent a dot; this comparison uses “Dot” for readability. Grok Bot is the computer-using agent accessed through the Cursor account system, with eligible Grok subscriptions linkable for access. Muse here means Meta’s personal agent, rather than an unrelated product with a similar name. The underlying model families and their APIs are separate products.

The comparison at a glance

Table 1. Product comparison
Dimension OpenAI Dot Grok Bot Meta Muse
Main documented design Persistent agent inside ChatGPT, connected to Work and Codex Persistent Bots with skills, routines and a shared account computer Personal agent with a VM, connectors and ongoing goals
Model disclosure GPT-6 Astra Cursor manages selection; no user model picker Muse Spark family; do not assume every task uses Spark 1.3
Entry route Eligible ChatGPT Pro or Business Premium; managed beta in other workspaces Paid Cursor, or an eligible linked Grok subscription Free limited tier; paid subscriptions
Published consumer prices relevant here Pro family: $100, $200 or $500/month Cursor Pro starts at $20/month; SuperGrok entry shown at $30/month Power $20/month; Maximum $100/month
Organization First personal dot; specialist dots are a separate preview Multiple Bots, conversations and shared files Ongoing personal context, tasks and artifacts
Local computer option Optional desktop access Desktop command execution and optional local network routing Mac computer use announced as available with permission
Important cost uncertainty Terms after introductory month Weekly usage allocation; on-demand spend can exceed a cap during a run Consumer “Muse tokens” are not an established API-token conversion
Main evaluation limitation Model scores do not establish Dot’s complete task success rate Model scores do not establish every managed Bot’s configuration Model/API evaluations do not establish the personal agent’s complete task success rate

Sources for the product identities and designs: OpenAI’s dots announcement, Grok Bot overview, Meta’s Muse announcement. Pricing, access and model qualifications are documented in the sections below.

What these agents can do, and what “always on” means

A persistent agent can retain a goal, revisit it after a scheduled trigger and use tools without a fresh prompt for every step. That makes it useful for recurring work: a weekly sales report, an inbox triage, a research watch or a project that needs several artifacts. It also creates more opportunities for stale assumptions and permission mistakes than a single answer does.

“Always on” should not be read as unlimited computation, guaranteed continuous execution or a promise that every task will finish. A service can pause work for quota, authentication, approvals, outages or missing information. None of the material reviewed establishes one common service-level guarantee across these three products. Treat completion notifications as a request to inspect a result, rather than proof that every requested external action succeeded.

Dot’s position inside ChatGPT

OpenAI identifies GPT-6 Astra as Dot’s model and describes a plugin ecosystem covering more than 4,000 apps, a cloud computer, Slack and Teams messaging, and voice interaction. The first personal dot precedes the specialist-dot rollout. OpenAI launch announcement

The practical attraction is continuity. A user who already gives ChatGPT research work and Codex engineering work can give Dot responsibility for coordinating related tasks. For example, a project brief might lead to source collection, a document draft and a coding task. That is a plausible workflow based on the product’s documented connections; it is not a measured success story from this review.

OpenAI’s Dot product page illustrates ongoing creative, business and technical work. Examples establish intended uses. They cannot tell you how often the product handles an unusual repository, a broken login flow or a contradictory brief correctly.

Grok Bot’s multiple-Bot design

Grok Bot supports multiple Bots that can work in parallel, exchange messages and use a persistent computer. Within one account, they share the computer’s files and signed-in apps; different users’ computers are isolated. This is a central architectural distinction. Grok Bot overview

You could create separate roles for reporting, research and engineering. Separate conversations make those roles easier to steer. Shared access also means role separation is weaker than creating three independently permissioned operating-system environments. A Bot name such as “Finance” does not, by itself, establish a finance-only data boundary.

The Bots documentation distinguishes configuration from accumulated context. Duplicating a Bot copies its profile, skills and routines rather than its learned memories or conversation history. Sharing a template likewise does not transfer the account’s computer or logins. Deleting a Bot leaves shared computer files and logins behind. These details matter for reuse, handoff and cleanup.

Muse’s personal and small-business scope

Meta introduced Muse as a personal agent with its own virtual machine, browser and connected services, accessible through its apps, web and WhatsApp. The launch describes a separate Sentinel that controls consequential interactions. Meta’s launch announcement

Its product introduction describes ongoing conversation, side chats, multiple tasks, memory and generated artifacts such as documents, PDFs and interactive work. That makes Muse relevant to administrative and creative work as well as personal errands.

On September 29, Meta announced Muse for Small Business. The documented connectors include Asana, Box, Canva, Dropbox, Figma, Granola, HighLevel, Intuit QuickBooks, Klaviyo, Lovable, Notion, Shopify, Slack, Stripe and Zoom, plus Facebook and Instagram business accounts. Meta says publishing, sending and spending require approval. This expands Muse’s documented business scope; it does not establish enterprise administration parity. Muse for Small Business

Pricing: subscription access, usage and the real bill

The first distinction is between a subscription that grants access, the amount of agent work included and any additional charges. A familiar chat-plan price does not establish an unlimited agent workload. An API’s price per million tokens does not establish the cost of using the corresponding consumer agent.

Current published entry prices

Table 2. Subscription prices and access
Product/access route Published monthly price What the evidence establishes What you should verify before buying
Dot through ChatGPT Pro $100 / $200 / $500 OpenAI documents three Pro tiers; Dot documentation identifies Pro eligibility broadly Your account’s rollout and entitlements, then post-introductory terms
Dot through Business Premium Workspace pricing Named eligible workspace tier Actual quoted seat price, billing basis and agent allowance
Dot in Enterprise/Edu/Healthcare Contract/workspace dependent Admin-enabled beta described in DevDay material Workspace availability, contract and enabled controls
Grok Bot through Cursor Pro From $20 Paid Cursor access; Pro listed at $20 Weekly Bot allowance and on-demand setting
Grok Bot through Cursor Teams From $40/user Teams pricing and access documented Shared budget configuration and administration
Grok Bot through linked SuperGrok Entry price shown at $30 Eligible personal subscriptions can be linked Linking eligibility and included Bot allocation
Muse Free $0 Limited free usage Actual account allowance and refresh terms
Muse Power $20 500 million Muse tokens/week Region, checkout price and account benefits
Muse Maximum $100 3 billion Muse tokens/week Same checks, plus whether higher allowance helps your workload

Sources: ChatGPT Pro tiers, ChatGPT pricing, DevDay recap, Cursor pricing, Grok Bot access and billing, Grok Bot pricing entry points, Muse subscriptions.

These are selected access routes, not an exhaustive list of every subscription upgrade. Taxes, currency conversion and app-store checkout can affect the bill. An unspecified workspace price is left unspecified rather than borrowed from a differently named plan.

Dot’s introductory month needs careful reading

OpenAI says the first dot is included in eligible plans without a separate fee. Its launch announcement distinguishes Dot conversations from delegated Work or Codex tasks: conversations do not consume ordinary ChatGPT usage, while delegated tasks count toward their usual allowances. The announcement also describes an expanded initial allowance. Dots launch

The current help page says Dot usage will not count toward eligible plan allowances for the next month and that later usage terms will be shared. Dot help

Taken together, these statements support an introductory benefit, with separate accounting still relevant for delegated work. They do not support “unlimited Dot forever,” or a claim that every downstream coding task is free. A buyer budgeting beyond October should leave later agent terms unresolved until OpenAI specifies them.

Pro 500’s higher allowance and Ultrafast access are also plan features. They do not prove a fixed speed or quota for every Dot task. OpenAI’s Pro-tier documentation says Ultrafast is exclusive to the $500 tier; buying credits on the $100 or $200 tier does not unlock it. It also documents a transition in allowances for existing Pro 200 users through October 29. Pro-tier help

Grok Bot’s on-demand setting can change the economics

The canonical billing help says paid Cursor plans include Grok Bot, and eligible individual Grok subscriptions can be linked. These include SuperGrok, SuperGrok Plus, SuperGrok Heavy and X Premium+; SuperGrok Lite is excluded. Included usage resets weekly and depends on agent work rather than a fixed message count. Cursor and linked Grok allowances do not stack.

On-demand use can continue after included usage runs out. Teams have it enabled by default. The documented monthly spending cap may be exceeded by a run already in progress, which is allowed to finish. Disable on-demand use if avoiding additional usage charges matters more than continuing a task. Grok Bot plans and billing

That last detail is material for procurement. A setting labeled “cap” should not automatically be modeled as an exact maximum liability. A long-running task can create a different exposure from a short request. Obtain the actual allocation shown in the account and establish a policy for long jobs rather than copying an unsupported number of “tasks per month” from a pricing roundup.

Muse publishes large allowances, but the unit needs context

Meta’s help center confirms a free limited tier, Power at $20 with 500 million Muse tokens weekly, and Maximum at $100 with 3 billion weekly. Benefits vary by region and account and are shown during onboarding. Subscriptions require the applicable adult age and a Meta account. Muse subscription help

Maximum provides six times the published weekly allowance for five times Power’s subscription price. That is arithmetic, not a finding that Maximum completes six times as many tasks. The reviewed documentation does not establish a conversion between these consumer Muse tokens, billable API tokens and completed tasks. Avoid valuing a Muse subscription by multiplying its allowance by the Meta Model API rate.

The free tier makes initial access cheaper, provided the account is eligible. A free attempt still costs setup and review time. A paid tier can be good value if it removes a quota bottleneck on useful work; it cannot repair a workflow that consistently produces wrong or unreviewable results.

Compare marginal cost and cost per accepted result

If you already pay for an eligible plan, the marginal subscription cost of trying its agent may be zero. If you need a new plan, charge the comparison with the new subscription cost. If other plan features are valuable to you, make that allocation explicit rather than assigning the entire bill to the agent or pretending the agent costs nothing.

A useful calculation is:

Cost per accepted result = (allocated subscription cost + extra usage + human setup, review and repair cost) ÷ accepted results.

For an illustrative month, a $20 plan producing ten accepted results has $2 of subscription cost per result before human time. If checking each result takes 15 minutes and that time is valued at $40/hour, review adds $10 per result. These are hypothetical assumptions, not measured figures for any of the products. They show why a lower entry price can coexist with a higher total cost.

Availability: country, plan and device are separate questions

Dot began rolling out on September 29. The help page excludes the EEA, Switzerland and UK from the initial Pro rollout, while Business Premium is described as available in supported ChatGPT regions. Enterprise beta requires admin enablement and starts disabled. Creation currently requires the desktop app or desktop web; mobile-app access follows availability, and mobile web is unsupported. Optional local access starts off. Texting is a limited US Pro beta. Dot setup and availability

Grok Bot is documented for Mac, Windows, Linux, iOS and Android. Its FAQ lists iOS/iPadOS 18 and Android 9 as minimum mobile versions. The August access expansion is historical context; use current billing documentation for current entitlements. An installable app does not prove that every country, organization or payment method is eligible.

Meta’s September 29 business announcement explicitly confirms Muse in the US and Canada. Regional Connect announcements contain additional geographic references, while subscription help says availability varies by region and account. The prudent conclusion is confirmed US/Canada access with further eligibility checked in the account, rather than a claim of worldwide availability. Business announcement, regional Connect announcement

Muse’s mobile, WhatsApp and web access was joined by a Mac desktop app. Meta’s localized Connect recap says permissioned Mac computer use is available. This makes descriptions of Muse as exclusively a phone or web agent outdated. It does not establish Windows or Linux desktop parity. Meta’s Connect recap

Available features versus announced features

Table 3. Available and announced features
Feature Evidence status on September 29
Dot’s first personal agent Rolling out; account access can lag announcement
Specialist dots with separate workplace identities and Microsoft Agent 365 collaboration Preview/pilot described by OpenAI; do not price it as a generally available personal feature
Grok Bot’s multiple Bots and routines Documented current product functionality
Grok Bot mobile push notifications Documentation says rollout may differ by account
Muse for Small Business connectors Announced as available today; inspect the app’s connector list
Muse on Meta AI glasses Announced for coming months
Muse’s own email address Announced future feature
Muse Charm Announced device; more details promised later in 2026
Muse Confidential VM Planned later in 2026; limited trusted testing described

Sources: OpenAI launch, Grok Bot notifications, Meta Connect news, Muse security design.

Roadmap features should contribute zero to a present-day task-completion score. They can matter strategically, but a buyer should record them as dependencies with uncertain delivery rather than silently including them in the current product.

Computer access, automation, memory and integrations

A computer gives an agent flexibility when there is no connector for a service. Browser interaction is also vulnerable to changed layouts, expired sessions, MFA and anti-bot checks. An API connector can be easier to audit because it exposes a specific operation, such as reading an invoice or creating a draft. It still needs correct permissions and correct arguments.

An integration count is therefore a weak buying metric. Ask whether the specific operation works: reading versus writing, drafts versus sending, your account tier, your file type and the exact service workspace. A vendor showing a familiar app icon does not establish that every action in that app is supported.

Local access expands what an agent can reach

Dot’s local work can use Work or Codex tasks, local skills and a local browser; Codex cloud environments must be created before delegation. Revoking local access stops further access, while disconnecting a connector does not erase information already obtained. Dot help

Grok Bot’s local command execution requires per-command approval by default and applies to the particular desktop. Optional local egress routes cloud-computer traffic through that desktop, exposing networks the desktop can reach; Enterprise admins can disable it. Cursor manages the model rather than providing a picker. Grok Bot settings

Muse’s documented Mac computer use likewise requires permission. Once enabled, computer access is a meaningful expansion of reach, separate from adding a cloud connector. Meta Connect recap

For all three, distinguish permission to read a project folder from permission to inspect an entire workstation. A local environment can contain private repositories, browser sessions and configuration secrets. Test the smallest useful scope first and verify the actual boundary. A natural-language role description alone does not establish an operating-system restriction.

Recurring work needs trigger and recovery semantics

Grok Bot documents skills, scheduled routines and event-driven automations. Browser teaching can record up to ten minutes, with no microphone audio; routines have a limit of 50 per Bot and 20 recent history entries. Event triggers rely on account integrations. Skills, routines and automations

Dot’s product material and Muse’s introduction both describe continuing work and recurring goals. Dot product page, Muse introduction

The harder evaluation is what happens after a failure. Does a routine retry an operation that already partly succeeded? Can it create duplicate invoices or reminders? Does it distinguish a failed notification from a failed underlying action? Record the trigger, timezone, retry behavior and last successful state. A recurring job that returns a confident summary while its source connection has expired is operationally worse than a visible failure.

Artifacts and portability

Grok Bot’s file documentation lists office documents, PDFs, code, data files, images, audio and video. Its desktop composer accepts six attachments at once, with 25 MB limits for documents, images and audio, and 200 MB for video. The shared /workspace supports handoffs between Bots. Files and results

For any agent, inspect the actual deliverable. A spreadsheet should preserve formulas and types. A PDF should open and contain the promised pages. A website should function beyond its screenshot. A slide deck should contain editable content if that was requested. File generation and file correctness are separate evaluation outcomes.

Exportability also deserves a trial. Keep copies of source inputs, accepted outputs, task instructions and permission decisions outside the agent account. That reduces the cost of switching services and makes a useful result independently inspectable. It does not require assuming that every vendor offers identical memory export or deletion controls.

Model specifications: useful context, separate from agent limits

The following numbers describe publicly documented model APIs. They are not promises about the context retained in a Dot, Grok Bot or Muse conversation, the memory of a long-running task, or the compute allocation of its cloud computer.

Table 4. API model specifications
API specification GPT-6 Astra Grok 4.7 Muse Spark 1.3
Documented context window 1,050,000 tokens 500,000 tokens 1,048,576 tokens
Maximum output in cited model documentation 128,000 tokens Not established here Not established here
API input modalities Text, images Text, images Text, images, video, PDF; audio has a support caveat
API output modality Text Text Text
Standard input price per million tokens $10 $2 $1.25
Cached input price per million $1 $0.50 $0.15
Output price per million $50 $6 $4.25
Consumer-agent linkage OpenAI identifies Astra as Dot’s model 4.7 trained for Bot use; managed selection does not pin every run Spark family identified; API 1.3 availability does not pin personal-agent routing

Sources: GPT-6 Astra model documentation, Grok 4.7 model documentation, Meta model documentation, Meta API pricing.

Meta explicitly says Spark 1.3 audio understanding is not fully supported and response quality may degrade. It recommends Spark 1.2 or Muse Voice Transcribe for audio. This caveat belongs beside the modality list. Meta model documentation

Astra’s documentation adds long-context pricing above 272,000 input tokens and separate speed/batch arrangements. Grok documents higher-context pricing above 200,000 input tokens. Meta’s standard Spark pricing states there is no long-context premium. Check the actual endpoint and mode before budgeting a developer workload. Astra pricing, Grok model pricing, Meta pricing

Meta also offers a Contributor tier, with lower rates in exchange for using prompts and completions for training. The published Spark 1.2/1.3 Contributor rates are $0.10 input, $0.002 cached input and $0.20 output per million tokens. Comparing those prices with another provider’s standard tier without disclosing the data-use tradeoff would be misleading. Meta API pricing

A large context window means a request can contain substantial material. It does not establish perfect retrieval, accurate reconciliation of conflicting sources or unlimited durable memory. Tool output, reasoning and system instructions consume resources too. Long tasks can compress or select context in product-specific ways that these API tables do not disclose.

Voice in an agent application does not contradict text-only model output in this table. Speech recognition, speech synthesis and real-time interaction can use other components. Likewise, an agent can work with a video through tools even when the named model endpoint’s native input types differ.

The reviewed material does not provide an apples-to-apples product specification for model parameter counts, training compute, cloud VM CPU/RAM/storage, maximum task duration, concurrent jobs or sustained task throughput. Those remain unknown here. “Not established” means insufficient evidence, rather than a missing capability or a zero score.

Benchmarks and evals: what the numbers establish

There are three levels of evidence: model capability evaluations, agent-system evaluations and completed-work measurements. A good model score can support an agent’s potential. A system evaluation includes permissions and tools. A real workflow measurement includes authentication, source quality, review time, failures and the bill.

No common, independently reproduced end-to-end evaluation of these three shipping agent products was established by the sources reviewed for this article. Accordingly, there is no defensible three-way accuracy percentage, speed league table or overall numerical ranking here.

GPT-6 Astra’s published capability results

OpenAI reports 72.6% on OSWorld 2 for Astra in its latency simulation, versus 65.7% for GPT-5.6 Sol, with approximately 40 versus 75 minutes. These are vendor-reported model/computer-use evaluation results under the stated setup. GPT-6 Astra announcement

Other selected Astra results in the same announcement are:

Table 5. Selected Astra capability results
Evaluation Published Astra result
Agents’ Last Exam 59.3%
AutomationBench 41.4%
BrowseComp 91.5%
Terminal-Bench 4.0 57.9%
DeepSWE v1.1 74.1%
MRCR v2, eight needles, 512K–1M context 96.3%

OpenAI says scores use maximum performance across efforts, with research/API configurations that can differ from production prompts and tools. Astra results and evaluation notes

The score is useful evidence that Astra can perform computer work in a controlled evaluation. It does not mean Dot completes 72.6% of a buyer’s everyday tasks, nor that every Dot task takes 40 minutes. Task mix, tools, partial-credit scoring, reasoning effort and latency modeling must accompany the number.

OpenAI’s later model announcement explains that its OSWorld 2 evaluation uses an offline task version and partial reward, and that other evaluations use different tool environments. The same announcement concerns GPT-6.1 Sol; its scores should not be relabeled as Dot/Astra scores. GPT-6.1 Sol evaluation notes

Dot-specific safety evaluations

OpenAI’s Dot appendix reports the following synthetic/adversarial outcomes:

Table 6. Dot safety evaluations
Dot evaluation Reported outcome
Bulk email-injection test 100 rollouts; 50,000 emails, including 16,600 attacks; zero scored successes
Iterative injection test 100 attack chains, 2,638 valid attempts; zero scored successes
Scope/permission adherence 45/49 cases, 91.8%; all 17 explicit permission-change cases passed
Challenging selected misalignment scenarios 0.84% severe misalignment
Longer chained tasks Moderate violations rose from 8.6% after five intervening tasks to 19.7% after ten

These are self-reported evaluations, not production incident rates. Zero observed success does not establish immunity. The chained-task result identifies a meaningful constraint-retention concern. Dot system-card appendix

For a buyer, the most relevant implication is to test instructions after several task switches and interruptions. A product may honor a clear rule immediately yet become less reliable after lengthy work. A short demo cannot establish durable permission discipline.

Grok 4.7’s published capability results

The September 21 model announcement reports these scores for Grok 4.7, generally at xHigh effort; the DeepSWE entry is explicitly at high effort:

Table 7. Selected Grok 4.7 results
Evaluation Published Grok 4.7 result
CursorBench 4.0 46.3%
DeepSWE v1.1 71.0%, high effort
EEBench 64.0%
AA Briefcase v1.1 1,657
Terminal-Bench 4.0 37.6%
Harvey Legal Agent Benchmark 19.6%
HealthBench Professional 56.7%

The source also reports 62.4% on LatchBio’s biosafety benchmark and a 3.3% risky-allowed rate on HackerBench v0.3. These measure different concerns from Dot’s email-injection tests. Grok 4.7 was trained for the Grok Bot environment, but these remain model evaluations. Preserve the footnotes and settings. Grok 4.7 announcement

A score of 1,657 is not 1,657 percent. Percentages across coding, engineering and legal evaluations also do not share a denominator. Averaging these entries would produce a number with no coherent meaning.

The model-selection documentation matters here: a managed service can change routing. A subscription name or the newest model announcement cannot, alone, prove which configuration handled a particular task. Record product and app versions where available, and distinguish documented model availability from verified per-run provenance.

Muse Spark 1.3’s evaluation evidence

Meta reports improved agentic and coding behavior for Spark 1.3. Its internal engineer comparisons with Spark 1.2 used about 20% fewer tool calls and 25% fewer tokens. Spark 1.3 is available through Muse Code and the Meta Model API. These are vendor-reported comparisons, not a measured personal-Muse saving against Dot or Grok Bot. Muse Spark 1.3 announcement

Meta’s methodology report describes Spark 1.3 at max reasoning against Spark 1.2, Opus 5 and GPT-5.6 Sol, with specified efforts. It covers GDPval-AA v2’s 220 tasks, OSWorld 2’s 108 workflows and partial scoring, DeepSearchQA’s 900 questions, AutomationBench’s 600 simulated tasks, and DeepSWE v1.1’s 113 tasks. OSWorld task versions differ for Spark 1.2; third-party configurations were best-effort rather than fully optimized. Meta evaluation methodology

That methodological disclosure helps readers interpret the associated results. It also prevents a clean current-model conclusion: its listed baselines are not Astra and Grok 4.7 in identical shipping-agent environments. Terminal-Bench versions and automation task versions should be matched before comparing scores from different releases.

Customer stories are another category of evidence

The vendor’s Grok Bot support case study describes handling a 175% increase in tickets without adding headcount and cites resolution costs as low as $0.20–$0.30. That is a vendor-presented operational case, with workload and accounting assumptions that a reader should inspect. It is not an audited universal cost per resolution or a controlled three-way study. Grok Bot support case study

Customer stories can reveal a useful workflow. They do not establish that a different business can reproduce the outcome. The same standard applies to demonstrations and partner testimonials from OpenAI and Meta. Evidence-based coverage labels the evidence class instead of treating an attractive example as a general rate of success.

Security, permissions and privacy

Agent security has several distinct questions: can an attacker redirect the model; can redirected behavior execute; can it expose secrets or private data; can the user understand and revoke permission; and what does the provider retain or use? A single safety percentage cannot answer all of them.

Dot’s safeguards and data policy

OpenAI describes an isolated Linux/Chrome cloud environment and a separation between code execution and safeguard controls. Secure sign-in pauses model interaction and submits credentials without exposing them to the model. Proactive research is technically constrained to read-only activity. Consumer training depends on data settings; Business, Enterprise and Edu data is not used for training by default. Sensitive actions have confirmation requirements, including saved-card purchases and permanent deletion. Dot safety, security and privacy

The practical benefit of a read-only background boundary is that an unsolicited research step has less opportunity to alter connected systems. It still accesses information and can form misleading conclusions. Read-only access also remains sensitive when the sources contain personal or confidential data.

Secure credential entry protects the credentials handled through that flow. A password pasted into an ordinary document or conversation is a different case. More generally, protecting an authentication token does not prevent an authorized session from being used to read or send information if the permitted action is too broad.

Grok Bot’s approvals and shared state

Grok Bot documents model-based Auto Review, user rules and admin-enforced rules; “Ask first” wins when rules conflict. Secure secret entry excludes the secret from model context. Privacy Mode and training choices are tied to the applicable Cursor account setup; legacy usage-based configurations have restrictions. Approvals, security and privacy

The shared computer is particularly important for business roles. Separate Bot conversations can share files and sessions. Use a design in which that sharing is acceptable, or obtain a stronger boundary through separate accounts/environments and explicit administration. Confirm which controls exist in the purchased tier instead of assuming a personal rule provides an enterprise policy.

Local egress deserves an explicit test because it can connect the cloud computer to resources reachable from a workstation. A company may want that convenience, but should know which destinations, authentication states and logs are involved. Default local command approval and a separate network-route setting are different controls.

Muse’s Sentinel and the limits of current confidentiality

Meta describes a runtime cell isolated from host safeguards. Sentinel controls connector actions and network egress outside that cell; real credentials replace surrogate tokens at the approved network boundary. Current operational policies restrict staff access but do not cryptographically prevent Meta access needed to operate, secure or support the service. Confidential VM is a future capability under limited testing.

Meta says conversations and VM data are not shared with its ad systems. Browsing performed as the user can nevertheless indirectly affect advertising. Sanitized inference trajectories can be used for training by default, with a settings opt-out. Muse security design

These distinctions are essential for an objective privacy comparison. A VM is an isolation mechanism. A confidential-computing promise describes a stronger, different provider-access boundary. A no-ad-system-sharing statement concerns a particular data flow. A training opt-out concerns another. Combining them into “Meta cannot see or use anything” would misstate the documentation.

Deleting an agent is not the same as reversing its work

Once an agent sends a message, creates a remote record or publishes a page, deleting the agent cannot automatically undo that action. Disconnecting access cannot make a recipient forget a message. A local reset also does not establish deletion from every connected service, backup or audit record.

Before recurring business use, document four separate procedures: pause future work, revoke credentials, remove stored agent data and undo external changes where possible. Test each with harmless records. This is an operational evaluation proposal, not a claim that the three products share identical deletion semantics.

Business and enterprise readiness

Grok Bot documents Teams and Enterprise controls, with more granular managed setup, network and audit options in Enterprise. Admin enablement and legacy-plan conditions matter. Its team and enterprise documentation should be read alongside the contract and tier-specific configuration.

OpenAI’s DevDay recap identifies Dot access in Business Premium and admin-enabled beta access for Enterprise, Edu and Healthcare. Specialist dots with their own workplace identities are an additional preview, rather than evidence that every personal Dot already has independently provisioned corporate credentials. DevDay recap

Muse’s new business connectors provide concrete workflow coverage. An accountant could inspect draft expense exceptions; a shop owner could review a campaign built from store and social data. Those are candidate uses for evaluation. The connector announcement does not, by itself, establish enterprise SSO, SCIM, retention policy, audit export or data-residency guarantees for every customer.

For procurement, ask for product-specific answers on identity, least-privilege access, logging, retention, training, residency, incident handling and support. A provider’s certification or API policy should not automatically be transferred to a different consumer product. The relevant question is which service and data flow the assurance actually covers.

There is also a responsibility question. If the agent prepares an accounting exception, a person can validate it against a ledger. If it changes the ledger, reconciliation and undo become necessary. If it changes an ad campaign, spend controls matter. Select the initial task according to how cheaply its outcome can be checked and repaired.

Which agent fits which buyer?

Table 8. Candidate selection by workflow
Buyer or task Most reasonable first candidate Evidence-based reason What could change the choice
Existing ChatGPT/Codex user coordinating research and engineering Dot Direct workflow relationship and named Astra model Rollout, later allowance and delegated-task cost
Existing Cursor user wanting several persistent work roles Grok Bot Included access and explicit multiple-Bot organization Shared sessions, billing settings and task accuracy
New user wanting a low-cost initial personal-agent trial Muse Documented free limited tier Region, quota and needed connectors
Small business using Shopify, QuickBooks, Canva and Meta business accounts Muse Specific connectors in September 29 announcement Exact supported operations, approval behavior and administration
Organization needing documented managed Bot network/setup controls Grok Bot Enterprise as a candidate Product-specific Enterprise documentation Contract scope, audit requirements and competing workspace controls
User outside an eligible consumer region Whichever account is verifiably eligible Access is a prerequisite Rollout or plan changes
Buyer prioritizing provider-inaccessible VM data today No winner established here Muse’s stronger Confidential VM is still a future rollout Product-specific released architecture and independent verification

These recommendations identify where to start a trial. They do not assert that one agent writes better reports, codes faster or makes fewer mistakes on your inputs. A connector advantage can disappear if the relevant operation is missing; an existing subscription advantage can disappear if review time dominates.

For coding, give special weight to reproducible output: repository changes, test results and an explanation of unresolved failures. For research, require dated primary sources and inspect whether each link supports the associated claim. For shopping, compare total delivered price and approval behavior, using a test that stops before payment. For administrative work, inspect the source record and the artifact rather than accepting the completion message.

A practical head-to-head evaluation you can run

The following is a proposed protocol. These tests were not run for this article, and the table contains no observed product results. Start with non-sensitive sample data and a few representative tasks; expand only when the initial results justify it.

Give each product the same inputs, acceptance criteria and permissions. Record the subscription, date, app version, relevant settings and available model identification. Use a fresh task context for direct comparisons, then a separate continuing context for memory tests. Do not let one product receive corrections that the others never see.

Table 9. Proposed evaluation protocol
Test Example input Acceptance criterion Useful measurement
Current research Compare three products using dated official sources Every material claim has a supporting link; unknowns are explicit Unsupported claims, source coverage, review minutes
Document production Source notes plus a required report outline Editable report preserves required facts and opens correctly Missing requirements, formatting repair
Spreadsheet reconciliation Synthetic transactions and a policy Correct totals, exceptions and traceable formulas Cell errors, false exceptions, accepted output
Coding repair Small repository with a reproducible failing test Fix addresses the failure without unrelated changes Test result, regression checks, human repair time
Browser workflow Test account with a draft record to create Correct fields saved once; nothing sent Completion, duplicates, authentication interruptions
Scheduled work One harmless recurring check Runs at expected time and reports source failure honestly Trigger reliability, duplicate actions
Permission retention Read-only rule, followed by interruptions and task changes Stops before an unapproved write Violations and clarity of approval requests
Untrusted-content handling Synthetic document containing a request to ignore the task Treats document instructions as untrusted; no external disclosure Boundary violations and task completion
Budget behavior Small task near an included-usage threshold Reports limits and follows the account’s configured spending behavior Extra charge, pause state, cap interpretation
Cleanup and recovery Harmless files, sessions and one external draft Verified pause/revocation and clearly identified residual data Cleanup steps, recoverability, external state

Predefine what counts as success. A beautifully written report with fabricated citations fails the research task. A correct answer delivered after exceeding an explicit spending limit fails that constraint. A coding fix that passes the target test but removes unrelated functionality needs further inspection.

Measure success rate, wall-clock time, human review/repair time and cost per accepted result separately. Also record severe failures individually. A weighted average can hide an unauthorized disclosure behind several successful formatting tasks, so do not let a convenience score erase a permission violation.

Repeated runs matter because a single success can be luck. Even a small number of repeats will reveal some variability, though it cannot establish a dependable population failure rate. Record the sample size and avoid describing a handful of runs as a comprehensive benchmark.

For a fair speed comparison, separate active computation from waiting for the user, login and queued service time. Both can matter to the buyer, but they answer different questions. For a fair cost comparison, include the same review standard for every product. Letting one product submit unchecked work gives it an artificial price advantage.

Finally, test a continuation: change one requirement after the first artifact, interrupt with another request, then return to the original task. Persistent agents are sold for that continuity. Evaluating only isolated first-turn tasks leaves their central behavior unmeasured.

What remains unknown, and what would justify a stronger verdict

The most important unresolved items are Dot’s later usage terms, exact per-account Grok Bot allocations and managed routing, consumer-token accounting for Muse, and consistent product-level task and concurrency limits. Public model prices cannot fill those gaps.

A stronger comparative verdict would require the three shipping products to complete the same tasks with comparable permissions, inputs, resource budgets and review standards. The report would need dated configurations, failed runs, output artifacts, billed costs and enough repeats to show variability. Independent reproduction would strengthen the evidence further.

The available documentation already supports a buying sequence. Start with the eligible agent attached to a subscription you use, unless another product has a necessary connector or control. Evaluate a task whose result you can verify cheaply. Grant additional access only when the measured benefit warrants it, and reassess introductory pricing when the terms change.

For related product detail, see Kingy AI’s OpenAI Dots guide and Grok Bot vs. OpenAI Dots buying guide. This three-way comparison adds Muse, its current subscription allowances and the September 29 business expansion; it preserves the distinction between published capability evidence and results that still need testing.

Source and methodology notes

The research prioritizes official product announcements, help pages, developer documentation, security descriptions and evaluation methodology. Vendor-reported scores and customer stories are attributed as such. Recommendations and the proposed test protocol are editorial analysis. No first-hand testing experience, independent audit or completed three-way benchmark is claimed.

The primary sources are linked beside the claims they support. Especially useful starting points are Dot setup, Dot safety, Grok Bot billing, Grok Bot security, Muse subscriptions, and Muse security.

This article is dated because rollout and usage terms are changing. A later announcement should update the relevant row and its source, rather than silently turning an earlier preview into an earlier available feature. Before subscribing, check the current checkout and the entitlements actually shown in your account.