Which AI coding agent is best in 2026?
Evidence cutoff: 17 August 2026
Edition: 1.1 — weekly maintained
Comparison lane: Grok Build terminal/headless, Codex local/CLI/IDE, and Claude Code terminal/IDE. Hosted and persistent-agent modes are scoped separately and are never merged into the local-product results.
Evidence labels: Official specification · Vendor-reported evaluation · Independent evaluation · Controlled original result · Inference
Bottom line: There is no universal winner, but there are defensible category leaders. Claude Code + Opus 5 is the point-score leader on the strongest current native-product test of ultra-long autonomous engineering (SWE-Marathon v1.1: 50.0% versus Codex + GPT-5.6 Sol at 42.5% and Grok Build + Grok 4.6 at 31.9%). The benchmark has only 20 distinct task clusters and does not establish statistical separation between configurations. In the common mini-swe-agent harness, Opus 5 and GPT-5.6 Sol have overlapping reported intervals; Sol is one point lower but markedly more efficient. Grok 4.6 leads CursorBench 3.2’s point score and cost efficiency in a shared IDE-style harness, although public data are insufficient to determine whether the small top-score gaps are meaningful. For most teams seeking the strongest all-round product and safest local default, Codex remains the medium-confidence overall recommendation; Claude Code is the terminal, provider-flexibility, and native long-horizon point leader; Grok Build is the value, customization, and shared-harness IDE pick.
Executive verdict
The useful question is not “Which brand won a benchmark?” It is “Which agent, on which surface, with which resolved model and permissions, is most likely to complete my repository task correctly, safely, quickly, and at an acceptable cost?” Once the comparison is normalized that way, the market looks less like a podium and more like three distinct operating systems for agentic software work.
Codex wins the overall recommendation. Its advantage is breadth plus containment: desktop, web, CLI, IDE, local and cloud workflows; a current family of fast-to-frontier models; native subagents and scheduled work; and an OS sandbox with workspace-limited writes and network disabled by default. Its main disadvantages are complexity, variable subscription limits, expensive top-tier API output, and the fact that the exact model used by some cloud tasks is not user-selectable. The overall verdict is medium confidence because current public coding benchmarks do not isolate the Codex product from the underlying model and harness.
Claude Code wins for terminal-native control and deployment flexibility, and it has the strongest current native-product long-horizon point result. It has mature project instructions, hooks, skills, plugins, subagents, worktrees, non-interactive execution, strong permission rules, excellent OpenTelemetry detail, and support for Anthropic’s API plus Bedrock, Vertex AI, Microsoft Foundry, and gateways. On SWE-Marathon v1.1, Claude Code + Opus 5 Max resolves 50.0% of 20 distinct ultra-long tasks across eight attempts per configuration, ahead of Codex + GPT-5.6 Sol Max at 42.5% and Grok Build + Grok 4.6 High at 31.9%. Those repeated trials do not turn 20 task clusters into 160 independent tasks, and the owner does not report a paired significance test. Claude Code’s most important product caveat is that native filesystem/network sandboxing is available but not enabled by default. Its default model also depends on plan and provider, so “Claude Code” is not one reproducible system unless the resolved model ID is captured.
Grok Build wins for open-ended agent experimentation, frontier-price value, and shared-harness IDE cost/performance. It exposes interactive and headless modes, ACP, custom models/providers, skills, plugins, hooks, MCP, subagents, worktrees, scheduled loops, and compatibility with common agent-instruction conventions. Grok 4.6 Extra High has CursorBench 3.2’s highest point score, 70.8%, at a reported $2.81 per task; Fable 5, Opus 5, and Grok 4.6 form a close point-score group, but Grok’s reported cost is dramatically lower on that board. xAI’s 12 August 2026 Grok 4.6 model card explicitly identifies Grok 4.6 as Grok Build’s default. Older Build documentation still names Grok 4.5, so reproducible testing must record the runtime-resolved model rather than trust a mutable alias. Build’s sandbox is off by default.
Scroll or swipe to compare →
| Decision | Winner | Confidence | Why |
|---|---|---|---|
| Best overall for most professional teams | Codex | Medium | Broadest coherent local/cloud product, strong current model ladder, safest documented local default. |
| Best terminal-first engineering workflow | Claude Code | Medium | Mature CLI ergonomics, provider choice, granular controls, strong observability. |
| Best out-of-box local containment | Codex | High | OS sandbox, workspace write boundary, and network-off default. |
| Best provider and data-plane flexibility | Claude Code | High | Anthropic, Bedrock, Vertex AI, Foundry, and gateways are officially supported. |
| Best custom-model experimentation | Grok Build | High | Custom model/provider support plus ACP and headless operation are first-class. |
| Best documented agent-native telemetry | Claude Code | Medium-high | Detailed opt-in OTel events for prompts, tools, permissions, edits, and MCP. |
| Best low-cost frontier API option in this snapshot | Grok 4.6 | Medium | $2 input / $6 output per MTok, but this is a model-price result, not a complete Build cost. |
| Best native-product ultra-long autonomous point result | Claude Code + Opus 5 | Medium | 50.0% on SWE-Marathon v1.1; only 20 task clusters, native stacks differ, and statistical separation was not established. |
| Best shared-harness IDE point score/value | Grok 4.6 | Medium | CursorBench 3.2: 70.8% at $2.81/task; top scores lack public CIs. |
| Best shared-harness long-repo efficiency near the top | Codex GPT-5.6 Sol | Medium-high | DeepSWE: 73% versus Opus 74%, with about half the output tokens and lower cost. |
Contents
- Scope: Grok Bot, products, and model results
- Product profiles
- Specification comparison
- Documented workflow fit
- Public benchmark evidence
- Fair-test summary
- Cost and value
- Security and procurement
- Buyer recommendations
- Decision tree
- Limitations
- Full evaluation protocol
- FAQ
- Sources and evidence ledger
Read this before comparing scores
Grok Bot, Grok Build, and persistent cloud agents
If you searched for “Grok Bot vs. Codex vs. Claude Code,” the coding product most people intend is Grok Build, not literal Grok Bot. Grok Build is xAI’s terminal/headless coding agent. Grok Bot is a separate persistent cloud-agent product: each named Bot has durable state and uses a persistent user-scoped cloud computer with a browser, filesystem, and terminal. Multiple Bots share that user-scoped computer. The products are related, but they are not interchangeable experimental units.
Scroll or swipe to compare →
| Surface | Persistence | Environment and tools | Comparison boundary |
|---|---|---|---|
| Grok Bot | Named agent with durable memory, files, browser sessions, and preferences | Shared user-scoped cloud computer with browser, filesystem, and terminal | Default model, plan entitlement, compute allocation, and quotas were not publicly disclosed at the cutoff. |
| Codex cloud and Work | Managed tasks and continuing work across Codex surfaces | OpenAI-managed environments plus web, desktop, CLI, IDE, and remote workflows | Not documented as the same named-agent/shared-computer persistence model as Grok Bot. |
| Claude Code web/cloud | Cloud sessions can continue after disconnection; remote monitoring/control varies by surface | Anthropic-managed or self-hosted environments, with terminal, IDE, web, CI, and mobile control | Plan, provider, runner, retention, and resolved model must be recorded; persistence is not assumed equivalent to Grok Bot. |
The main comparison therefore uses Grok Build terminal/headless, Codex local/CLI/IDE, and Claude Code terminal/IDE. Persistent and hosted modes are discussed as a separate product-design lane, not averaged into the benchmark verdicts.
Model results are not product results
A model answering a patch-generation benchmark through a minimalist harness is not the same system as an interactive agent reading repository instructions, selecting files, running tests, asking for approval, compacting context, and revising a patch. Agent performance is a joint result of model capability, prompt, tool schema, scaffolding, environment, budget, retries, product defaults, and benchmark validity.
This guide therefore separates five evidence classes:
- Official specification: a vendor’s documentation of a feature, limit, price, or policy.
- Vendor-reported evaluation: a vendor’s benchmark number, useful but not independent.
- Independent evaluation: a result produced by a benchmark owner or third party with enough configuration detail to audit.
- Controlled original result: a run produced under the protocol in this guide. This edition contains no paid three-way original run because account access and paid testing were not authorized.
- Inference: an explicitly labeled analyst conclusion drawn from the cited specifications or evaluations, not a vendor fact or measured Kingy.ai result.
An evidence label is not an insult. Official documentation is the best source for supported features and contractual controls; it is not independent proof of task quality. A benchmark owner can report a careful result; it still may not represent an IDE product with different tools. The goal is traceability, not false symmetry.
Product profiles
Grok Build
Grok Build is an interactive terminal agent, a headless CLI, and an Agent Client Protocol endpoint. Headless execution supports JSON and streaming JSON, resumable sessions, and scripting. Official documentation covers macOS, Linux, WSL, and native Windows PowerShell installation. The product supports browser or device authentication, enterprise OIDC, external token providers, and API keys.
Its workflow surface is already substantial: project and user skills, plugins and marketplaces, hooks, MCP, custom agents, parallel subagents, worktrees, background tasks, scheduled loops, session import/export, cross-session memory, plan mode, permissions, sandbox profiles, AGENTS.md compatibility, Claude Code configuration compatibility, and custom model/provider support.
The current default is Grok 4.6. xAI’s dated Grok 4.6 model card explicitly calls it the default in Grok Build. Older Build overview and Grok 4.5 documentation still identify 4.5 as the Build model, so xAI’s documentation set is inconsistent even though the latest dated primary source resolves the current product default. Any serious evaluation must still capture the resolved model at runtime; an alias such as grok-build-latest is not a stable experimental variable.
The current xAI model catalog positions Grok 4.6 at a 500K context window and $2 per million input tokens and $6 per million output tokens. Grok 4.5 is also listed at 500K and the same standard input/output price, with $0.30 cached input. A separate grok-build-0.1 API model has 256K context and lower prices. That model is not the Grok Build agent product. Long-context surcharges and tool charges can materially alter an API workload.
Security defaults deserve attention. Build’s permission mode is Ask, but its OS sandbox is off by default. xAI documents Landlock on Linux and Seatbelt on macOS, with profiles ranging from off through workspace, read-only, and strict. Some network restrictions differ by operating system; built-in profiles do not permanently protect every credential path. An enterprise baseline should pin an appropriate sandbox, explicit denies, and disabled bypass modes through managed requirements.
The consumer pricing page lists Free, SuperGrok at $30 monthly, and SuperGrok Plus at $100 monthly, with other tiers not fully exposed in the retrieved page text. The Grok FAQ describes a shared weekly usage pool expressed as a percentage rather than a guaranteed number of Build tasks. That makes subscription economics difficult to normalize. API cost can be modeled; a subscription “task price” cannot be honestly derived without recorded usage depletion.
Grok Build is best suited to experienced users who want to modify the agent layer, connect alternative models, embed an ACP agent, or run structured headless workflows. It is less suitable for a team that expects strong local containment without configuration or needs the longest public enterprise evidence trail.
OpenAI Codex
Codex spans the ChatGPT desktop app, ChatGPT Work/web, CLI, IDE extension, local sessions, and cloud environments. Official integrations include common VS Code-family editors, Xcode and JetBrains routes, GitHub workflows, Slack, Linear, MCP, skills, plugins, hooks, subagents, browser and web tools, scheduled tasks, an SDK, an app server, and non-interactive codex exec.
The current local model guidance uses GPT-5.6 Sol at medium reasoning for the Power setting, with Sol, Terra, and Luna forming a capability/cost ladder. Each is documented with a 1.05M context window and 128K maximum output. Standard API prices are $5/$30 per million input/output tokens for Sol, $2/$12 for Terra, and $0.20/$1.20 for Luna, with cached input at one tenth of standard input. Requests above the documented long-context threshold attract higher rates. The headline context is not the same as effective repository capacity: agent instructions, tool schemas, search results, history, and compaction consume or transform it.
Codex cloud deserves separate treatment. Its runtime is OpenAI-managed, and the default cloud model currently cannot be changed by the user. That is acceptable for a product-default trial but invalid for a model-controlled comparison. Local desktop, CLI, and IDE use shared configuration and allow more explicit model selection.
The Codex plan ladder currently includes Free, Go at $8 monthly, Plus at $20, Pro 5x at $100, Pro 20x at $200, Business at $20 per user per month annually or $25 monthly, and sales-led Enterprise/Edu. Official message estimates vary by model and plan and are expressed as broad rolling-window ranges; additional weekly limits may apply. The same task can consume very different capacity depending on context, reasoning, tools, and caching. API-key use enables CLI, SDK, and IDE workflows but not every cloud collaboration feature.
Codex’s strongest differentiator is its documented default containment. Local commands run in an OS-enforced sandbox—Seatbelt on macOS, Bubblewrap on Linux/WSL2, and a native Windows mechanism—with writes normally limited to the active workspace. Network is off by default, and actions beyond the boundary require permission. Sandbox and approval policy remain separate controls. Beta permission profiles and managed requirements can constrain filesystem and network behavior across teams.
Enterprise features include workspace controls, SSO, SCIM, groups, RBAC, separate permissions for local and cloud Codex surfaces, compliance logging, data residency options, and OpenTelemetry export. Business, Enterprise, Edu, and API inputs are not used for training by default. Consumer Plus/Pro settings are different and should be reviewed before sensitive repository use.
Codex is the strongest default choice for teams that want one product across exploratory desktop work, repository CLI sessions, parallel local agents, and managed cloud delegation. It is less attractive when the organization needs to route the same agent through several hyperscaler model providers, or when top-tier API output cost dominates the budget.
Claude Code
Claude Code is a terminal-first agent with IDE, desktop, browser/cloud, mobile monitoring, remote control, Slack, GitHub Actions, GitLab CI/CD, automatic review, and computer-use integrations. Its CLI requirements are unusually explicit: current documentation lists macOS 13+, Windows 10 1809+ or Windows Server 2019+, Ubuntu 20.04+, Debian 10+, Alpine 3.19+, x64 or ARM64, at least 4 GB of RAM, a supported shell, and internet access.
The workflow system is deep: CLAUDE.md instructions and memory, MCP, skills, plugins and marketplaces, multiple hook types, built-in and custom subagents, background subagents, agent teams, cross-session messaging, worktrees, scheduled tasks, goals, an Agent SDK, non-interactive operation, cloud sessions, and self-hosted cloud-session runners on eligible plans.
There is no single Claude Code default model. Current documentation maps Max, Team Premium, Enterprise pay-as-you-go, and Anthropic API use to Opus 5 by default; Pro, Team Standard, and Enterprise subscription seats default to Sonnet 5. Fable 5 can be available without being the default. Aliases such as opus, sonnet, best, and default are mutable and provider-dependent. A reproducible run records the resolved full model ID.
Fable 5, Opus 5, and Sonnet 5 are documented with 1M-context configurations and 128K maximum output, but plan and provider conditions matter. Claude Code compacts conversations before the active window fills. General Claude plan context numbers should not be silently substituted for the coding runtime’s model-specific context.
Current API prices are $10/$50 per million input/output tokens for Fable 5, $5/$25 for Opus 5, $2/$10 for Sonnet 5, and $1/$5 for Haiku 4.5, with separate cache-write and cache-read rates. Sonnet 5’s $2/$10 price is introductory through 31 August 2026; the announced standard price thereafter is $3/$15. That imminent change is a reason this guide is maintained weekly.
Subscription pricing currently places Pro at $20 monthly or $200 annually, Max tiers at $100 and $200 monthly, Team Standard at $20 per seat annually or $25 monthly, Team Premium at $100 annually or $125 monthly, and a current self-serve Enterprise offer at $20 per seat plus API-rate usage. Claude surfaces share rolling five-hour and weekly usage pools. Fixed task counts are not disclosed.
Claude Code starts with strict read-only permissions: edits and Bash require approval. Native Bash sandboxing can enforce filesystem and network limits with Seatbelt on macOS and Bubblewrap on Linux/WSL2, but it is opt-in. If the sandbox cannot start, the documented default is to warn and continue unsandboxed unless fail-closed behavior is configured. Windows does not have a native equivalent in the documented Claude Code sandbox.
Claude Code is particularly strong for organizations that need provider choice. It can run through Anthropic and supported third-party providers, including Amazon Bedrock, Google Vertex AI, Microsoft Foundry, or gateways. That flexibility is not free: identity, retention, managed settings, residency, audit coverage, model aliases, and pricing can change with the route. Its customer-controlled OTel export is one of the clearest agent-specific observability designs in this comparison.
Specification comparison
Scroll or swipe to compare →
| Dimension | Grok Build | Codex | Claude Code |
|---|---|---|---|
| Comparable local surface | TUI/headless/ACP | Desktop/CLI/IDE | CLI/IDE/Desktop local |
| Hosted/persistent surface | Grok Bot is separate | Codex cloud/Work | Web/cloud/remote |
| Current default identity | Grok 4.6 per 12 August model card; older Build pages still name 4.5 | Power: GPT-5.6 Sol medium; cloud default not selectable | Plan/provider-dependent Opus 5 or Sonnet 5 |
| Headline current context | 500K for Grok 4.5/4.6 | 1.05M for GPT-5.6 family | Up to 1M in qualifying model/runtime paths |
| Native non-interactive mode | Yes | Yes | Yes |
| Skills/plugins/hooks | Yes | Yes | Yes |
| MCP | Yes | Yes | Yes |
| Subagents/worktrees | Yes | Yes | Yes |
| Custom model/provider | Yes | Supported model routing; Bedrock path documented | Multiple first- and third-party provider paths |
| Local OS sandbox default | Off | On; network off | Native sandbox available, off by default |
| Default permission stance | Ask | Workspace work allowed within sandbox; escalate beyond | Read-only; edit/Bash ask |
| Fixed subscription task count | Not publicly guaranteed | Not publicly guaranteed; broad estimates plus weekly limits | Not publicly guaranteed; rolling and weekly pools |
The table reveals a convergence in visible features. All three can read and edit, execute a shell, use MCP, orchestrate subagents, work in isolated Git branches, and automate. The decisive differences now live in defaults, model resolution, ecosystem reach, observability, governance, provider routing, and the friction of using these capabilities together.
Documented workflow fit and reliability
Installation and first useful run
Claude Code publishes the clearest minimum client requirements, which lowers deployment ambiguity. Its terminal-first mental model is direct: install the client, authenticate, open a repository, review the read-only exploration, then approve edits or shell actions. Teams that already work through shells, dotfiles, and repository instructions can standardize quickly. The tradeoff is that a safe enterprise baseline requires an explicit second step: turn on native sandboxing, make failure fail closed, and decide whether provider routing changes central policy availability.
Codex offers more entry points and therefore more choice at onboarding. A developer can begin in the desktop app, an IDE extension, or the CLI and later move work into cloud or remote workflows. That continuity is a strength for mixed-seniority teams because the visible diff, terminal, worktree, and task model can be introduced gradually. It is also a source of configuration complexity: administrators should document which surfaces are approved, how local and cloud permissions differ, which model/effort settings are the norm, and when an API key is appropriate instead of ChatGPT sign-in.
Grok Build’s terminal installation and headless design are approachable for an experienced engineer. ACP and custom-provider support make it especially easy to treat the agent as a component rather than only an application. The first-run hazard is security expectation: Ask permissions can feel like containment even while the OS sandbox remains off. Onboarding documentation should make the sandbox profile, network behavior, writable paths, and bypass policy visible before the first private repository session.
Repository understanding
All three products can search files, read project instructions, inspect Git state, run local tools, and maintain a plan. Their advertised context windows are large enough to tempt users into “read the whole repository” prompts. That is rarely the best workflow. Effective comprehension depends on search, file selection, tool-output discipline, memory, and compaction. A 1M-token ceiling can still produce weak reasoning if irrelevant generated files and logs crowd out architectural invariants.
The most portable repository setup is a short root instruction file that points to authoritative build, test, style, and architecture documents, with nested instructions only where behavior genuinely differs. Codex’s AGENTS.md convention, Claude Code’s CLAUDE.md, and Grok Build’s compatibility features can all consume a concise project contract. Avoid copying a massive wiki into every prompt. Give the agent a deterministic validation command and make unknowns discoverable in the repository.
Codex is particularly attractive for visual, exploratory repository work because desktop and IDE surfaces can combine files, diffs, terminal output, browser context, and multiple agents. Claude Code is strongest when the repository contract is expressed as shell commands, hooks, permissions, and reusable subagents. Grok Build is strongest when comprehension must feed a custom host through ACP or a nonstandard model route. None has a public, controlled “repository understanding” product score for the current trio, so these are workflow-fit judgments.
Editing, diff review, and validation
The reliable agent loop is the same across brands: inspect; state a falsifiable plan; make a bounded edit; run the closest tests; inspect the diff; run broader validation in proportion to risk; report evidence and limitations. Products differ in how naturally they enforce it.
Codex’s workspace sandbox and approval boundary encourage local iteration without repeated permission prompts for ordinary repository work. Its desktop and IDE surfaces are well suited to comparing parallel worktrees and reviewing a patch in context. The risk is over-delegation: cloud or multi-agent features can create several plausible patches, increasing review load if ownership and acceptance criteria are vague.
Claude Code’s read-only default makes the transition from analysis to mutation explicit. Hooks can require or record validation, and worktrees isolate parallel sessions. Its permissions are expressive enough to create narrow command allowlists. Native sandboxing should still be enabled because a permission prompt and an OS boundary address different failure modes.
Grok Build supports plan mode, Ask/Auto/Always-approve, worktrees, hooks, subagents, and background tasks. Plan mode must not be treated as a complete write lock: official documentation notes that shell redirection and the behavior of subagents require separate control. The product is flexible enough to build a rigorous loop, but administrators must compose the controls deliberately.
Automation and CI
For unattended work, pin everything that can move: full model ID, client version, container image, dependencies, tool permissions, network domains, timeout, retry count, maximum cost, and output schema. Mutable aliases are useful for interactive use and dangerous for regression baselines.
Grok Build’s streaming JSON, resumable headless sessions, ACP endpoint, scheduled loops, and custom-model support make it the most obviously embeddable of the three. Claude Code offers non-interactive execution, an Agent SDK, GitHub Actions, GitLab CI/CD, cloud sessions, scheduled work, and self-hosted runners. Codex provides codex exec, SDK/app-server routes, GitHub Action, MCP server behavior, scheduled tasks, and local/cloud delegation.
No CI agent should receive unrestricted production credentials or permission to merge based solely on its own test report. Use short-lived credentials, protected branches, environment separation, signed or attributable commits, deterministic checks, and human approval for high-impact changes. Treat issue text, build logs, dependency metadata, and web pages as untrusted prompt input.
Failure recovery
Agent failures usually fall into five classes:
- Specification failure: the task is ambiguous or the hidden requirement is not discoverable.
- Context failure: the agent reads the wrong files, loses an invariant during compaction, or overweights stale instructions.
- Tool failure: the shell, network, dependency mirror, browser, or MCP server behaves differently from the expected environment.
- Reasoning failure: the proposed change does not satisfy the requirement despite adequate evidence.
- Governance failure: permissions are too broad, approval fatigue leads to unsafe acceptance, or telemetry is insufficient to reconstruct the run.
The best product is the one that makes the dominant local failure visible and recoverable. Codex’s default isolation reduces the blast radius of tool and governance failures. Claude Code’s detailed event export and permission model make it easier to reconstruct agent actions. Grok Build’s headless session and extensibility make custom recovery logic easier to build. Public benchmarks mostly score final success and underweight diagnosis quality, so a pilot should record whether the agent identifies its own failing assumption, reverts a bad patch, and recovers without a complete restart.
Evidence-based product-fit scorecard
The following qualitative ratings are analyst judgments derived from documented capabilities, defaults, current evaluations, and evidence maturity. They are not benchmark measurements. “Leader” means the strongest documented fit in this comparison; “Strong” means a credible default; “Conditional” means configuration, provider, or evidence gaps materially affect the recommendation.
Scroll or swipe to compare →
| Dimension | Grok Build | Codex | Claude Code | Interpretation |
|---|---|---|---|---|
| Native local containment default | Conditional | Leader | Conditional | Grok/Claude sandboxes require enablement; Codex defaults contained. |
| Terminal-native workflow | Strong | Strong | Leader | Claude’s CLI and policy model are exceptionally mature. |
| Desktop/local/cloud breadth | Conditional | Leader | Strong | Codex has the most coherent breadth; Grok Bot is separate. |
| Provider/custom-model flexibility | Leader | Strong | Leader | Grok custom providers; Claude hyperscaler/gateway breadth. |
| Enterprise documentation maturity | Conditional | Leader | Leader | Grok controls are credible but product-specific coverage is thinner. |
| Agent-native observability | Conditional | Strong | Leader | Claude’s detailed OTel design leads; Codex is strong. |
| Current native long-horizon evidence | Conditional | Strong | Leader | Claude has the SWE-Marathon point lead; one benchmark is not universal. |
| Shared IDE-harness cost/performance | Leader | Strong | Conditional | CursorBench favors Grok value; product UX is not measured. |
| Reproducibility of default identity | Conditional | Strong | Conditional | Older Grok docs conflict with the current card; Claude defaults vary; Codex cloud is opaque. |
A regulated bank should heavily weight containment, provider contract, audit, and residency. A solo tool builder may weight ACP, custom models, price, and headless execution. The scorecard exposes those tradeoffs without pretending that an unexplained decimal is a measurement.
What the public benchmarks actually say
Why a league table would be misleading
Three recent facts should reduce confidence in small score differences without erasing real category results.
First, OpenAI’s July 2026 audit of SWE-bench Pro estimated that roughly 30% of tasks had material issues and withdrew its earlier recommendation of the benchmark. Broken tests, ambiguous specifications, and environment defects can reward the wrong behavior. OpenAI had already stopped emphasizing SWE-bench Verified because of increasing contamination.
Second, Anthropic demonstrated that infrastructure configuration alone can move Terminal-Bench 2.0 results by about six percentage points—larger than many leaderboard gaps. Container setup, timeouts, network policy, package mirrors, terminal state, and retry logic are part of the measurement.
Third, “native product” and “shared harness” answer different questions. SWE-Marathon uses Claude Code, Codex, and Grok Build as deployed stacks, so model, agent, tools, prompts, and effort all contribute. DeepSWE, CursorBench, APEX-SWE, and FrontierCode hold the agent scaffold more constant, so they compare model configurations rather than the native products. Both views matter; they must not be merged into a single unexplained score.
The most relevant disclosed results
SWE-Marathon v1.1 — native-product, ultra-long autonomous work. Twenty distinct task clusters are run eight times per configuration, for 160 correlated trials per row. Claude Code + Opus 5 Max resolves 50.0%; Claude Code + Opus 4.8 Max 48.8%; Claude Code + Fable 5 with fallback 45.0%; Codex + GPT-5.6 Sol Max 42.5%; Grok Build + Grok 4.6 High 31.9%; Claude Code + Sonnet 5 Max 30.0%; and Grok Build + Grok 4.5 High 29.4%. Claude Code + Opus 5 is the point-score leader, by 7.5 percentage points over Codex and 18.1 over Grok Build. The owner does not publish configuration-level confidence intervals or a paired significance test, so the result does not establish statistical separation. Confidence in the directional category result is medium: the generalization sample is 20 tasks, native agents differ, and effort settings are not identical.
DeepSWE v1.1 — common mini-swe-agent, long-horizon repository work. Across 113 tasks, Opus 5 Max reports 74% ±4, $11.84 per task, 118K output tokens, and 99 steps. GPT-5.6 Sol Max reports 73% ±3, $8.39, 60K tokens, and 61 steps. Grok 4.6 Extra High reports 67% ±2, $5.50, 71K tokens, and 87 steps. Fable 5 Max with fallback reports 70% ±4 at $21.63. The benchmark reports overlapping Opus/Sol intervals but does not publish an equivalence test; the public evidence therefore does not establish a decisive difference. Sol uses roughly 49% fewer output tokens, 38% fewer steps, and 29% lower reported cost than Opus while remaining near the top.
CursorBench 3.2 — shared Cursor harness, ambiguous multi-file IDE work. Grok 4.6 Extra High leads the point scores at 70.8%, $2.81 per task, 41,136 tokens, and 46 steps. Fable 5 Max records 70.5% at $17.32; Opus 5 Max 70.0% at $8.23; Grok 4.6 High 69.9% at $2.34; and GPT-5.6 Sol Max 67.2% at $5.69. Cursor does not expose a public task count or confidence intervals on the board and cautions against overreading small differences. Public data are insufficient to determine whether the small point gaps are meaningful; Grok is nevertheless the clear reported cost-efficiency leader.
APEX-SWE — shared Terminus-2 harness, integration and observability. On 200 hidden cases, Opus 5 Max records 63.7% ±6.4, Fable 5 Max 58.8% ±6.4, and Grok 4.6 High 56.4% ±6.2. The reported uncertainty bands overlap, so Opus has the best point estimate without a categorical statistical win. A current GPT-5.6 Sol row was not visible on the owner board, making this an incomplete three-way comparison. Pass@1 is the metric; the public static evidence does not establish that only one physical rollout was performed.
Terminal-Bench 3.0 — diverse terminal work. A current mixed-provenance snapshot reports Opus 5 Max at 43.5%, GPT-5.6 Sol Max at 34.6%, Fable 5 Max with fallback at 34.1%, Grok 4.6 High at 26.0%, Opus 4.8 at 21.1%, Grok 4.5 at 15.7%, and Sonnet 5 at 14.6%. The benchmark launched with 74 tasks across seven domains. Its methodology is strong, but exact agent/harness and attempt settings are not consistently normalized across the newest rows, which are easiest to verify in xAI’s model card rather than a static owner export. Opus 5 is therefore the point leader in the published snapshot, not a clean native-product or common-harness winner.
FrontierCode 1.1 Extended — shared Cognition scaffold. The current 150-task snapshot reports Fable 5 Max with fallback at 64.9%, Opus 5 Max 63.6%, Grok 4.6 High 61.3%, GPT-5.6 Sol Max 60.6%, Opus 4.8 59.6%, GPT-5.5 56.7%, and Sonnet 5 56.2%. The private extended set limits auditability and no confidence intervals are disclosed. It supports the conclusion that current leaders form a close cohort rather than one universal champion.
Artificial Analysis Coding Agent Index — independent native-agent composite. The current v1.3 method uses 321 tasks across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA, with exactly three attempts per task and equal-weighted component scores. It is one of the strongest independent product-configuration methods. Its dynamic chart did not expose a complete, auditable current Grok 4.6 row in static output during this review, and published v1.1 and v1.3 numbers cannot be combined. The method is retained; stale revision scores are not imported into the headline.
RuBench v2 — a configuration warning, not a three-way result. This 25-task study ran product configurations three times and found that one Claude Code configuration silently substituted Fable 5 with Opus 4.8 on 20% of tasks; it also identified tool-call contamination in some GPT-5.6 runs. Grok Build was not included, so it cannot rank the trio. Its lesson is still decisive: the deployed product configuration, including fallbacks and tool traces, is the experimental unit.
Scroll or swipe to compare →
| Workload | Point leader | Direct answer | Critical caveat |
|---|---|---|---|
| Native ultra-long autonomous engineering | Claude Code + Opus 5 point leader | 50.0% on SWE-Marathon v1.1 | 20 task clusters; statistical separation not established |
| Shared-harness long-repo engineering | Opus 5 by 1 point | Reported Opus/Sol intervals overlap; no decisive difference established | Product UX is not measured |
| Shared-harness IDE work | Grok 4.6 point leader | 70.8% and best reported cost efficiency | No public CIs/task count; point-gap meaning is unknown |
| Integration and observability | Opus 5 point estimate | 63.7% ±6.4 on APEX-SWE | Intervals overlap; no current Sol row |
| Diverse terminal tasks | Opus 5 point leader | 43.5% current snapshot | Mixed row provenance and incomplete harness normalization |
| Production-code shared scaffold | Fable 5 point estimate | 64.9% on FrontierCode Extended | Close cohort; private tasks; fallback disclosed |
Task-level performance verdicts
The evidence supports category point leaders and buyer recommendations, not a single measured crown. Claude Code + Opus 5 has the current public native-product point lead for ultra-long autonomous work. Opus 5 and GPT-5.6 Sol have overlapping reported intervals in shared-harness long-repository work, with Sol clearly more efficient in the reported run. Grok 4.6 leads the shared IDE-harness point score and value. For narrower bug fixing, feature implementation, refactoring, test generation, and repository explanation, the suites do not expose a clean, current, product-level three-way slice with enough power to name a universal task winner.
Product-fit recommendations combine these measurements with documented defaults:
Scroll or swipe to compare →
| Task or environment | Recommended product | Confidence | Basis |
|---|---|---|---|
| Safely explore an unfamiliar local repository | Codex | Medium-high | Sandbox and network-off defaults reduce accidental reach. |
| Terminal-centric implementation with explicit approvals | Claude Code | Medium | Mature terminal workflow and strict read-only permission default. |
| Prototype an ACP integration or swap model providers | Grok Build | High | ACP and custom model/provider support are documented first-class features. |
| Ultra-long native autonomous engineering | Claude Code + Opus 5 | Medium | SWE-Marathon v1.1 point lead of 7.5 points over Codex; separation not statistically established. |
| Efficient shared-harness repository engineering | GPT-5.6 Sol | Medium-high | One point behind Opus on DeepSWE; far fewer tokens/steps and lower cost. |
| Ambiguous multi-file IDE-style work | Grok 4.6 | Medium | Highest CursorBench point score and lowest cost among leading rows. |
| Parallel local feature work across worktrees | Codex or Claude Code | Medium | Both have mature workflows; no controlled native speed winner. |
| Enterprise deployment across AWS/GCP/Azure model routes | Claude Code | High | Broad officially documented provider paths. |
| Mixed desktop, web, CLI, mobile, and managed cloud delegation | Codex | Medium-high | Coherent cross-surface product breadth. |
| Lowest documented frontier-model output rate in this snapshot | Grok 4.6 API | Medium | $6/MTok output; excludes tools, long context, subscription economics, and product overhead. |
| Regulated repository with strong local containment | Codex, subject to contract | Medium | Strong default isolation; eligibility and third-party scope still require review. |
How we would test them fairly
This edition synthesizes public evidence; it does not pretend that differently configured leaderboards are an original Kingy.ai test. A result capable of changing the verdict would use 30–60 frozen repository tasks across bug fixing, features, refactors, test generation, build repair, security fixes, repository explanation, and long-horizon work.
Two lanes are required. A product-default lane measures the agent a buyer actually receives, including its resolved model, tools, permissions, sandbox, compaction, and retries. A controlled-model lane matches environment, tool schema, timeout, retry policy, prompt, network, and economic/opportunity budgets as closely as the products allow. The lanes must never be averaged into one unexplained score.
The primary outcome is verified task success against task-specific tests, a repository-wide regression suite, and a hidden semantic oracle. Secondary outcomes include first-pass success, human intervention, time, fully loaded cost, tokens, tool calls, patch churn, introduced security findings, mutation score, recovery, and repeatability. Use at least three independent runs per product-task pair, treat repository as a cluster, report uncertainty and effect sizes, and blind human reviewers to product identity.
The complete preregisterable protocol appears in the evaluation appendix, after the buyer decision material.
Cost and value
API prices are only marginal prices
Scroll or swipe to compare →
| Current model example | Input / MTok | Cached read / MTok | Output / MTok | Context headline |
|---|---|---|---|---|
| Grok 4.6 | $2.00 | Not consistently disclosed in current catalog | $6.00 | 500K |
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 | 1.05M |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | 1.05M |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | 1.05M |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | Up to 1M |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 introductory | Up to 1M |
At a hypothetical 200K uncached input and 30K output, standard token charges would be roughly $0.58 for Grok 4.6, $1.90 for GPT-5.6 Sol, $0.76 for Terra, $0.076 for Luna, $1.75 for Opus 5, and $0.70 for Sonnet 5 at the introductory rate. This illustration excludes long-context multipliers, cache writes, tools, web search, code execution, retries, compaction, and agent overhead. It is not a measured task cost.
The cheapest model is not necessarily the cheapest completed task. If a low-cost run needs retries, human repair, or produces regressions, its cost per verified success rises. Conversely, an expensive model can be economical on high-value work if it succeeds on the first attempt. A procurement model should calculate:
Fully loaded cost per verified success = subscription allocation + marginal usage + engineer intervention + compute/CI + review/rework + incident risk.
Subscription economics cannot be normalized today
xAI expresses paid usage as a shared weekly pool. OpenAI publishes broad rolling-window message estimates that vary by plan and model, and warns that weekly limits may apply. Anthropic uses shared rolling five-hour and weekly pools without fixed task counts. None publishes a stable, comparable “coding tasks per month” entitlement, so such a table would be invented precision.
Instead, pilot teams should log beginning and ending quota, productive agent minutes, tasks attempted, verified successes, and overage spend. After two representative weeks, calculate success per included subscription dollar. Keep personal and enterprise plans separate because training defaults, governance, support, and capacity differ.
Security, privacy, and enterprise readiness
Default local posture
Codex has the strongest documented out-of-box containment. It combines an OS sandbox, workspace-scoped writes, and network disabled by default. That does not make every configuration safe: approvals, permission profiles, connectors, MCP annotations, browser tools, and explicit network access can expand reach.
Claude Code has the strictest permission prompt default—read-only until edits or Bash are approved—but the native filesystem/network sandbox is optional. Teams should enable it and configure fail-closed behavior. Built-in read/edit/write tools remain governed by permission rules, while Bash subprocesses inherit sandbox restrictions.
Grok Build asks before sensitive actions by default, but its OS sandbox is off. Managed deployments should pin a workspace or strict profile, protect credential paths with explicit denies, restrict bypass modes, and test operating-system-specific network behavior. Plan mode is not a complete write barrier: shell redirection and subagent behavior require separate controls.
Data use and retention
All three vendors state that qualifying commercial/API data is not used for model training by default or without opt-in: see xAI’s enterprise terms, OpenAI’s business-data commitments, and Anthropic’s Claude Code data-use documentation. Consumer subscriptions have different controls. Do not infer that a personal plan has the same contract as a business workspace.
Standard provider retention is often around 30 days for API content, but surface-specific application state, abuse logs, workspace retention, and local transcripts vary. Grok Build maintains local session state under ~/.grok/. Claude Code stores local transcripts under ~/.claude/projects/ with a configurable cleanup period. Codex local files and exported telemetry remain under customer control, while cloud and ChatGPT data follow the applicable workspace and endpoint policies.
Zero-data-retention arrangements have exceptions. They do not erase local histories, customer SIEM exports, repository commits, third-party MCP data, or cloud-provider logs. Confirm the exact model, endpoint, feature, and account in the contract; OpenAI likewise documents endpoint-specific data controls, while Claude Code documents surface-specific storage and retention.
Enterprise controls
Codex and Claude Code have the most detailed public enterprise documentation. Codex is strong in workspace RBAC, separate local/cloud permissions, managed requirements, compliance logging, residency options, and default isolation. Claude Code is strong in granular managed settings, provider-native IAM, rich OTel, and multiple data-plane options. Some Anthropic server-managed settings require direct Anthropic connectivity and do not apply through third-party providers.
Grok Build documents enterprise OIDC and managed requirements, bypass restrictions, ZDR, SSO/SCIM/RBAC, and xAI audit APIs. Public coverage is thinner on Build-specific tool-event telemetry and the precise product scope of some platform-wide enterprise controls. Buyers should require order-form confirmation for CMEK, residency, audit fields, and regulated workloads.
Procurement control matrix
This matrix is a diligence starting point, not a compliance certification. “Documented” means a public vendor page describes the control; it does not prove that the control is included in every plan, provider route, region, or integration.
Scroll or swipe to compare →
| Procurement question | Grok Build | Codex | Claude Code |
|---|---|---|---|
| Local containment default | OS sandbox off; permission prompts on | OS sandbox on, workspace write, network off | Read-only permissions by default; native sandbox optional |
| Central policy | Enterprise OIDC and managed requirements | Managed configuration | Server-managed settings, with route-specific limits |
| Commercial training posture | Enterprise terms | Business-data commitments | Claude Code data usage |
| Retention / ZDR | API security and ZDR; verify Build scope | Endpoint-specific controls | Surface-specific retention; verify provider route |
| Audit and telemetry | Management audit API; Build event scope needs confirmation | Compliance logging plus customer-controlled exports; verify workspace and surface | OpenTelemetry usage monitoring and provider-native logs |
| Provider / data-plane choice | xAI-hosted and custom-model paths; confirm contract per route | OpenAI local/cloud surfaces; confirm workspace, endpoint, and region | Anthropic, Bedrock, Vertex AI, Foundry, and gateways |
| Regulated-workload decision | Contract and order-form review required | Contract, eligible service, and exact workspace path required | Contract, eligible provider route, and exact service scope required |
Certifications and BAAs apply only to named services and scopes. Local endpoints, plugins, connectors, web search, MCP servers, gateways, and customer infrastructure are not automatically covered. No product should receive a blanket “HIPAA compliant” label without the qualifying account, executed agreement, supported path, retention mode, and customer configuration.
Recommendations by buyer
Solo developer
Start with the agent already included in a service you use, then run the same five real tasks for a week. Choose Codex if you value a polished desktop-to-CLI continuum and safe defaults. Choose Claude Code if you live in the terminal and want fine-grained workflow control. Choose Grok Build if you want to experiment with the agent stack or use custom models. Do not buy three expensive tiers before learning whether your bottleneck is model quality, quota, or review time.
Startup
Codex is the default recommendation for a small team needing speed across product work, review, and cloud delegation. Claude Code becomes preferable when the startup already standardizes on AWS, GCP, or Azure model routing, or wants deep terminal automation and OTel. Grok Build is attractive for cost-sensitive experimentation but should be deployed with a pinned sandbox and an explicit evaluation of documentation maturity.
Enterprise engineering organization
Run a two-vendor pilot rather than a brand vote. Shortlist Codex for containment, workspace governance, and broad surfaces; shortlist Claude Code for provider flexibility, SIEM integration, and policy depth. Include Grok Build where ACP, custom models, or xAI economics matter, but require product-specific contractual answers. Measure verified success and intervention on internal tasks, not only developer preference.
Regulated organization
Select the contracted path before the model. A local tool sending prompts to a consumer account is not equivalent to a regulated enterprise workspace. Confirm BAA eligibility, ZDR or modified retention, inference and storage regions, key management, audit coverage, endpoint controls, and third-party exclusions. Codex currently has the strongest default local containment, but the final choice can change with the required cloud and identity architecture.
Open-source maintainer
Prioritize patch discipline, reproducible tests, and protection against prompt injection in issues and repository content. Codex’s network-off sandbox is valuable for unknown repositories. Claude Code’s permission and hook system can implement a strong review workflow. Grok Build’s sandbox must be enabled. Require agents to show diffs and test evidence; never auto-merge solely on an agent’s self-report.
Migration and switching costs
The products share enough conventions to make a staged migration possible, but not free. Repository instructions differ—AGENTS.md, CLAUDE.md, skills, hooks, plugin formats, permission policies, and MCP configuration. A portable baseline should keep universal build/test commands in ordinary repository documentation, isolate vendor-specific instructions in thin adapters, use standard MCP transports, and avoid aliases for production automation.
Preserve a neutral task manifest containing repository commit, prompt, acceptance tests, environment digest, and budget. When switching, rerun a representative sample and compare outcomes. Do not assume that a prompt tuned for one tool schema transfers unchanged. Review retention and local transcript paths before decommissioning a client.
Decision tree
- Do you require the safest documented local default? Choose Codex for the pilot.
- Do you require Bedrock, Vertex AI, Foundry, or gateway routing? Choose Claude Code for the pilot.
- Do you require first-class custom-model support or ACP embedding? Choose Grok Build for the pilot.
- Do you need one integrated desktop/local/cloud collaboration product? Prefer Codex.
- Do you optimize for terminal automation and agent-event observability? Prefer Claude Code.
- Is marginal frontier-model API price the dominant constraint? Evaluate Grok 4.6, but measure product overhead and success rate.
- Is the workload regulated? Stop and validate the exact account, endpoint, retention, provider, integrations, and contract before testing sensitive code.
Limitations
This edition did not execute paid three-way product tests. It does not claim a controlled original performance result. The public benchmark section uses disclosed results as historical evidence and explicitly avoids updating older model scores into current-product rankings.
Official documentation can be inconsistent or change without notice. xAI’s dated 12 August model card resolves Grok Build’s current default as Grok 4.6, but older Build pages still naming 4.5 illustrate why the updater must prefer the latest dated primary source and preserve runtime-resolution evidence. Prices, aliases, quotas, supported surfaces, retention terms, and security controls are time-sensitive. Weekly verification reduces staleness but cannot guarantee that a vendor has published every runtime change or that a page was available at the scheduled check.
Feature existence is not usability. This guide did not measure installation reliability, latency under load, IDE polish, rate-limit behavior, support responsiveness, or accessibility. Security conclusions describe documented defaults and controls, not penetration-test results.
The overall winner is a decision recommendation, not a statistically significant task-performance ranking. Buyers with high stakes should execute the preregistered protocol on their own repositories.
Appendix: Evaluation framework for a real head-to-head
The following protocol is designed to produce a result that could legitimately update the verdict. It is a proposed benchmark, not a result from this edition.
Experimental unit and sample size
Use at least 30 tasks, ideally 45 to 60, drawn from repositories that were not public before the candidate models’ knowledge cutoffs. Stratify the task set across:
- Small bug fixes with deterministic failing tests.
- Cross-file feature additions with acceptance tests.
- Refactors with behavioral equivalence checks.
- Test generation measured by mutation score, not raw line coverage.
- Repository comprehension and precise explanation.
- Dependency or build-system repair.
- Security-relevant fixes with adversarial tests.
- Long-horizon tasks requiring planning, validation, and recovery.
The repository is the blocking unit. Avoid placing near-duplicate issues from one repository into both tuning and evaluation strata. Freeze every repository in an immutable image with a commit hash and test oracle.
Two comparison lanes
Run two distinct lanes and never average them together without reporting both.
Product-default lane: Use each product as a normal buyer would receive it. Record the resolved model, default reasoning, permissions, sandbox, tools, compaction, and retries. This measures product value but does not isolate model quality.
Controlled-model lane: Match the environment, tool schema, timeout, retry policy, token or dollar budget, network access, and prompt as closely as the products allow. Pin full model IDs. This better isolates system capability but may remove product advantages.
If a product prevents model pinning, record that fact and classify the run as product-default. Do not pretend the variable was controlled.
Configuration controls
For every run capture:
- Product and client version.
- Surface: CLI, IDE, desktop local, or cloud.
- Full resolved model ID and provider.
- Reasoning/effort setting.
- Context and compaction events.
- Permission and sandbox mode.
- Network policy and allowed domains.
- MCP servers, plugins, hooks, skills, and project instructions.
- Container image digest, operating system, CPU, memory, disk, and region.
- Prompt, task assets, seed if exposed, timeout, retry rule, and maximum turns.
- Input, cached input, output, tool charges, subscription depletion, and wall time.
- Every patch, command, test result, approval, and error.
Use a fresh worktree or container for each run. Prewarm or clear caches consistently. Randomize product order within each task to reduce time-of-day and service-load effects. Blind human graders to product identity.
Budgets
Equalize at least two budgets:
- Equal economic budget: the same maximum marginal API-equivalent spend per task, excluding fixed subscription fees but reporting them separately.
- Equal opportunity budget: the same timeout, maximum agent turns, retry count, and permitted tools.
Token equality alone is not fair because models price and tokenize differently. Dollar equality alone is not enough when one product includes subscription capacity. Report both resource use and the outcome frontier.
Primary metrics
The primary metric should be verified task success: all task-specific tests pass, the repository-wide regression suite passes, and a hidden semantic oracle accepts the behavior. Partial credit must be defined before the run.
Secondary metrics:
- First-pass success.
- Regression-free success.
- Human intervention count and minutes.
- Wall-clock time to verified success.
- Marginal and fully loaded cost per successful task.
- Tokens and tool calls per success.
- Patch size and unnecessary churn.
- Static-analysis and security findings introduced.
- Test quality through mutation score.
- Recovery rate after an induced failure.
- Reproducibility: same task succeeds across repeats.
Publish an overall score only when weights were preregistered, and keep every raw component visible so buyers can apply their own priorities.
Repeats and statistics
Use a minimum of three independent runs per product-task pair for stochastic systems; five is preferable for the final shortlist. Report pass@1 from the first run and empirical success across repeats. Do not report best-of-N as though it were pass@1.
For paired binary success, use a paired bootstrap over tasks for confidence intervals and a McNemar or randomization test for pairwise differences. For time and cost, report medians, interquartile ranges, paired bootstrap intervals, and survival curves when timeouts occur. Correct multiple pairwise comparisons with Holm’s method. Report effect sizes and confidence intervals, not only p-values.
Treat repository as a cluster if several tasks come from one repository. Predefine how infrastructure failures, model refusals, quota errors, and invalid tasks are handled. Publish the result with and without tasks adjudicated invalid.
Human review rubric
Human reviewers should score only dimensions not fully captured by tests:
- Correctness and completeness.
- Scope discipline and unnecessary changes.
- Maintainability and alignment with repository conventions.
- Security and privacy risk.
- Test appropriateness.
- Explanation quality and truthful uncertainty.
Use at least two reviewers for a stratified sample. Measure inter-rater agreement. Resolve disagreements using the written rubric, not product reputation.
FAQ
Is Grok Bot the same as Grok Build?
No. Grok Build is the coding agent compared in the local lane. Grok Bot is a persistent named cloud agent with durable state and a shared user-scoped computer. It belongs in a separate persistent-agent comparison.
Which agent writes the best code?
The current public evidence cannot answer that universally. Historical vendor tables produce different leaders by suite, and current products use newer or plan-dependent models. Run a controlled task set on your repositories.
Which is best overall?
Codex is this guide’s medium-confidence overall recommendation because it combines broad surfaces, strong current models, local/cloud orchestration, and the strongest documented containment default. Claude Code is extremely close for terminal-first and provider-flexible teams.
Which is cheapest?
At standard API rates in this snapshot, Grok 4.6 has a low frontier output price. GPT-5.6 Luna is much cheaper but targets a different speed/capability point. Subscription value cannot be ranked without usage logs and verified success rates.
Which has the largest context window?
Codex GPT-5.6 models are documented at 1.05M; qualifying Claude Code model paths at 1M; Grok 4.5/4.6 at 500K. Headline context is not effective repository capacity and does not prove better long-horizon performance.
Which is safest by default?
Codex, based on documented local defaults: OS sandbox, workspace-scoped writes, and network off. Claude Code has strong read-only permissions but optional native sandboxing. Grok Build’s sandbox is off by default.
Which is best for enterprise cloud-provider choice?
Claude Code. It officially supports Anthropic plus Bedrock, Vertex AI, Microsoft Foundry, and gateways. Confirm which managed settings and data controls apply on each route.
Can I use these tools with private code?
Yes under appropriate terms and controls, but “private” is not a configuration. Use a qualifying commercial account, verify training and retention settings, restrict tools and network, inspect local transcripts, and review MCP/connectors.
Are benchmark leaderboards useful?
Yes, as one input. They are most useful when tasks are valid, configuration is disclosed, the model is current, the harness resembles your workflow, and confidence intervals are reported. Small unnormalized gaps should not drive procurement.
How often should this comparison be refreshed?
Weekly for model aliases, prices, quotas, defaults, deprecations, and major security changes; immediately after a model launch or pricing transition. Benchmark verdicts should change only when comparable evidence changes.
Source notes
The public claim and source ledger records evidence classes, supported claims, access dates, and status. The companion benchmark normalization ledger exposes every quoted score with revision, unit, harness, configuration, attempts, uncertainty, fallback, cost, tokens, steps, and comparison rule. The task manifest, methodology, limitations, and weekly maintenance rules remain in the reproducibility package.
Primary product claims are sourced to current official documentation from xAI Build, OpenAI Codex, and Claude Code. Current evaluation claims trace to SWE-Marathon, DeepSWE, CursorBench, APEX-SWE, Terminal-Bench, FrontierCode, and the Artificial Analysis Coding Agent Index. Benchmark cautions are supported by OpenAI’s coding-evaluation audit, OpenAI’s SWE-bench Verified note, and Anthropic’s infrastructure-noise analysis.
Weekly maintenance promise
This edition is attached to an active Monday 09:00 America/Vancouver verification workflow. The updater checks model defaults and aliases, product pages, context and price tables, plan limits, retirement dates, security controls, benchmark-owner notices, and both public ledgers. It may update only this existing Kingy.ai post after factual, source, link, metadata, and layout gates pass; it must preserve the post ID, slug, Blog category, featured image, canonical URL, and rollback safeguards. Material verdict changes require human review. It does not create a new post, run paid tests, use private accounts or repositories, purchase credits, contact vendors, erase history, or convert vendor-reported numbers into independent results.
Continue exploring AI tools
Compare additional model specifications in the Kingy.ai AI model database, or read more independent analysis in the Kingy.ai Blog.
