AI News

Grok 4.7 Benchmarks: Specs, Pricing and Frontier Model Comparisons

Grok 4.7 makes a strong case for developers who care about the bill as much as the score. In Cursor’s published coding evaluation, it beats GPT-5.6 Sol at the compared settings while costing less per task. Claude Fable 5.1 still scores higher. That is a useful result for buyers, even without a claim that Grok wins everything.

Our recommendation is to shortlist Grok 4.7 for repository work and long-running agents, then measure its completion rate on your own jobs. Start with high reasoning effort and compare it with xhigh before paying for extra thinking on every request.

Sources checked September 21, 2026. This is an analysis of published evaluations and official documentation. Kingy.ai did not run a new Grok 4.7 model test for this article. Benchmark settings and source differences are identified below.

Looking for access instructions? Read our Grok 4.7 launch, pricing and availability guide.

Grok 4.7 specifications: what is confirmed

The Grok 4.7 overview, model detail page and model directory now provide the specifications that were missing from early release speculation.

SpecificationConfirmed detail
API model IDgrok-4.7
Context window500,000 tokens
Inputs and outputsText and image input; text output
Knowledge cutoffMay 2026
Reasoning effortLow, medium, high and xhigh; high is the default
Output capxAI documents no separate text output limit; this does not mean unlimited context, runtime or free output
Developer capabilitiesFunction calling, structured outputs, web search, X search and code execution
API supportResponses and Chat Completions; the model detail page lists Batch API as unsupported
Parameter countNot disclosed in the launch announcement or model pages checked

The context window is an input capacity, not a guarantee that every detail will be recovered correctly. If your workflow depends on a clause buried in a large contract or a function in a sprawling repository, test that retrieval explicitly. A long-context specification alone does not settle the question.

Image input also does not make this an image generator. Grok 4.7 produces text; xAI’s image, video and voice products use separate models and APIs. Current events require enabled search tools rather than an assumption that the base model knows everything happening on X.

The published benchmark table

The following results come from xAI’s launch comparison. Treat them as vendor-published results, not a Kingy.ai reproduction. The column headings matter: most Grok 4.7 results use xhigh, while Grok 4.6 uses high. DeepSWE is the marked exception.

EvaluationGrok 4.7 xhighGrok 4.6 highGPT-5.6 Sol maxClaude Fable 5.1 max
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%*65.2%72.7%70.0%
Terminal-Bench 4.038.0%20.3%37.3%57.9%
AA Briefcase v1.11,6571,5461,4871,678
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%
EEBench64.0%53.0%39.4%56.4%

* Grok 4.7’s DeepSWE result uses high effort. AA Briefcase is reported as a score, not a percentage. Higher is better within each row; the scales are not interchangeable.

The table gives no basis for a universal winner. Grok leads these compared entries on legal work and electrical engineering; Fable leads the terminal test; Sol leads DeepSWE. The health and legal rows describe performance on particular evaluation sets, not permission to rely on a model for unsupervised clinical or legal decisions.

For professional work, xAI’s separate GDPval chart reports 1,695 for Grok 4.7 xhigh, 1,735 for Fable 5.1 max and 1,542 for GPT-6 Astra max. That is one reported task-based comparison with Astra. It does not establish that Grok beats Astra across coding, research or general reasoning.

Grok 4.7 vs GPT-5.6, Claude and Gemini on CursorBench

Cursor’s live evaluation page lets us compare the models on the same benchmark version and inspect average task cost. It also includes Gemini 3.8 Flash and Claude Opus 5, which are absent from xAI’s main launch table.

Model and effortCursorBench 4.0Average cost / task
Claude Fable 5.1 max51.8%$17.28
Claude Opus 5 max46.6%$11.95
Grok 4.7 xhigh46.3%$6.01
Grok 4.7 high43.9%$4.69
GPT-5.6 Sol max41.7%$8.23
Grok 4.6 xhigh41.4%$6.10
Gemini 3.8 Flash high39.6%$4.70

Source: Cursor, checked September 21, 2026. These are selected rows, not the entire leaderboard. Costs reflect the tokens consumed in this evaluation, including the provider’s applicable input, cache and output rates. They are not subscription charges or a quote for your own task.

Against Sol max, Grok xhigh gains 4.6 percentage points while reducing the reported average task cost by approximately 27%. Against Opus 5 max, Grok’s score is 0.3 points lower and its task cost is approximately 50% lower. Cursor cautions that small score differences may not be statistically meaningful, so that 0.3-point gap should not drive a purchase by itself.

Fable max buys a larger score advantage at a substantially higher task cost. It is worth trying when an additional successful completion would save expensive human work. Grok deserves attention when you have enough volume for lower average costs to matter. Those are workload decisions, not rival fan clubs.

The Gemini result supports a narrower conclusion: Grok scores higher than Gemini 3.8 Flash high on this coding evaluation. It says nothing by itself about which model is better at video understanding, broad research or another Gemini configuration.

For more context on the alternatives, see Kingy.ai’s Claude Fable 5.1 review and Gemini 3.8 Flash comparison.

Grok 4.7 vs Grok 4.6: compare the same effort

The launch table’s 40.4-to-46.3 jump mixes high and xhigh. Cursor’s matching settings give a clearer upgrade comparison. At high, the score rises from 40.4% to 43.9%. At xhigh, it rises from 41.4% to 46.3%. Those are gains of 3.5 and 4.9 percentage points, respectively.

Equal effort labels still do not mean equal compute budgets across different model families. Within the Grok comparison, however, they remove one obvious source of confusion.

There is also a practical reason to resist setting every request to xhigh. On Grok 4.7, moving from high to xhigh adds 2.4 points on CursorBench and raises average cost from $4.69 to $6.01, about 28%. Whether that trade is worthwhile depends on the cost of an incorrect or unfinished answer. For a quick refactor, high may be enough; for a difficult failure investigation, xhigh is a sensible candidate to test.

API pricing, caching and the long-context threshold

According to xAI’s pricing documentation, the standard global API rates are:

Prompt sizeInput / 1MCached input / 1MOutput / 1M
Below 200,000 tokens$2.00$0.50$6.00
200,000 tokens and above$4.00$1.00$12.00

Once the standard model’s prompt reaches the long-context threshold, the higher rates apply to all tokens in that request. Do not budget a 400,000-token request by charging only the final 200,000 tokens at the higher price.

As a simple calculation, 100,000 uncached input tokens plus 10,000 billed output tokens cost $0.26 at the standard short-context rates. This illustration excludes extra tool charges and assumes the output total already includes any billed reasoning. An agent can make many requests before finishing one job, so a single-request example is not a project estimate.

Cache behavior can also change the comparison. Anthropic lists Fable 5.1 cache reads at $0.25 per million tokens, below Grok’s $0.50 short-context cached-input rate, despite Fable’s higher uncached and output prices. For agents that repeatedly read the same large prefix, inspect the full token mix.

Grok 4.7 Fast is a separate serving option in Cursor and Grok Build, not a public xAI API model. Faster token generation also does not guarantee a proportionate reduction in a job’s total time: search, tool execution, tests and retries consume time too.

What the evaluations can and cannot tell you

Cursor describes its benchmark as work derived from real development sessions, with ambiguous requests and tasks that span files and tools. That makes it relevant to coding agents. It remains a particular environment with particular grading rules.

Cursor is also a Grok development and distribution partner. Its published evaluation is valuable corroboration of the numbers, but it should not be described as an entirely unaffiliated assessment. Cursor previously disclosed training-data overlap for Grok 4.5 and said that data had been removed for future models. That historical disclosure does not establish contamination in Grok 4.7.

Do not combine CursorBench 3.2 with 4.0, or Terminal-Bench 2.1 with 4.0, and call the resulting ranking an upgrade chart. Changed tasks can lower scores even when the model improves. Use the same benchmark version, the same tools where possible, and explicit effort settings.

For a purchase decision, run a small set of tasks you already know how to judge. Record whether the output passes existing tests, whether the agent finishes without rescue, what a reviewer must repair, total elapsed time and the complete bill. Include failures in the average. A cheap successful demonstration can hide an expensive failure rate.

Frequently asked questions

Is Grok 4.7 better than Claude?

There is no single answer across Claude models and tasks. CursorBench places Fable 5.1 max above Grok 4.7 xhigh; Opus 5 max is close in score and more expensive per evaluated task. Other evaluations produce different orderings.

Does Grok 4.7 beat GPT-6 Astra?

xAI reports a higher Grok score in its GDPval chart at the named settings. The main launch table compares Grok with GPT-5.6 Sol. Neither observation proves a general victory over Astra.

Is the 2.1-trillion-parameter claim confirmed?

The official model and launch pages checked for this article do not disclose a parameter count. Treat a numerical architecture claim as unconfirmed unless a primary technical source publishes it.

Should I switch from Grok 4.6?

Grok 4.7 is worth a controlled trial on existing tasks: its published coding results improve at matched high and xhigh settings, and the standard API rate card is unchanged. Keep your current setup available until the new model passes your own regression checks.