Updated August 26, 2026: added query-matched benchmark answers, separated Grok 4 from Grok 4 Heavy, and labelled model variants, tools, and test conditions using xAI's official releases.
Verdict: Grok 4's published results are strong, but the headline figures mix Grok 4, Grok 4 Heavy, tool-enabled runs, no-tool runs, and later evaluation tables. Treat them as a scorecard to interpret—not a single proof that Grok 4 is best for every workload.
Grok 4 benchmark scores at a glance
The table below uses numbers xAI published in its Grok 4 launch and later Grok 4 Fast comparison. The later table is particularly useful because it presents Grok 4 scores in plain text under stated no-tool conditions.
| Benchmark | Model | Published score | Conditions/source | What it measures |
|---|---|---|---|---|
| USAMO 2025 | Grok 4 Heavy | 61.9% | xAI Grok 4 launch | Olympiad-style mathematical proofs |
| Humanity's Last Exam, text-only subset | Grok 4 Heavy | 50.7% | xAI launch; Heavy and test-time compute | Broad expert-level knowledge and reasoning |
| ARC-AGI V2 | Grok 4 | 15.9% | xAI Grok 4 launch | Novel abstraction and reasoning |
| GPQA Diamond | Grok 4 | 87.5% | Later xAI comparison table | Graduate-level science questions |
| AIME 2025 | Grok 4 | 91.7% | Later xAI table; no tools | Competition mathematics |
| HMMT 2025 | Grok 4 | 90.0% | Later xAI table; no tools | Competition mathematics |
| HLE | Grok 4 | 25.4% | Later xAI table; no tools | Expert-level multidisciplinary questions |
| LiveCodeBench, Jan–May | Grok 4 | 79.0% | Later xAI comparison table | Competitive code generation |
These figures are xAI-reported. A serious comparison should reproduce the same benchmark version, prompt, tool policy, pass count, and scoring method before treating small differences as meaningful.
What was Grok 4 Heavy's USAMO score?
xAI reported 61.9% on USAMO 2025 for Grok 4 Heavy. USAMO is not a multiple-choice arithmetic test; it evaluates proof-oriented mathematical reasoning. The "Heavy" label matters because xAI describes Heavy as using parallel test-time compute to consider multiple hypotheses. A Grok 4 Heavy result should not be silently attributed to the standard Grok 4 model.
What was Grok 4's AIME 2025 score?
In xAI's later Grok 4 Fast comparison, standard Grok 4 scored 91.7% on AIME 2025 without tools. That table also reported 90.0% on HMMT 2025. These later plain-text figures are easier to audit than chart-only launch graphics, but they still reflect xAI's own evaluation setup.
If another page reports 100%, check whether it is referring to a different model, pass count, thinking-token budget, or tool condition. Benchmark names are often reused across materially different setups.
What was Grok 4's GPQA score?
xAI's later comparison table reports 87.5% on GPQA Diamond for Grok 4. GPQA Diamond is a difficult subset of graduate-level science questions. It is more relevant to scientific reasoning than to everyday chat quality, writing style, or software engineering reliability.
Humanity's Last Exam: 25.4% or 50.7%?
Both numbers appear in official xAI material, but they describe different conditions:
- 25.4% is the Grok 4 no-tool result in a later comparison table.
- 50.7% is the Grok 4 Heavy result on the text-only subset highlighted at launch.
The gap is a clean example of why a score without its model variant and test condition is incomplete. Heavy uses additional test-time compute, and the launch also discusses tool-enabled HLE evaluation.
ARC-AGI V2 and tool use
xAI reported 15.9% on ARC-AGI V2 for Grok 4 and described the model as natively trained to use tools such as code execution and web search. Tool use is part of the product's intended capability, but it complicates comparisons. A tool-enabled research run answers a different question from a closed-book no-tool benchmark.
When judging a model for work, use the condition that resembles the actual job. For research, tool use may be a feature. For a controlled reasoning comparison, it can be a confounder.
What about Math-500 and other requested scores?
Kingy's query data shows demand for a Grok 4 Math-500 score. We did not find a plain-text Math-500 result in the two primary xAI releases used for this update. Rather than copy a number from an unsourced chart, social post, or secondary recap, this page leaves the value unpublished until a primary xAI table or model card establishes it.
That editorial restraint is deliberate. A missing sourced number is better than a confident benchmark error.
How to read Grok 4 benchmark claims
- Name the model. Grok 4, Grok 4 Heavy, Grok 4 Fast, and later Grok versions are not interchangeable.
- Name the condition. Record tools, pass count, test-time compute, date range, and benchmark version.
- Separate vendor and independent results. xAI's numbers are useful primary evidence for what xAI claims, not independent replication.
- Match the benchmark to the workload. GPQA does not measure product UX; AIME does not measure factual freshness; LiveCodeBench does not measure maintainability.
- Avoid tiny-difference theatre. A one-point lead may disappear under a different prompt, seed, or scoring harness.
Does Grok 4's benchmark performance make it the best model?
No benchmark table can answer that alone. Grok 4 may be a strong fit when native tools, live search, large context, or xAI integration matter. Cost, latency, reliability, safety, data controls, and task-specific evals can reverse the decision.
For production selection, build a small evaluation set from your own work: representative prompts, acceptance criteria, cost limits, and failure categories. Benchmarks are a useful prior. Your eval is the decision.
FAQ
What was Grok 4 Heavy's USAMO 2025 score?
xAI reported 61.9% for Grok 4 Heavy.
What was Grok 4's AIME 2025 score?
xAI's later no-tool comparison table reports 91.7% for standard Grok 4.
What was Grok 4's GPQA Diamond score?
xAI's later comparison table reports 87.5%.
What was Grok 4's Humanity's Last Exam score?
The standard Grok 4 no-tool result was 25.4% in a later table. xAI highlighted 50.7% for Grok 4 Heavy on the text-only subset at launch.
Are the scores independently verified?
The values in this article are xAI-published. Independent replication may use different harnesses or conditions, so compare like with like.
Official sources
The Kingy Brief
Follow The Kingy Brief.
One consequential launch, one pricing, limit, or shutdown change, one hands-on test, one exact prompt or Test Pack, and one try / watch / skip verdict.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
