AI News

Grok 4 Benchmarks: USAMO, AIME, GPQA & HLE Scores

Updated August 26, 2026: added query-matched benchmark answers, separated Grok 4 from Grok 4 Heavy, and labelled model variants, tools, and test conditions using xAI's official releases.

Verdict: Grok 4's published results are strong, but the headline figures mix Grok 4, Grok 4 Heavy, tool-enabled runs, no-tool runs, and later evaluation tables. Treat them as a scorecard to interpret—not a single proof that Grok 4 is best for every workload.

Grok 4 benchmark scores at a glance

The table below uses numbers xAI published in its Grok 4 launch and later Grok 4 Fast comparison. The later table is particularly useful because it presents Grok 4 scores in plain text under stated no-tool conditions.

Benchmark Model Published score Conditions/source What it measures
USAMO 2025 Grok 4 Heavy 61.9% xAI Grok 4 launch Olympiad-style mathematical proofs
Humanity's Last Exam, text-only subset Grok 4 Heavy 50.7% xAI launch; Heavy and test-time compute Broad expert-level knowledge and reasoning
ARC-AGI V2 Grok 4 15.9% xAI Grok 4 launch Novel abstraction and reasoning
GPQA Diamond Grok 4 87.5% Later xAI comparison table Graduate-level science questions
AIME 2025 Grok 4 91.7% Later xAI table; no tools Competition mathematics
HMMT 2025 Grok 4 90.0% Later xAI table; no tools Competition mathematics
HLE Grok 4 25.4% Later xAI table; no tools Expert-level multidisciplinary questions
LiveCodeBench, Jan–May Grok 4 79.0% Later xAI comparison table Competitive code generation

These figures are xAI-reported. A serious comparison should reproduce the same benchmark version, prompt, tool policy, pass count, and scoring method before treating small differences as meaningful.

What was Grok 4 Heavy's USAMO score?

xAI reported 61.9% on USAMO 2025 for Grok 4 Heavy. USAMO is not a multiple-choice arithmetic test; it evaluates proof-oriented mathematical reasoning. The "Heavy" label matters because xAI describes Heavy as using parallel test-time compute to consider multiple hypotheses. A Grok 4 Heavy result should not be silently attributed to the standard Grok 4 model.

What was Grok 4's AIME 2025 score?

In xAI's later Grok 4 Fast comparison, standard Grok 4 scored 91.7% on AIME 2025 without tools. That table also reported 90.0% on HMMT 2025. These later plain-text figures are easier to audit than chart-only launch graphics, but they still reflect xAI's own evaluation setup.

If another page reports 100%, check whether it is referring to a different model, pass count, thinking-token budget, or tool condition. Benchmark names are often reused across materially different setups.

What was Grok 4's GPQA score?

xAI's later comparison table reports 87.5% on GPQA Diamond for Grok 4. GPQA Diamond is a difficult subset of graduate-level science questions. It is more relevant to scientific reasoning than to everyday chat quality, writing style, or software engineering reliability.

Humanity's Last Exam: 25.4% or 50.7%?

Both numbers appear in official xAI material, but they describe different conditions:

  • 25.4% is the Grok 4 no-tool result in a later comparison table.
  • 50.7% is the Grok 4 Heavy result on the text-only subset highlighted at launch.

The gap is a clean example of why a score without its model variant and test condition is incomplete. Heavy uses additional test-time compute, and the launch also discusses tool-enabled HLE evaluation.

ARC-AGI V2 and tool use

xAI reported 15.9% on ARC-AGI V2 for Grok 4 and described the model as natively trained to use tools such as code execution and web search. Tool use is part of the product's intended capability, but it complicates comparisons. A tool-enabled research run answers a different question from a closed-book no-tool benchmark.

When judging a model for work, use the condition that resembles the actual job. For research, tool use may be a feature. For a controlled reasoning comparison, it can be a confounder.

What about Math-500 and other requested scores?

Kingy's query data shows demand for a Grok 4 Math-500 score. We did not find a plain-text Math-500 result in the two primary xAI releases used for this update. Rather than copy a number from an unsourced chart, social post, or secondary recap, this page leaves the value unpublished until a primary xAI table or model card establishes it.

That editorial restraint is deliberate. A missing sourced number is better than a confident benchmark error.

How to read Grok 4 benchmark claims

  1. Name the model. Grok 4, Grok 4 Heavy, Grok 4 Fast, and later Grok versions are not interchangeable.
  2. Name the condition. Record tools, pass count, test-time compute, date range, and benchmark version.
  3. Separate vendor and independent results. xAI's numbers are useful primary evidence for what xAI claims, not independent replication.
  4. Match the benchmark to the workload. GPQA does not measure product UX; AIME does not measure factual freshness; LiveCodeBench does not measure maintainability.
  5. Avoid tiny-difference theatre. A one-point lead may disappear under a different prompt, seed, or scoring harness.

Does Grok 4's benchmark performance make it the best model?

No benchmark table can answer that alone. Grok 4 may be a strong fit when native tools, live search, large context, or xAI integration matter. Cost, latency, reliability, safety, data controls, and task-specific evals can reverse the decision.

For production selection, build a small evaluation set from your own work: representative prompts, acceptance criteria, cost limits, and failure categories. Benchmarks are a useful prior. Your eval is the decision.

FAQ

What was Grok 4 Heavy's USAMO 2025 score?

xAI reported 61.9% for Grok 4 Heavy.

What was Grok 4's AIME 2025 score?

xAI's later no-tool comparison table reports 91.7% for standard Grok 4.

What was Grok 4's GPQA Diamond score?

xAI's later comparison table reports 87.5%.

What was Grok 4's Humanity's Last Exam score?

The standard Grok 4 no-tool result was 25.4% in a later table. xAI highlighted 50.7% for Grok 4 Heavy on the text-only subset at launch.

Are the scores independently verified?

The values in this article are xAI-published. Independent replication may use different harnesses or conditions, so compare like with like.

Official sources