AI News

Can Claude Fable 5.1 and Mythos 5.1 Do Real Science?

Verdict: Claude Fable 5.1 is a serious candidate for computational research, but it is not yet defensible to call it the best AI model for scientific research. Anthropic reports a 52.6% score on Terminal-Bench-Science 0.1—far above the public leaderboard—but that Fable 5.1 run is not yet on the independent leaderboard and has not been replicated. Mythos 5.1’s protein-design result is more consequential and less verifiable: outside organizations reportedly ran the assays, yet Anthropic has not released the campaign’s sequences, prompts, protocols, raw measurements, or named validation reports.

The most supportable conclusion on launch day is narrower. Fable 5.1 looks unusually capable at long, tool-using computational tasks. Mythos 5.1 may extend that capability into advanced biology. Neither has yet cleared the standard that matters in science: a third party reproducing the work from released methods and artifacts.

Kingy.ai’s recommendation: Use Fable 5.1 as a high-end research copilot only inside a workflow with locked sources, executable checks, recorded environments, and expert review. Treat Mythos 5.1’s biology projects as Anthropic-reported evidence—not operational guidance—until qualified independent teams can inspect and reproduce them.

The short answer

  • Best new candidate for scientific terminal work: Fable 5.1, based on Anthropic’s internal Terminal-Bench-Science result.
  • Best independently listed result on that benchmark today: Claude Opus 5 at 30.0% on the public leaderboard.
  • Most interesting science artifact released: the Venus elevation map, because the 3.9 GB GeoTIFF, README, licence, and validation summary are public.
  • Most important claim still missing reproducibility materials: Mythos 5.1’s protein-binder campaign.
  • Best first-party evidence for theoretical science: Gemini 3.1 Deep Think, although its physics and chemistry tests are not directly comparable with Terminal-Bench-Science.
  • Best open-weight comparison in the public science-terminal set: Kimi K3, whose model files are public and whose official leaderboard score is 7.1%.

This is a launch-day evidence audit, not a paid cross-model benchmark. Kingy.ai did not run Fable 5.1, Mythos 5.1, or their rivals through the seven-test protocol described below. Where a result is absent, we write “not measured”—not a surrogate score.

One model, two access and safety systems

Fable 5.1 and Mythos 5.1 are not two independently trained frontier models. Anthropic says they use the same underlying model. Fable 5.1 is the generally available product with production safeguards; Mythos 5.1 is the more permissive configuration offered only to approved organizations under restricted access.

That distinction makes raw-versus-system comparisons essential. A published “Fable 5.1” result may include Claude Code, a shell, files, a package repository, classifiers, interventions, or fallback to another model. A Mythos result may use the same base model with fewer interruptions. Search, retrieval, code execution, orchestration, and proprietary data can change the outcome as much as the model weights.

Anthropic’s launch post and 212-page system card are unusually detailed about this boundary. They also reveal why a single benchmark number is not a clean measure of the underlying model: on intervention-affected biology tasks, Fable sometimes fell back to Opus 5, and some intercepted attempts received zero credit.

Scientific benchmark comparison—with the caveats attached

The table below separates like-for-like public Terminal-Bench-Science results from vendor tests that measure different things.

Model or system Scientific evidence available Reported result Methodology warning Evidence status
Fable 5.1 + Claude Code Terminal-Bench-Science 0.1 52.6% Anthropic internal run; 10 trials per task; production interventions/fallbacks; not on public leaderboard at review time Anthropic-reported
Claude Opus 5 + Claude Code Terminal-Bench-Science 0.1 30.0% Public leaderboard uses 3 trials per task; agent and harness are part of the system Independently listed benchmark run
GPT-5.6 Sol + Codex CLI Terminal-Bench-Science 0.1 22.4% Public leaderboard; system includes Codex CLI and terminal tools Independently listed benchmark run
Fable 5 + Claude Code Terminal-Bench-Science 0.1 21.4% Public score differs from Anthropic system-card figure of 24.7%; run protocol matters Independently listed benchmark run
Kimi K3 + Kimi Code Terminal-Bench-Science 0.1 7.1% Open weights, but the listed result is still an agent-system score Independently listed benchmark run
Grok 4.6 + Grok Build Terminal-Bench-Science 0.1 7.1% Agent-system score, not a bare-model test Independently listed benchmark run
Mythos 5.1 Anthropic internal life-science and chemistry suites 44.1%–90.3% across selected suites Several suites are private or internally developed; protocols and graders differ Anthropic-reported
Gemini 3.1 Deep Think Physics, chemistry, condensed-matter, academic reasoning 87.7% IPhO; 82.8% IChO; 50.5% CMT First-party benchmark suite; not a terminal reproduction test Google-reported
Gemini 3.7 Flash General agentic and multimodal tests No comparable TBS score Optimized for agentic scale; launch page does not publish this science benchmark Not comparable
Meta Muse Spark 1.1 General coding/agent claims No comparable TBS score Superseded by Muse Spark 1.2; no science-specific evidence located Not comparable
DeepSeek V4-Pro General API/model release evidence No comparable TBS score No public run in the inspected Terminal-Bench-Science leaderboard Not measured
Qwen3.8-Max General model documentation No comparable TBS score No public run in the inspected Terminal-Bench-Science leaderboard Not measured

Sources: the open Terminal-Bench-Science leaderboard and announcement, its Apache-licensed repository, Google DeepMind’s Gemini 3.1 Deep Think page, and each provider’s model documentation. Numbers from incompatible evaluations should not be ranked as if they shared a denominator.

Why the 52.6% is impressive—and not conclusive

Terminal-Bench-Science is closer to real computational work than a multiple-choice exam. Scientists proposed tasks; accepted tasks run in containers; the agent receives natural-language instructions and must create artifacts such as code, proofs, simulations, or processed data that hidden tests then grade. The project says 70 tasks survived a pipeline that began with 920 proposals.

Anthropic reports that Fable 5.1 completed 52.6% under its internal protocol, compared with 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. The system card specifies 10 trials per task—700 task attempts—and standard errors around 3.5 to 4.5 percentage points. The public leaderboard, however, lists three trials per task and did not include Fable 5.1 when this article was reviewed. It listed Opus 5 at 30.0%, GPT-5.6 Sol at 22.4%, Kimi K3 at 7.1%, and Grok 4.6 at 7.1%.

That does not make Anthropic’s result false. It means the strongest claim has one source, a different sampling protocol, and no public run record yet. Anthropic also sponsors the benchmark with API credits, and Anthropic personnel appear among project contributors or advisers. The benchmark is openly governed and independently maintained, but “independent” should not be mistaken for “entirely arm’s-length from every vendor.”

Contamination cannot be proved from public evidence. It also cannot be dismissed. The benchmark repository, proposals, and development process are public, and Fable 5.1’s stated knowledge cutoff is June 2026. Anthropic separately warns that its June 2026 ArXivMath set may overlap training data. A clean follow-up should use newly held-out tasks created after the cutoff, hashed before model access, and run by a neutral party.

Anthropic science-claim evidence ledger

Claim What Anthropic says What is public Independent check or replication What remains unverifiable Evidence strength
Terminal-Bench-Science gain Fable 5.1 scored 52.6% Benchmark, tasks, harness, public leaderboard, system-card method summary Benchmark organization is external; Fable 5.1 result itself not publicly listed or replicated Exact run logs, prompts, failed attempts, intervention trace, costs Medium
Protein-binder design Nearly 50% hit rate across 12 targets; three targets reportedly beat the best Adaptyv competition affinities by 10× Launch description; Adaptyv’s separate competition data are public Anthropic says two external organizations performed tests, but they are unnamed in the release Sequences, target list, prompts, tool versions, assay protocols, raw curves, negatives, full selection process Low–medium
Venus elevation map One-third of Venus mapped at 300 m grid; 22%–23% lower RMS error on reported validations 3.9 GB GeoTIFF, README, DOI, CC BY 4.0 licence, limitations and validation summary No independent replication found; artifact can be inspected Training code, weights, prompts, end-to-end pipeline, exact environment Medium–high for artifact; medium for method
Biology-model GPU kernels Seven open-source models accelerated 1.4×–2.5× on H100 with identical outputs Launch summary only; Anthropic says code will be open-sourced None located Kernel code, commits, test vectors, tolerance definition, profiler traces, hardware/software environment Low
Cost and speed Fable 5.1 agentic workloads can cost up to 45% less; cache reads cut to $0.25/M Token prices and summary of four weeks of Anthropic usage No independent science-workflow cost study located Workload mix, latency distribution, retries, total tool costs, human time Medium for prices; low–medium for savings
Biology/chemistry suites Mythos 5.1 leads or remains near Opus 5 on selected suites Detailed system-card tables and some benchmark descriptions Many suites or graders are private/internal; no 5.1 replication located Raw outputs, per-item errors, leakage analysis, complete evaluator code Low–medium

“Evidence strength” describes auditability, not a judgement about whether the claim is true.

The protein-binder result is promising, but not reproducible yet

Anthropic reports that Mythos 5.1 used open-source protein tools to design binders for 12 targets, with nearly half of tested designs registering as hits. It also says that three targets produced affinities ten times better than the best entries in comparable Adaptyv competitions, and that the designs shown in its launch graphic were lab-confirmed.

Wet-lab testing is stronger than a model grading its own output. The missing chain of custody still matters. Anthropic does not name the two external testing organizations in the launch post, link signed reports, release the designed sequences, disclose the complete target and negative set, or publish assay protocols and raw measurements. Without those materials, outsiders cannot estimate selection bias, check whether “nearly 50%” uses all attempted designs as the denominator, or reproduce the comparison.

The Adaptyv ProteinBase competition archive shows that public, standardized wet-lab leaderboards are feasible. It does not independently authenticate Anthropic’s new campaign. Until a full campaign record is deposited, the correct wording is Anthropic-reported wet-lab validation by external organizations, not “independently replicated protein discovery.”

This article does not provide protein sequences, pathogen work, synthesis instructions, or operational laboratory protocols. Mythos’s advanced biology capability is an evidence-review subject here, not a how-to.

The Venus map is the strongest released artifact

The Zenodo record for the Venusian elevation product is the clearest example of something an independent researcher can actually download. It contains a 3.9 GB GeoTIFF covering roughly 35% of Venus at a 300 m grid, a README, citation information, a CC BY 4.0 licence, and a DOI. The accompanying notes report 23% lower RMS error on covered radargram areas and 22% lower error across 21 held-out boxes compared with the baseline used by the project.

The deposit also states important limitations: smaller features can be damped, low-confidence regions are flagged, and the work is a research product rather than a peer-reviewed map. The credited author is an Arizona State University researcher working through Anthropic’s STEM Fellows program, with Claude credited in the record; Anthropic itself is not presented as endorsing the scientific conclusions.

What is missing is enough to prevent end-to-end reproduction: no training code, weights, prompts, preprocessing scripts, or frozen environment. An expert can inspect the map and repeat some comparisons against public inputs, but cannot rebuild it from the release alone. That makes this a released result with partial method transparency, not a fully reproducible computational paper.

The GPU-kernel claim needs code, not adjectives

Anthropic reports that Mythos 5.1 wrote custom GPU kernels and caching logic for seven open-source biology models, achieving 1.4× to 2.5× speedups on an Nvidia H100 while preserving identical outputs. It estimates 30% to 60% lower inference cost and says the work will be open-sourced “soon.”

At review time, the launch post did not link the promised repository. A credible reproduction needs exact commits, compiler and CUDA versions, tensor shapes and dtypes, warm-up policy, memory constraints, tolerance rules, profiler traces, and a test suite. “Identical” can mean bitwise identity or agreement within a tolerance; those are not interchangeable. Cost estimates based on cloud list prices are useful planning inputs, not measured total-cost studies.

Until the code appears, this is the weakest of the headline project claims.

What the system-card science evaluations really show

Mythos 5.1 is strong across several Anthropic evaluations, but it does not sweep them. It scores 90.3% on the human-solvable BioMystery set, while Opus 5 reaches 91.4%. On the harder BioMystery set, Opus 5 scores 51.8% and Mythos 5.1 scores 44.1%. Mythos 5.1 leads the reported ProteinGym Hard comparison at 49.3%, organic chemistry V2 at 69.2%, and protocol troubleshooting at 70.2%. Opus 5 leads protocol understanding at 80.0% versus Mythos 5.1 at 77.2%.

These results support “competitive with the frontier across selected life-science tasks.” They do not support “universally best for biology.” Several sets are internal, some graders were updated, tool access varies, and not every rival appears in every table. The system card also reports that Mythos 5.1 abstains less in closed-book factual testing, producing both more correct and more incorrect answers. That trade-off is especially relevant in literature synthesis, where confident unsupported claims can be costlier than a refusal.

Capability versus evidence-strength matrix

Model Likely computational capability Science-specific public evidence Reproducibility access Main limitation
Fable 5.1 Very high Strong internal result; public benchmark not yet updated Closed model; open benchmark Strongest score not independently logged
Mythos 5.1 Very high, fewer safety interruptions Detailed vendor tables and project claims Restricted model; most project artifacts closed Access and artifact gap
Opus 5 High Best public TBS 0.1 score reviewed Closed model; public benchmark record Expensive; still an agent-system result
GPT-5.6 Sol High Public TBS run plus broad official tool support Closed model; hosted shell/code tools Science-specific evidence thinner than Fable’s launch package
Gemini 3.1 Deep Think Very high theoretical reasoning Strong physics, chemistry and math vendor evaluations Closed specialized mode Different tests; limited like-for-like reproduction evidence
Gemini 3.7 Flash High-throughput agentic work Little science-specific evidence Closed API/product Speed is not scientific validity
Grok 4.6 High general agent capability Public TBS score of 7.1% Closed model; public benchmark record Low score on the one comparable science-terminal run
Muse Spark 1.1 General multimodal/agentic No meaningful science-specific set located Closed API Superseded by 1.2
Kimi K3 Strong open agent model Public TBS score of 7.1% Open weights and model card Reproduction requires substantial hardware
DeepSeek V4-Pro Strong cost-oriented frontier API No comparable science run located API; no open weights for this endpoint Science conclusion would be extrapolation
Qwen3.8-Max Large-context frontier API No comparable science run located API documentation; not an open-weight release Science conclusion would be extrapolation

Provider references: GPT-5.6 Sol, Gemini 3.7 Flash, Grok 4.6, Muse Spark, Kimi K3 weights and model card, DeepSeek V4-Pro, and Qwen3.8-Max.

Reproduction results by model: what Kingy measured in v1

Model Literature audit Figure/table extraction Equation-to-code Public-data reproduction Confounder test Blind hypothesis review Long terminal task V1 status
Fable 5.1 Not run
Mythos 5.1 Not run; restricted access
Opus 5 Not run
GPT-5.6 Sol Not run
Gemini 3.1 Deep Think Not run
Gemini 3.7 Flash Not run
Grok 4.6 Not run
Muse Spark 1.1 Not run; superseded
Kimi K3 Not run
DeepSeek V4-Pro Not run
Qwen3.8-Max Not run

Why publish an empty table? Because it marks the exact boundary between audited claims and original evidence. No Kingy.ai hands-on ranking exists yet. The launch materials do not provide comparable time, token, retry, citation, or human-intervention logs for the full model set.

Citation and numerical-error comparison

Model/system Citation precision Citation recall Unsupported-claim rate Numerical agreement with a hidden reference What can be said today
Fable 5.1 Not measured Not measured Not measured Not measured by Kingy Terminal artifact pass rate is high in Anthropic’s run; it is not a citation audit
Mythos 5.1 Not measured Not measured Not measured Not measured by Kingy Fewer abstentions in closed-book testing may increase both useful answers and errors
Opus 5 Not measured Not measured Not measured Not measured by Kingy Public terminal score exists; scientific numerical fidelity is a different metric
GPT-5.6 Sol Not measured Not measured Not measured Not measured by Kingy Public terminal score exists; official tool availability does not prove source quality
Gemini 3.1 Deep Think Not measured Not measured Not measured Not measured by Kingy High first-party theoretical-science scores; no common citation protocol
Other required rivals Not measured Not measured Not measured Not measured by Kingy Current evidence is insufficient for a numerical or citation ranking

Any table that turns benchmark pass rates into citation accuracy or numerical agreement would be misleading. Those are separate dependent variables and need separate hidden-answer tests.

The safe hands-on protocol that could settle this

Kingy.ai’s next evidence round should freeze all prompts, source packets, containers, reference answers, tool permissions, and stopping rules before revealing model identities to graders.

  1. Authoritative-source literature synthesis. Give each model the same curated corpus. Score claim-level citation precision and recall, source-entailment, and unsupported claims.
  2. Scientific figure and table extraction. Use public papers with concealed numerical answers. Score exact cells, units, uncertainty, axis interpretation, and derived values.
  3. Equation-to-code implementation. Require executable functions and test them against analytical solutions, edge cases, dimensional checks, and randomized inputs.
  4. Published computational reproduction. Reproduce one result from public data in a fixed container. Measure numerical agreement, environment completeness, retries, and human intervention.
  5. Statistical analysis with planted traps. Include missing-not-at-random data, batch effects, leakage, collinearity, multiple comparisons, and a tempting but wrong causal story.
  6. Hypothesis generation and ranking. Have qualified experts blind-score novelty, plausibility, discriminating power, feasibility, and awareness of prior work. Keep advanced biology non-operational.
  7. Long-running scientific terminal task. Run every available model in the same fixed environment with a wall-clock cap and hidden tests.

For every trial, log model version, system prompt, tools, retrieval sources, container digest, time, tokens, cost, retries, interventions, failures, and the final artifact. Use multiple seeds. Report medians and uncertainty, not only the best run.

Timeline: artifacts released versus independently replicated

Date Event Artifact state Independent replication state
June 2026 Fable 5.1 stated knowledge cutoff Model documentation Not applicable
Before launch Terminal-Bench-Science task development and public repository activity Benchmark and harness public Public runs exist for earlier models
29 Aug 2026 Venus elevation map v1 deposited on Zenodo GeoTIFF, README, DOI and licence public None found by 1 Sep
1 Sep 2026 Fable 5.1 and Mythos 5.1 announced Launch post and system card public No independent 5.1 replication found
1 Sep 2026 Protein-binder and GPU-kernel projects described Summary only; kernel code promised None found
Next required milestone Neutral TBS 0.1/0.2 run plus complete logs Pending Pending
Next required milestone Protein campaign and kernel repositories released Pending Pending

The absence of a replication three days after the Venus deposit or on launch day is not suspicious. It is simply too early. “No replication found” is a timestamped status, not a verdict.

Best model by scientific workflow

Scientific workflow Best current choice Why Required guardrail
Long computational task in a terminal Fable 5.1 as the leading candidate; Opus 5 as the public-score leader 52.6% Anthropic-reported versus 30.0% publicly listed Re-run under one neutral harness before declaring a winner
Theoretical physics, chemistry, or advanced math Gemini 3.1 Deep Think Strong first-party domain evaluations and specialized reasoning Expert proof-checking; do not equate olympiad scores with research discovery
Literature synthesis No winner yet No comparable claim-level citation audit Locked corpus, claim-level citation grading, unsupported-claim threshold
Reproducing a public computational paper Fable 5.1 is the most promising unverified choice Terminal evidence is directionally relevant Container digest, hidden numerical checks, complete logs
High-volume multimodal extraction Gemini 3.7 Flash as a cost/latency candidate Multimodal and agentic-scale design Hidden-answer extraction test; sample-level QA
Open-weight experimentation Kimi K3 Public weights, model card, and a comparable public TBS run Budget for substantial hardware; freeze inference stack
Advanced biology research No general-purpose recommendation Mythos 5.1 evidence is promising but restricted and incomplete Qualified institution, legal authority, safeguards, biosafety review, non-operational scope
Budget-sensitive API research DeepSeek V4-Pro or Qwen3.8-Max as evaluation candidates Attractive frontier positioning and long context Do not select until they pass the same science protocol
New Meta evaluation Muse Spark 1.2, not 1.1 1.1 is already superseded Treat as a fresh model; no inherited science score

What would change the verdict

Five releases would materially strengthen the case:

  1. A neutral, versioned Fable 5.1 run on Terminal-Bench-Science with per-task logs, costs, interventions, and the public three-trial protocol.
  2. The full protein-binder campaign record: all attempted sequences, targets, selection steps, tool versions, protocols, raw measurements, negatives, and named external reports.
  3. End-to-end Venus training code, frozen environment, preprocessing scripts, prompts, and a second group’s reconstruction.
  4. The seven GPU-kernel patches with test vectors, tolerances, profiler traces, and independent hardware results.
  5. A controlled citation, numerical-reproduction, confounder, and expert-blind study across the full comparison set.

Final verdict

Can Claude Fable 5.1 and Mythos 5.1 do real science? They can evidently contribute to real scientific work, and Fable 5.1 may be a major step forward for computational research agents. The public evidence does not yet show that either system can independently produce reproducible, correctly sourced, numerically valid research at a level that survives neutral expert scrutiny.

The Venus map is real and downloadable, but only partly reproducible. Terminal-Bench-Science is rigorous and open, but Fable 5.1’s headline run is still Anthropic-reported. The protein-binder work reportedly reached external wet labs, but its decisive evidence remains private. The GPU speedups remain a promise until the code lands.

So the answer to the reader’s purchasing or workflow decision is practical: pilot Fable 5.1, do not trust it alone, and measure the outputs that science actually depends on. For Mythos 5.1, wait for institutional access and substantially better artifact disclosure. Benchmarks can nominate a model for a trial; reproduction earns it a place in the lab.

FAQ

Is Fable 5.1 the best AI model for scientific research?

Not yet on public evidence. Anthropic reports the strongest Terminal-Bench-Science result, but the public leaderboard’s highest listed score at review time belongs to Opus 5. Fable 5.1 needs a neutral logged run and broader citation and reproduction testing.

Are Fable 5.1 and Mythos 5.1 different models?

Anthropic says they share the same underlying model. Fable 5.1 is the generally available safeguarded system; Mythos 5.1 is a more permissive configuration for approved organizations.

Was Mythos 5.1’s protein design independently validated?

Anthropic reports wet-lab tests by two external organizations. The organizations and full campaign record were not linked in the launch materials, and no independent replication was found. “Externally tested” is therefore supportable; “independently replicated” is not.

Is the Venus elevation map public?

Yes. The GeoTIFF, README, DOI, and licence are public on Zenodo. The training code, weights, prompts, and full reconstruction pipeline are not included, so the result is inspectable but not end-to-end reproducible.

How does GPT-5.6 Sol compare?

GPT-5.6 Sol has a public Terminal-Bench-Science score of 22.4% and broad hosted research tools. Fable 5.1’s Anthropic-reported 52.6% is much higher, but the two numbers come from different published run records and need a neutral same-protocol retest.

What is the safest way to use AI for scientific research?

Use a locked authoritative source set, keep executable tests and reference answers hidden, record the environment and every intervention, require claim-level citations, and have a qualified expert review the result. Do not treat model confidence as evidence.


Disclosure: This is an independent evidence audit. Kingy.ai did not receive model access, payment, or unpublished results from Anthropic for this article. No paid cross-model hands-on benchmark was run for v1. All Fable 5.1 and Mythos 5.1 project claims are labelled Anthropic-reported unless a separately released artifact supports stronger wording. Prices, model availability, leaderboards, and artifacts can change after publication.

Last reviewed: September 1, 2026.

Related Kingy.ai coverage: Claude Fable 5 and Mythos 5 launch, Fable 5 benchmark caveats, the Kimi K3 distillation evidence, AI-agent security controls, and the Fable Jacobian evidence audit.