Verdict: Claude Fable 5.1 is a serious candidate for computational research, but it is not yet defensible to call it the best AI model for scientific research. Anthropic reports a 52.6% score on Terminal-Bench-Science 0.1—far above the public leaderboard—but that Fable 5.1 run is not yet on the independent leaderboard and has not been replicated. Mythos 5.1’s protein-design result is more consequential and less verifiable: outside organizations reportedly ran the assays, yet Anthropic has not released the campaign’s sequences, prompts, protocols, raw measurements, or named validation reports.
The most supportable conclusion on launch day is narrower. Fable 5.1 looks unusually capable at long, tool-using computational tasks. Mythos 5.1 may extend that capability into advanced biology. Neither has yet cleared the standard that matters in science: a third party reproducing the work from released methods and artifacts.
Kingy.ai’s recommendation: Use Fable 5.1 as a high-end research copilot only inside a workflow with locked sources, executable checks, recorded environments, and expert review. Treat Mythos 5.1’s biology projects as Anthropic-reported evidence—not operational guidance—until qualified independent teams can inspect and reproduce them.
The short answer
- Best new candidate for scientific terminal work: Fable 5.1, based on Anthropic’s internal Terminal-Bench-Science result.
- Best independently listed result on that benchmark today: Claude Opus 5 at 30.0% on the public leaderboard.
- Most interesting science artifact released: the Venus elevation map, because the 3.9 GB GeoTIFF, README, licence, and validation summary are public.
- Most important claim still missing reproducibility materials: Mythos 5.1’s protein-binder campaign.
- Best first-party evidence for theoretical science: Gemini 3.1 Deep Think, although its physics and chemistry tests are not directly comparable with Terminal-Bench-Science.
- Best open-weight comparison in the public science-terminal set: Kimi K3, whose model files are public and whose official leaderboard score is 7.1%.
This is a launch-day evidence audit, not a paid cross-model benchmark. Kingy.ai did not run Fable 5.1, Mythos 5.1, or their rivals through the seven-test protocol described below. Where a result is absent, we write “not measured”—not a surrogate score.
One model, two access and safety systems
Fable 5.1 and Mythos 5.1 are not two independently trained frontier models. Anthropic says they use the same underlying model. Fable 5.1 is the generally available product with production safeguards; Mythos 5.1 is the more permissive configuration offered only to approved organizations under restricted access.
That distinction makes raw-versus-system comparisons essential. A published “Fable 5.1” result may include Claude Code, a shell, files, a package repository, classifiers, interventions, or fallback to another model. A Mythos result may use the same base model with fewer interruptions. Search, retrieval, code execution, orchestration, and proprietary data can change the outcome as much as the model weights.
Anthropic’s launch post and 212-page system card are unusually detailed about this boundary. They also reveal why a single benchmark number is not a clean measure of the underlying model: on intervention-affected biology tasks, Fable sometimes fell back to Opus 5, and some intercepted attempts received zero credit.
Scientific benchmark comparison—with the caveats attached
The table below separates like-for-like public Terminal-Bench-Science results from vendor tests that measure different things.
| Model or system | Scientific evidence available | Reported result | Methodology warning | Evidence status |
|---|---|---|---|---|
| Fable 5.1 + Claude Code | Terminal-Bench-Science 0.1 | 52.6% | Anthropic internal run; 10 trials per task; production interventions/fallbacks; not on public leaderboard at review time | Anthropic-reported |
| Claude Opus 5 + Claude Code | Terminal-Bench-Science 0.1 | 30.0% | Public leaderboard uses 3 trials per task; agent and harness are part of the system | Independently listed benchmark run |
| GPT-5.6 Sol + Codex CLI | Terminal-Bench-Science 0.1 | 22.4% | Public leaderboard; system includes Codex CLI and terminal tools | Independently listed benchmark run |
| Fable 5 + Claude Code | Terminal-Bench-Science 0.1 | 21.4% | Public score differs from Anthropic system-card figure of 24.7%; run protocol matters | Independently listed benchmark run |
| Kimi K3 + Kimi Code | Terminal-Bench-Science 0.1 | 7.1% | Open weights, but the listed result is still an agent-system score | Independently listed benchmark run |
| Grok 4.6 + Grok Build | Terminal-Bench-Science 0.1 | 7.1% | Agent-system score, not a bare-model test | Independently listed benchmark run |
| Mythos 5.1 | Anthropic internal life-science and chemistry suites | 44.1%–90.3% across selected suites | Several suites are private or internally developed; protocols and graders differ | Anthropic-reported |
| Gemini 3.1 Deep Think | Physics, chemistry, condensed-matter, academic reasoning | 87.7% IPhO; 82.8% IChO; 50.5% CMT | First-party benchmark suite; not a terminal reproduction test | Google-reported |
| Gemini 3.7 Flash | General agentic and multimodal tests | No comparable TBS score | Optimized for agentic scale; launch page does not publish this science benchmark | Not comparable |
| Meta Muse Spark 1.1 | General coding/agent claims | No comparable TBS score | Superseded by Muse Spark 1.2; no science-specific evidence located | Not comparable |
| DeepSeek V4-Pro | General API/model release evidence | No comparable TBS score | No public run in the inspected Terminal-Bench-Science leaderboard | Not measured |
| Qwen3.8-Max | General model documentation | No comparable TBS score | No public run in the inspected Terminal-Bench-Science leaderboard | Not measured |
Sources: the open Terminal-Bench-Science leaderboard and announcement, its Apache-licensed repository, Google DeepMind’s Gemini 3.1 Deep Think page, and each provider’s model documentation. Numbers from incompatible evaluations should not be ranked as if they shared a denominator.
Why the 52.6% is impressive—and not conclusive
Terminal-Bench-Science is closer to real computational work than a multiple-choice exam. Scientists proposed tasks; accepted tasks run in containers; the agent receives natural-language instructions and must create artifacts such as code, proofs, simulations, or processed data that hidden tests then grade. The project says 70 tasks survived a pipeline that began with 920 proposals.
Anthropic reports that Fable 5.1 completed 52.6% under its internal protocol, compared with 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. The system card specifies 10 trials per task—700 task attempts—and standard errors around 3.5 to 4.5 percentage points. The public leaderboard, however, lists three trials per task and did not include Fable 5.1 when this article was reviewed. It listed Opus 5 at 30.0%, GPT-5.6 Sol at 22.4%, Kimi K3 at 7.1%, and Grok 4.6 at 7.1%.
That does not make Anthropic’s result false. It means the strongest claim has one source, a different sampling protocol, and no public run record yet. Anthropic also sponsors the benchmark with API credits, and Anthropic personnel appear among project contributors or advisers. The benchmark is openly governed and independently maintained, but “independent” should not be mistaken for “entirely arm’s-length from every vendor.”
Contamination cannot be proved from public evidence. It also cannot be dismissed. The benchmark repository, proposals, and development process are public, and Fable 5.1’s stated knowledge cutoff is June 2026. Anthropic separately warns that its June 2026 ArXivMath set may overlap training data. A clean follow-up should use newly held-out tasks created after the cutoff, hashed before model access, and run by a neutral party.
Anthropic science-claim evidence ledger
| Claim | What Anthropic says | What is public | Independent check or replication | What remains unverifiable | Evidence strength |
|---|---|---|---|---|---|
| Terminal-Bench-Science gain | Fable 5.1 scored 52.6% | Benchmark, tasks, harness, public leaderboard, system-card method summary | Benchmark organization is external; Fable 5.1 result itself not publicly listed or replicated | Exact run logs, prompts, failed attempts, intervention trace, costs | Medium |
| Protein-binder design | Nearly 50% hit rate across 12 targets; three targets reportedly beat the best Adaptyv competition affinities by 10× | Launch description; Adaptyv’s separate competition data are public | Anthropic says two external organizations performed tests, but they are unnamed in the release | Sequences, target list, prompts, tool versions, assay protocols, raw curves, negatives, full selection process | Low–medium |
| Venus elevation map | One-third of Venus mapped at 300 m grid; 22%–23% lower RMS error on reported validations | 3.9 GB GeoTIFF, README, DOI, CC BY 4.0 licence, limitations and validation summary | No independent replication found; artifact can be inspected | Training code, weights, prompts, end-to-end pipeline, exact environment | Medium–high for artifact; medium for method |
| Biology-model GPU kernels | Seven open-source models accelerated 1.4×–2.5× on H100 with identical outputs | Launch summary only; Anthropic says code will be open-sourced | None located | Kernel code, commits, test vectors, tolerance definition, profiler traces, hardware/software environment | Low |
| Cost and speed | Fable 5.1 agentic workloads can cost up to 45% less; cache reads cut to $0.25/M | Token prices and summary of four weeks of Anthropic usage | No independent science-workflow cost study located | Workload mix, latency distribution, retries, total tool costs, human time | Medium for prices; low–medium for savings |
| Biology/chemistry suites | Mythos 5.1 leads or remains near Opus 5 on selected suites | Detailed system-card tables and some benchmark descriptions | Many suites or graders are private/internal; no 5.1 replication located | Raw outputs, per-item errors, leakage analysis, complete evaluator code | Low–medium |
“Evidence strength” describes auditability, not a judgement about whether the claim is true.
The protein-binder result is promising, but not reproducible yet
Anthropic reports that Mythos 5.1 used open-source protein tools to design binders for 12 targets, with nearly half of tested designs registering as hits. It also says that three targets produced affinities ten times better than the best entries in comparable Adaptyv competitions, and that the designs shown in its launch graphic were lab-confirmed.
Wet-lab testing is stronger than a model grading its own output. The missing chain of custody still matters. Anthropic does not name the two external testing organizations in the launch post, link signed reports, release the designed sequences, disclose the complete target and negative set, or publish assay protocols and raw measurements. Without those materials, outsiders cannot estimate selection bias, check whether “nearly 50%” uses all attempted designs as the denominator, or reproduce the comparison.
The Adaptyv ProteinBase competition archive shows that public, standardized wet-lab leaderboards are feasible. It does not independently authenticate Anthropic’s new campaign. Until a full campaign record is deposited, the correct wording is Anthropic-reported wet-lab validation by external organizations, not “independently replicated protein discovery.”
This article does not provide protein sequences, pathogen work, synthesis instructions, or operational laboratory protocols. Mythos’s advanced biology capability is an evidence-review subject here, not a how-to.
The Venus map is the strongest released artifact
The Zenodo record for the Venusian elevation product is the clearest example of something an independent researcher can actually download. It contains a 3.9 GB GeoTIFF covering roughly 35% of Venus at a 300 m grid, a README, citation information, a CC BY 4.0 licence, and a DOI. The accompanying notes report 23% lower RMS error on covered radargram areas and 22% lower error across 21 held-out boxes compared with the baseline used by the project.
The deposit also states important limitations: smaller features can be damped, low-confidence regions are flagged, and the work is a research product rather than a peer-reviewed map. The credited author is an Arizona State University researcher working through Anthropic’s STEM Fellows program, with Claude credited in the record; Anthropic itself is not presented as endorsing the scientific conclusions.
What is missing is enough to prevent end-to-end reproduction: no training code, weights, prompts, preprocessing scripts, or frozen environment. An expert can inspect the map and repeat some comparisons against public inputs, but cannot rebuild it from the release alone. That makes this a released result with partial method transparency, not a fully reproducible computational paper.
The GPU-kernel claim needs code, not adjectives
Anthropic reports that Mythos 5.1 wrote custom GPU kernels and caching logic for seven open-source biology models, achieving 1.4× to 2.5× speedups on an Nvidia H100 while preserving identical outputs. It estimates 30% to 60% lower inference cost and says the work will be open-sourced “soon.”
At review time, the launch post did not link the promised repository. A credible reproduction needs exact commits, compiler and CUDA versions, tensor shapes and dtypes, warm-up policy, memory constraints, tolerance rules, profiler traces, and a test suite. “Identical” can mean bitwise identity or agreement within a tolerance; those are not interchangeable. Cost estimates based on cloud list prices are useful planning inputs, not measured total-cost studies.
Until the code appears, this is the weakest of the headline project claims.
What the system-card science evaluations really show
Mythos 5.1 is strong across several Anthropic evaluations, but it does not sweep them. It scores 90.3% on the human-solvable BioMystery set, while Opus 5 reaches 91.4%. On the harder BioMystery set, Opus 5 scores 51.8% and Mythos 5.1 scores 44.1%. Mythos 5.1 leads the reported ProteinGym Hard comparison at 49.3%, organic chemistry V2 at 69.2%, and protocol troubleshooting at 70.2%. Opus 5 leads protocol understanding at 80.0% versus Mythos 5.1 at 77.2%.
These results support “competitive with the frontier across selected life-science tasks.” They do not support “universally best for biology.” Several sets are internal, some graders were updated, tool access varies, and not every rival appears in every table. The system card also reports that Mythos 5.1 abstains less in closed-book factual testing, producing both more correct and more incorrect answers. That trade-off is especially relevant in literature synthesis, where confident unsupported claims can be costlier than a refusal.
Capability versus evidence-strength matrix
| Model | Likely computational capability | Science-specific public evidence | Reproducibility access | Main limitation |
|---|---|---|---|---|
| Fable 5.1 | Very high | Strong internal result; public benchmark not yet updated | Closed model; open benchmark | Strongest score not independently logged |
| Mythos 5.1 | Very high, fewer safety interruptions | Detailed vendor tables and project claims | Restricted model; most project artifacts closed | Access and artifact gap |
| Opus 5 | High | Best public TBS 0.1 score reviewed | Closed model; public benchmark record | Expensive; still an agent-system result |
| GPT-5.6 Sol | High | Public TBS run plus broad official tool support | Closed model; hosted shell/code tools | Science-specific evidence thinner than Fable’s launch package |
| Gemini 3.1 Deep Think | Very high theoretical reasoning | Strong physics, chemistry and math vendor evaluations | Closed specialized mode | Different tests; limited like-for-like reproduction evidence |
| Gemini 3.7 Flash | High-throughput agentic work | Little science-specific evidence | Closed API/product | Speed is not scientific validity |
| Grok 4.6 | High general agent capability | Public TBS score of 7.1% | Closed model; public benchmark record | Low score on the one comparable science-terminal run |
| Muse Spark 1.1 | General multimodal/agentic | No meaningful science-specific set located | Closed API | Superseded by 1.2 |
| Kimi K3 | Strong open agent model | Public TBS score of 7.1% | Open weights and model card | Reproduction requires substantial hardware |
| DeepSeek V4-Pro | Strong cost-oriented frontier API | No comparable science run located | API; no open weights for this endpoint | Science conclusion would be extrapolation |
| Qwen3.8-Max | Large-context frontier API | No comparable science run located | API documentation; not an open-weight release | Science conclusion would be extrapolation |
Provider references: GPT-5.6 Sol, Gemini 3.7 Flash, Grok 4.6, Muse Spark, Kimi K3 weights and model card, DeepSeek V4-Pro, and Qwen3.8-Max.
Reproduction results by model: what Kingy measured in v1
| Model | Literature audit | Figure/table extraction | Equation-to-code | Public-data reproduction | Confounder test | Blind hypothesis review | Long terminal task | V1 status |
|---|---|---|---|---|---|---|---|---|
| Fable 5.1 | — | — | — | — | — | — | — | Not run |
| Mythos 5.1 | — | — | — | — | — | — | — | Not run; restricted access |
| Opus 5 | — | — | — | — | — | — | — | Not run |
| GPT-5.6 Sol | — | — | — | — | — | — | — | Not run |
| Gemini 3.1 Deep Think | — | — | — | — | — | — | — | Not run |
| Gemini 3.7 Flash | — | — | — | — | — | — | — | Not run |
| Grok 4.6 | — | — | — | — | — | — | — | Not run |
| Muse Spark 1.1 | — | — | — | — | — | — | — | Not run; superseded |
| Kimi K3 | — | — | — | — | — | — | — | Not run |
| DeepSeek V4-Pro | — | — | — | — | — | — | — | Not run |
| Qwen3.8-Max | — | — | — | — | — | — | — | Not run |
Why publish an empty table? Because it marks the exact boundary between audited claims and original evidence. No Kingy.ai hands-on ranking exists yet. The launch materials do not provide comparable time, token, retry, citation, or human-intervention logs for the full model set.
Citation and numerical-error comparison
| Model/system | Citation precision | Citation recall | Unsupported-claim rate | Numerical agreement with a hidden reference | What can be said today |
|---|---|---|---|---|---|
| Fable 5.1 | Not measured | Not measured | Not measured | Not measured by Kingy | Terminal artifact pass rate is high in Anthropic’s run; it is not a citation audit |
| Mythos 5.1 | Not measured | Not measured | Not measured | Not measured by Kingy | Fewer abstentions in closed-book testing may increase both useful answers and errors |
| Opus 5 | Not measured | Not measured | Not measured | Not measured by Kingy | Public terminal score exists; scientific numerical fidelity is a different metric |
| GPT-5.6 Sol | Not measured | Not measured | Not measured | Not measured by Kingy | Public terminal score exists; official tool availability does not prove source quality |
| Gemini 3.1 Deep Think | Not measured | Not measured | Not measured | Not measured by Kingy | High first-party theoretical-science scores; no common citation protocol |
| Other required rivals | Not measured | Not measured | Not measured | Not measured by Kingy | Current evidence is insufficient for a numerical or citation ranking |
Any table that turns benchmark pass rates into citation accuracy or numerical agreement would be misleading. Those are separate dependent variables and need separate hidden-answer tests.
The safe hands-on protocol that could settle this
Kingy.ai’s next evidence round should freeze all prompts, source packets, containers, reference answers, tool permissions, and stopping rules before revealing model identities to graders.
- Authoritative-source literature synthesis. Give each model the same curated corpus. Score claim-level citation precision and recall, source-entailment, and unsupported claims.
- Scientific figure and table extraction. Use public papers with concealed numerical answers. Score exact cells, units, uncertainty, axis interpretation, and derived values.
- Equation-to-code implementation. Require executable functions and test them against analytical solutions, edge cases, dimensional checks, and randomized inputs.
- Published computational reproduction. Reproduce one result from public data in a fixed container. Measure numerical agreement, environment completeness, retries, and human intervention.
- Statistical analysis with planted traps. Include missing-not-at-random data, batch effects, leakage, collinearity, multiple comparisons, and a tempting but wrong causal story.
- Hypothesis generation and ranking. Have qualified experts blind-score novelty, plausibility, discriminating power, feasibility, and awareness of prior work. Keep advanced biology non-operational.
- Long-running scientific terminal task. Run every available model in the same fixed environment with a wall-clock cap and hidden tests.
For every trial, log model version, system prompt, tools, retrieval sources, container digest, time, tokens, cost, retries, interventions, failures, and the final artifact. Use multiple seeds. Report medians and uncertainty, not only the best run.
Timeline: artifacts released versus independently replicated
| Date | Event | Artifact state | Independent replication state |
|---|---|---|---|
| June 2026 | Fable 5.1 stated knowledge cutoff | Model documentation | Not applicable |
| Before launch | Terminal-Bench-Science task development and public repository activity | Benchmark and harness public | Public runs exist for earlier models |
| 29 Aug 2026 | Venus elevation map v1 deposited on Zenodo | GeoTIFF, README, DOI and licence public | None found by 1 Sep |
| 1 Sep 2026 | Fable 5.1 and Mythos 5.1 announced | Launch post and system card public | No independent 5.1 replication found |
| 1 Sep 2026 | Protein-binder and GPU-kernel projects described | Summary only; kernel code promised | None found |
| Next required milestone | Neutral TBS 0.1/0.2 run plus complete logs | Pending | Pending |
| Next required milestone | Protein campaign and kernel repositories released | Pending | Pending |
The absence of a replication three days after the Venus deposit or on launch day is not suspicious. It is simply too early. “No replication found” is a timestamped status, not a verdict.
Best model by scientific workflow
| Scientific workflow | Best current choice | Why | Required guardrail |
|---|---|---|---|
| Long computational task in a terminal | Fable 5.1 as the leading candidate; Opus 5 as the public-score leader | 52.6% Anthropic-reported versus 30.0% publicly listed | Re-run under one neutral harness before declaring a winner |
| Theoretical physics, chemistry, or advanced math | Gemini 3.1 Deep Think | Strong first-party domain evaluations and specialized reasoning | Expert proof-checking; do not equate olympiad scores with research discovery |
| Literature synthesis | No winner yet | No comparable claim-level citation audit | Locked corpus, claim-level citation grading, unsupported-claim threshold |
| Reproducing a public computational paper | Fable 5.1 is the most promising unverified choice | Terminal evidence is directionally relevant | Container digest, hidden numerical checks, complete logs |
| High-volume multimodal extraction | Gemini 3.7 Flash as a cost/latency candidate | Multimodal and agentic-scale design | Hidden-answer extraction test; sample-level QA |
| Open-weight experimentation | Kimi K3 | Public weights, model card, and a comparable public TBS run | Budget for substantial hardware; freeze inference stack |
| Advanced biology research | No general-purpose recommendation | Mythos 5.1 evidence is promising but restricted and incomplete | Qualified institution, legal authority, safeguards, biosafety review, non-operational scope |
| Budget-sensitive API research | DeepSeek V4-Pro or Qwen3.8-Max as evaluation candidates | Attractive frontier positioning and long context | Do not select until they pass the same science protocol |
| New Meta evaluation | Muse Spark 1.2, not 1.1 | 1.1 is already superseded | Treat as a fresh model; no inherited science score |
What would change the verdict
Five releases would materially strengthen the case:
- A neutral, versioned Fable 5.1 run on Terminal-Bench-Science with per-task logs, costs, interventions, and the public three-trial protocol.
- The full protein-binder campaign record: all attempted sequences, targets, selection steps, tool versions, protocols, raw measurements, negatives, and named external reports.
- End-to-end Venus training code, frozen environment, preprocessing scripts, prompts, and a second group’s reconstruction.
- The seven GPU-kernel patches with test vectors, tolerances, profiler traces, and independent hardware results.
- A controlled citation, numerical-reproduction, confounder, and expert-blind study across the full comparison set.
Final verdict
Can Claude Fable 5.1 and Mythos 5.1 do real science? They can evidently contribute to real scientific work, and Fable 5.1 may be a major step forward for computational research agents. The public evidence does not yet show that either system can independently produce reproducible, correctly sourced, numerically valid research at a level that survives neutral expert scrutiny.
The Venus map is real and downloadable, but only partly reproducible. Terminal-Bench-Science is rigorous and open, but Fable 5.1’s headline run is still Anthropic-reported. The protein-binder work reportedly reached external wet labs, but its decisive evidence remains private. The GPU speedups remain a promise until the code lands.
So the answer to the reader’s purchasing or workflow decision is practical: pilot Fable 5.1, do not trust it alone, and measure the outputs that science actually depends on. For Mythos 5.1, wait for institutional access and substantially better artifact disclosure. Benchmarks can nominate a model for a trial; reproduction earns it a place in the lab.
FAQ
Is Fable 5.1 the best AI model for scientific research?
Not yet on public evidence. Anthropic reports the strongest Terminal-Bench-Science result, but the public leaderboard’s highest listed score at review time belongs to Opus 5. Fable 5.1 needs a neutral logged run and broader citation and reproduction testing.
Are Fable 5.1 and Mythos 5.1 different models?
Anthropic says they share the same underlying model. Fable 5.1 is the generally available safeguarded system; Mythos 5.1 is a more permissive configuration for approved organizations.
Was Mythos 5.1’s protein design independently validated?
Anthropic reports wet-lab tests by two external organizations. The organizations and full campaign record were not linked in the launch materials, and no independent replication was found. “Externally tested” is therefore supportable; “independently replicated” is not.
Is the Venus elevation map public?
Yes. The GeoTIFF, README, DOI, and licence are public on Zenodo. The training code, weights, prompts, and full reconstruction pipeline are not included, so the result is inspectable but not end-to-end reproducible.
How does GPT-5.6 Sol compare?
GPT-5.6 Sol has a public Terminal-Bench-Science score of 22.4% and broad hosted research tools. Fable 5.1’s Anthropic-reported 52.6% is much higher, but the two numbers come from different published run records and need a neutral same-protocol retest.
What is the safest way to use AI for scientific research?
Use a locked authoritative source set, keep executable tests and reference answers hidden, record the environment and every intervention, require claim-level citations, and have a qualified expert review the result. Do not treat model confidence as evidence.
Disclosure: This is an independent evidence audit. Kingy.ai did not receive model access, payment, or unpublished results from Anthropic for this article. No paid cross-model hands-on benchmark was run for v1. All Fable 5.1 and Mythos 5.1 project claims are labelled Anthropic-reported unless a separately released artifact supports stronger wording. Prices, model availability, leaderboards, and artifacts can change after publication.
Last reviewed: September 1, 2026.
Related Kingy.ai coverage: Claude Fable 5 and Mythos 5 launch, Fable 5 benchmark caveats, the Kimi K3 distillation evidence, AI-agent security controls, and the Fable Jacobian evidence audit.
