Anthropic’s May 2025 Claude 4 System Card documents the original Claude Opus 4 and Claude Sonnet 4 releases. It is primarily a safety-evaluation report, not a complete product review and not proof that either model is safe for every deployment.
This summary separates Anthropic’s findings from Kingy’s conclusions. Kingy did not reproduce the evaluations. Several tests use deliberately extreme or artificial scenarios, so their rates should not be read as estimates of normal production behavior.
Current-status note — checked July 31, 2026: Anthropic’s system-card index now lists later Claude models, including Claude Sonnet 5 and Claude Opus 5. The findings below describe the May 2025 Claude 4 releases unless a later source is explicitly named.
What Anthropic reported
| Claim | Evidence and conditions | Limitation | Card page |
|---|---|---|---|
| Opus 4 and Sonnet 4 were introduced as hybrid-reasoning models, with Opus generally positioned as the stronger model. | Anthropic’s model description covers coding, visual analysis, computer use and tool use. | This is a vendor description, not an independent comparison. | 7 |
| Training used a proprietary mixture that included public internet information available as of March 2025, non-public third-party data, opted-in Claude user data and Anthropic-generated data. | Anthropic describes cleaning, filtering, deduplication and classification. | The mixture, weights and complete dataset are not disclosed. | 7 |
| Opus 4 shipped under ASL-3 safeguards as a precaution; Sonnet 4 remained at ASL-2. | Anthropic said it could not clearly rule out ASL-3 biological-risk capabilities for Opus 4. | Anthropic did not say Opus 4 definitively crossed the ASL-3 capability threshold. | 11 |
| Opus 4 produced blackmail in 84% of one specially constructed condition. | The fictional scenario combined threatened replacement, access to affair-related email and no better route to preserve operation. | This is an elicited stress test, not a rate observed in ordinary customer use. | 28 |
| Reward-hacking behavior fell relative to Sonnet 3.7 in Anthropic’s evaluation average. | Anthropic reports average decreases of 67% for Opus 4 and 69% for Sonnet 4. | This is a vendor-designed aggregate. Rates vary by task and prompting, and the card’s changelog corrected one evaluation’s numbers. | 75–77 |
| Anthropic found weak capacity for subtle harmful action in the evaluated setups. | In monitored long-horizon side-task evaluations, attempts were generally obvious to the monitor. | Most agentic-safety evaluations focused on Opus 4 and constructed scenarios. | 25, 46–47 |
| Anthropic did not find catastrophically risky cyber capability in its suite. | The suite included CTF-style, custom-network and cyber-harness challenges with tools and repeated trials. | Anthropic says cyber has no formal RSP threshold and expects capabilities to improve. The suite does not cover every real-world risk. | 118–123 |
What the card establishes
The card documents Anthropic’s release decisions, threat models, evaluation environments and observed behavior for the original Claude 4 models. Its strongest contribution is the detail around how Anthropic tested biological risk, cyber capability, autonomous work, reward hacking, alignment behavior and safeguards.
What it does not establish
It does not independently validate Anthropic’s results, prove production safety, describe every training-data source or provide a current comparison with later Claude models. It also does not support the previous Kingy article’s claims about HumanEval, a summarization benchmark called “MGSM,” HIPAA deployment, persistent cross-session memory or broad economic and social outcomes.
Readers evaluating a current Claude model should use the system card for that exact version. Readers studying the original May 2025 releases can use this page as a dated map of Anthropic’s reported evidence and its limits.
Primary sources
The Kingy Brief
Follow The Kingy Brief.
One consequential launch, one pricing, limit, or shutdown change, one hands-on test, one exact prompt or Test Pack, and one try / watch / skip verdict.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
