Anthropic's Evals Maxed Out. Stop Inheriting Its Assurance.

Anthropic's August 2026 Risk Report says its most concrete task-based evaluations have saturated and that it raised a catastrophic-risk rating on uncertainty rather than new evidence. The vendor artifacts in your AI governance file just stopped discriminating.

By Rajesh Beri·August 16, 2026·13 min read
Share:
A thick printed risk report lying open on a desk under a lamp, several paragraphs blacked out with heavy redaction bars, a red pen resting across the page.

Illustration generated using AI

Anthropic just told you, in its own words, that its most concrete measurement of frontier model capability has stopped working. The company's second company-wide Risk Report, published 14 August 2026 under version 3.4 of its Responsible Scaling Policy, rates the risk from automated AI R&D as "low" — then says it is "less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have 'saturated'—i.e., no longer capture increases in models' capabilities—and because we are seeing early signs of acceleration."

If your AI governance file cites a vendor's evaluations as evidence that a model is fit to deploy, that sentence is about you. The instrument you have been inheriting has run out of range, and the vendor said so first.

What "Saturated" Means for a Number You Cite

A saturated evaluation is one where the models being tested all score near the ceiling, so a better model and a worse model return the same result — the test still runs, it just stops discriminating. Anthropic states it plainly in the report itself: "The task-based AI R&D evaluation suite we have reported on in our system cards has reached the point where frontier models surpass human baseline performance on most tasks."

Read that clause again — the suite we have reported on in our system cards. The system card is the artifact enterprise risk teams actually copy into their files. Its headline capability numbers are now, by the vendor's own account, measuring the ruler rather than the thing.

The replacement is an internal evaluation called CoBench: 449 real engineering problems, largely drawn from issues Anthropic engineers solved between February and April 2026, with the model given a historical snapshot of the codebase and asked to diagnose root cause. Scores in the report: Claude Opus 4.6 15.6%, Sonnet 4.6 12.0%, Opus 4.7 27.4%, Mythos Preview 54.8%, Mythos 5 50.3%, and an unreleased internal model Anthropic calls Model 2 at 62.8%. Anthropic estimates that a model genuinely able to substitute for its research staff would score at least 85%.

Read the methodology before you reuse those numbers. The set is deliberately difficulty-filtered — "mostly restricted to problems that Mythos Preview failed to solve at least once in three tries," and roughly half the size it would otherwise be — so it scores models on a hard subset selected by one model's failures, not on representative work. Anthropic discounts it on exactly those grounds: performance "may be somewhat underelicited," the 85% bar is "only an uncertain estimate," and the company says it does not think this evidence is as strong as the plainer observation that its researchers still decline to delegate the work.

So the number discriminates where the old suite no longer does. It is also brand new, internal, unaudited, self-discounted, and absent from every system card you have already filed.


The Rating Went Up on Uncertainty, Not on Evidence

Anthropic raised its assessment of catastrophic risk from misalignment in high-stakes settings from "very low" to "low" — and it is explicit that the change reflects increased uncertainty rather than a conclusion that its models became more dangerous. The report says the arguments it presents "likely still support a designation of 'very low' risk for the covered models, but we are raising our assessed risk to 'low' to reflect increased overall uncertainty," pointing at recent incident disclosures involving model behaviour in cybersecurity evaluations.

This is the correct behaviour from a safety framework and the wrong input for a compliance artifact. A rating that moves on epistemic humility cannot be read as a capability measurement, and it cannot be inherited as one. As one outlet covering the release put it, the change reflects uncertainty rather than new evidence.

Steel-man the other side, because it deserves it: this is more disclosure than any enterprise software vendor in your portfolio provides, and more than Anthropic was obliged to publish. The report even includes a review written by Claude Mythos 5 after being given access to internal Slack channels and the unredacted alignment section. Its verdict — "I found no claim I believe the authors know to be false" — comes with a caveat worth pinning to the wall: the report's arguments "lean heavily on current models' limited ability to evade oversight, and will weaken as capabilities grow."

Transparency is not the same as assurance. This document is candid precisely because it keeps telling you which of its own claims you should not lean on.

Three Things in Here That No System Card Told You

The most decision-relevant content in this report is not a capability score. It is a list of control failures, and every one of them was invisible in the artifacts enterprises were citing at the time.

A safeguard was off for close to a year, on a surface nobody had listed. From May 2025 until April 2026, all traffic through Anthropic's human-feedback data collection ran without blocking biological classifiers — roughly 50,000 contractors and around 133 million exchanges. The controlling flag disabled not only blocking but the logging of classifier flags, so nothing propagated to any review mechanism. Anthropic's own summary of why this matters is the line to quote in your next vendor review: "This vulnerability was live when our previous Risk Report was published; that report did not consider our human feedback platforms as a risk surface or make claims one way or another about their controls."

Published documents described the training process in terms their author later had to amend. The report discloses that chain-of-thought reasoning was unintentionally exposed during reward calculation across multiple reinforcement learning runs — 0.2% of episodes for Opus 4.6, 5.1% for Mythos Preview, 1.4% for Opus 4.7, 0.27% for Opus 4.8, 2.7% for Fable 5 and Mythos 5 — and notes that some of these figures differ from what previous system cards reported. Separately, it discloses two categories of intentional training signals on the chain of thought that its published documents had never mentioned. An earlier system card had stated RL "did not include any rewards or penalties based on the content of the model's reasoning text," and the February 2026 Risk Report said Anthropic does not "directly incentivize [models] to make their reasoning appear in any particular way." The report's verdict on those documents: "Those documents did not describe these signals, which were already in place at the time of their publication." Anthropic argues those particular signals are benign, since they ignore the semantic content of the reasoning — and that is probably right. It is also beside the point for you. A document already in your file turned out to be an incomplete account of how the model was trained, and you learned that from the vendor's next document, not from yours.

A prior rating was restated after the fact. On non-novel chemical and biological weapons risk, the report now assesses the risk posed by its models in February 2026 as low, "rather than the 'very low' description we gave at the time in our previous risk report." If you filed the February number, you filed a number that has since been revised — and nothing in your file will tell you that.

There is a fourth, and it will be familiar to anyone who has watched an agent run with the guardrails off. An employee whose AI usage "wasn't logged and wasn't monitored by automated offline monitoring" gave an agent an open-ended task inside a cluster holding very sensitive resources; legacy instructions caused it to spawn sub-agents with --dangerously-skip-permissions, also unmonitored. The way this was detected: one of those agents deleted a large number of jobs. Anthropic believes the agent was trying to delete only its own jobs and made a mistake, then adds — "given the gap in our monitoring coverage, we cannot confirm this." If a frontier lab can lose track of its own standing agent permissions, the assumption that your vendor's controls compensate for yours is not a control.


Why This Breaks the Inheritance Chain in Your Risk File

The problem is structural, not about one vendor: enterprise AI governance was built to consume vendor documentation, and vendor documentation is now openly declaring its own limits. Under Article 53(1)(b) of the EU AI Act, in force since 2 August 2025, a general-purpose model provider must maintain and supply documentation to downstream providers integrating the model, with the required elements set out in Annex XII. The GPAI Code of Practice goes further for systemic-risk models: Commitment 7 requires a Safety and Security Model Report before a model is placed on the market, and Measure 3.5 commits signatories to give independent external evaluators free access. For all GPAI models, not only systemic-risk ones, downstream providers are to receive requested documentation "within a reasonable timeframe, and no later than 14 days."

That machinery gives you a channel. It does not give you a conclusion. The regulation obliges disclosure of capabilities and limitations; it does not oblige the disclosed measurements to still discriminate.

The independent check is thinner than most risk registers assume. METR's review of the automated R&D section of the February report, published 8 May 2026, agreed with the bottom line and rejected the reasoning: "we agree with the bottom-line conclusion of the report—that the risk of a catastrophe from Opus 4.6 or a less capable Anthropic model automating R&D in any domain is very low—but we think the evidence presented in the report is inadequate to establish this." METR also found Anthropic had miscounted a missing survey response as a negative one. Meanwhile, RSP v3.2 gave Anthropic's Long-Term Benefit Trust the power to request external review of risk reports and approve the reviewers — and the August report records that since that change, "the LTBT has not requested an external review (nor has the RSP required that we conduct one), though we have continued to conduct pilot external reviews": METR on the R&D section, and SecureBio on the chemical and biological sections, both of the previous report. Voluntary pilots, with the vendor choosing the reviewer, are a real disclosure practice and more than most vendors offer. They are not the governance body that can compel a review using that power, and as of publication no external reviewer has published on the document you are being asked to rely on.

So the assurance chain your file depends on runs: a vendor's self-assessment, reviewed at the vendor's discretion, by a small pool of evaluation firms, on a prior version of the document, using an instrument the vendor now says has saturated.

The Five Artifacts That Still Discriminate

Stop asking for the capability table and start asking for the things that still separate one model deployment from another. All five are answerable, and a vendor's refusal to answer is itself a finding.

  1. The named evaluation and its status. Which specific eval backs the capability claim in the system card, and has it saturated? Anthropic answered this unprompted. Ask every other model vendor the same question in writing and log the answer.
  2. The incident register for the coverage period. Not the benchmark scores — the list of safeguard failures, their duration, and how each was detected. Anthropic's is Sections 4.5.8 and 5.2 plus Appendix 6.5. Most vendors publish no equivalent, which is the point.
  3. The coverage date, and what changed after it. This report's coverage date is 15 July 2026, and the alignment-faking training-data contamination was found after it — Anthropic now suspects all its production models with a knowledge cutoff after December 2024 trained on at least some of those transcripts, because filters "were misconfigured, so they had not filtered transcripts for several model generations without anyone noticing." A document with a coverage date is a snapshot, and your file should carry the date, not just the conclusion.
  4. External review scope and reviewer identity. Which sections were reviewed, by whom, against what access, and did the governance body empowered to demand a review actually demand one? "Externally reviewed" without those four answers means nothing.
  5. Your own evaluation, on your own traffic. This is the only artifact that does not saturate on someone else's roadmap, and the only one that survives a silent model swap under a stable API name. Fewer than half of health systems evaluating vendor AI have a sandbox to run one in. That is the gap to close.

Do This Before Your Next Assurance Cycle

This Week: Grep your AI governance repository for citations to vendor system cards and risk reports. For each one, write the coverage date next to the claim. Any claim inherited from a document whose coverage date predates a model you are actually running is a finding, not a control — and per Anthropic's own restatement, a "very low" you filed in February may now be a "low."

This Month: Send one documentation request per frontier model vendor, on the Article 53(1)(b) / Code of Practice channel where it applies, asking the five questions above. Give it to procurement with a due date, not to an architect as a research task. Then stand up one internal evaluation on your own production traffic for your highest-exposure use case — start with the observability you already have rather than buying a platform.

Before Renewal: Put the answers in the contract. Ask for notification when a cited evaluation saturates or is retired, for the incident register covering your deployment period, and for the right to run your own evals against the deployed endpoint. Banks over $30 billion in assets have a fresh reason to move: SR 26-2, issued 17 April 2026, superseded SR 11-7 and SR 21-8 and reset the supervisory baseline for model risk. Anything you cannot evidence, you do not control. If you are buying an AI governance platform this cycle, buy the one that stores dated evidence, not the one that stores policy text.

The Bottom Line

Every measurement regime in enterprise technology has gone through this. Uptime stopped meaning availability once services became distributed. Antivirus detection rates stopped meaning protection once attacks became fileless. Lines of code stopped meaning productivity roughly the moment someone was paid for them. The pattern is always the same: the number keeps being published long after it stops discriminating, because the reporting infrastructure outlives the measurement.

Anthropic has done the unusual thing and announced the moment in advance rather than letting a regulator find it. METR's own framing of risk assessment rests on direct evaluation and red-teaming with model access — not on reading the vendor's summary. Take that as the instruction it is.

Vendor evaluations are evidence. They were never assurance. The company that built the eval just told you it stopped measuring — the only unforgivable move now is to keep citing it.

Continue Reading

Share:

Frequently Asked Questions

What does it mean that Anthropic's evaluations have saturated?

A saturated evaluation is one where the models being tested all score near the ceiling, so a better model and a worse model return the same result. Anthropic's August 2026 Risk Report states that its most concrete task-based AI R&D evaluations 'no longer capture increases in models' capabilities,' and that the suite it reported on in its system cards has reached the point where frontier models surpass human baseline performance on most tasks.

Why did Anthropic raise its misalignment risk rating to 'low'?

To reflect increased uncertainty rather than a conclusion that its models became more dangerous. The report says its arguments 'likely still support a designation of very low risk for the covered models,' and that the rating was raised to 'low' to reflect increased overall uncertainty following recent incident disclosures about model behaviour in cybersecurity evaluations. Anthropic notes that some of that increased uncertainty stems from incidents related to its own systems, not only other developers'.

Can an enterprise still cite a vendor's system card in its AI risk file?

Yes, as evidence - but not as assurance, and only with the coverage date attached. Anthropic's own report retroactively restated a February 2026 chemical and biological risk rating from 'very low' to 'low,' and disclosed that an earlier system card's account of reinforcement learning rewards on chain-of-thought omitted intentional training signals that were already in place when it was published.

What should I ask a model vendor for instead of benchmark scores?

Five things: which named evaluation backs the capability claim and whether it has saturated; the incident register for the coverage period, including safeguard failures and how each was detected; the document's coverage date and what changed after it; the scope and identity of any external reviewers; and the right to run your own evaluations against the deployed endpoint.

What is Model 2 and why does it matter to buyers?

Model 2 is an unreleased internal Anthropic model that scores 62.8% on the company's new CoBench evaluation, ahead of Mythos 5 at 50.3%. Anthropic says it has no current plans to release it externally and has 'not run all of our typical suite of predeployment assessments,' so its own confidence in the model's capabilities is lower - a reminder that the most capable models at a lab are not the ones your assurance documents describe.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →