Skip to content
NewsResearch

Anthropic raises its misalignment risk level and discloses an unreleased internal model stronger than its public frontier

· by Pondero Newsdesk

The short version

Anthropic's second risk report, published August 14, elevated its misalignment risk rating from very low to low and acknowledged that an unreleased internal model outscores Mythos 5 on a new benchmark that replaced a safety evaluation suite the company says can no longer track capability gains.

Anthropic raises its misalignment risk level and discloses an unreleased internal model stronger than its public frontier

Anthropic's second risk report, published August 14, 2026, carried two findings that reinforce each other: the company raised its formal misalignment risk rating from "very low" to "low," and it disclosed that an internal model not yet cleared for release outscores its current public frontier on the benchmark it built to replace an evaluation suite that saturated.

What the report found

The 186-page document covers activity through July 15, 2026, per Anthropic's risk report page. Three substantive findings sit alongside the risk-label change.

First, Anthropic disclosed that an unreleased model it calls Model 2 scored 62.8% on CoBench, its new internal evaluation of 449 real engineering problems. Mythos 5, its current publicly deployed frontier model, scored 50.3% on the same benchmark. Model 2 has not completed Anthropic's full predeployment assessment suite. The company stated it has no current plans to release Model 2 externally, per TECHi's analysis.

Second, Anthropic disclosed that 133 million human-feedback exchanges with approximately 50,000 contractors ran without its blocking biological classifiers between May 2025 and April 2026. The company said it remediated the gap and found no evidence of concerning misuse, per Unite.AI.

Third, the report rated automated AI R&D risk as "low" and then immediately added that it is "less confident in this assessment than we were in prior risk reports," citing two reasons: benchmark saturation and early signs of acceleration.

The benchmark that saturated

The hedged confidence on AI R&D risk traces to a structural problem Anthropic described alongside it. The task-based evaluation suite it had used to monitor whether models were approaching dangerous capability thresholds has saturated. Frontier models now exceed human baseline performance on most tasks in the old suite, making it unable to distinguish capability gains between model generations.

CoBench, the 449-question replacement, retains meaningful headroom at current performance levels. Anthropic estimated that genuine research-staff substitution would require 85% performance on CoBench. Model 2 sits at 62.8%. Mythos 5 sits at 50.3%. The old instruments that were supposed to catch a dangerous threshold can no longer measure incremental progress toward it. CoBench now carries that function, but two generations of models have arrived while the old suite was already blind.

Why it matters

The misalignment risk change from "very low" to "low" was framed as an uncertainty adjustment rather than a signal that models became more dangerous. Anthropic's own language, that its prior arguments "likely still support" the original lower label, indicates eroded confidence rather than a reversal of evidence.

What compounds the story is the timing. The same week Anthropic published its findings, OpenAI announced it was holding its largest planned frontier reinforcement-learning run, citing preliminary evidence that its Astra model may have crossed the "Critical cybersecurity capability" threshold under its own Preparedness Framework, per SiliconAngle. OpenAI's Preparedness Framework is being rewritten, and the team responsible for it was dissolved in late July.

Two labs. One week. Anthropic's instruments for detecting dangerous thresholds saturated before those thresholds were reached. OpenAI's instruments triggered. The safety governance infrastructure both companies built around responsible scaling is under simultaneous stress from opposite directions.

For AI-tool operators and enterprises building on these APIs: system cards and vendor risk assessments that cite the old evaluation suites are now measuring a benchmark that frontier models have already outgrown. Demand current CoBench-equivalent scores or equivalent replacement metrics when assessing deployment risk.

What to watch next

Anthropic said Model 2 is in a staged internal rollout. Whether its full predeployment assessment clears is the next concrete data point. If it does, the 62.8% vs. 50.3% CoBench gap becomes the headline of whatever release follows. If it does not clear, that finding will say more about where the safety thresholds actually sit than the risk report's language does.

Sources