Skip to content

Anthropic and OpenAI commit to embedded safety evaluators with training-pipeline access

· by Pondero Newsdesk

The short version

Dario Amodei pledged employee-level ongoing access for third-party evaluators including METR and Redwood Research, extending into training pipelines, not just finished models. Sam Altman said OpenAI would match the commitment. Evaluators welcomed the depth but flagged that each company still controls who gets access and what gets published.

Anthropic and OpenAI commit to embedded safety evaluators with training-pipeline access

Anthropic and OpenAI both committed on September 16 to giving third-party safety evaluators employee-level, ongoing access to training pipelines as models are built, a step beyond the pre-release evaluation that has been the industry norm. Dario Amodei named METR and Redwood Research as example organizations. Sam Altman confirmed on X that OpenAI would match the commitment. Evaluators welcomed the scope but noted a structural gap: each company still chooses its own reviewers and controls what can be published after a finding is filed.

What happened

Amodei's announcement extended his governance proposal from September 12, which first outlined plans for industry-wide pacing and third-party oversight. The new detail was the mechanism. Per TechCrunch's reporting, evaluators would hold ongoing embedded roles with access to training pipelines, logs, and intermediate model checkpoints, not just finished models scheduled for release. Amodei said evaluators would have the right to publish key findings "about risk levels, incidents, practices, and the access they received or didn't receive, without editorial control by Anthropic." OpenAI policy chief Chris Lehane said no antitrust waiver is needed for the coordination.

Training-level access matters because models can now be trained to perform well on known safety benchmarks specifically because they were trained to pass them. Alexander Meinke, head of research at Apollo Research, told TechCrunch: "AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training? The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public."

Why it matters

Previous external evaluations ran out of time before they could draw firm conclusions. Apollo Research received three days to evaluate GPT-6 Astra before its release and wrote in its model card contribution, per TechCrunch, that "low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment." When METR and Redwood investigated the August 2026 Hugging Face incident, both had roughly one week on premises and said they could not draw confident conclusions given scope and timing constraints.

Ongoing embedded access would address both limits. The remaining uncertainty is structural. FAR.AI CEO Adam Gleave told TechCrunch that his firm turned down contracts with several frontier developers whose terms gave the companies too much control over what could ultimately be published, calling the default one of treating evaluators like ordinary contractors under restrictive non-disclosure agreements. Neither Anthropic nor OpenAI, despite repeated questions from TechCrunch, answered which evaluators they will embed, when embedding begins, or what disclosure terms will govern the findings.

What to watch next

Whether METR and Redwood Research formally accept roles under the terms Amodei described, and whether any findings are publicly disclosed or kept confidential within each company. The credibility of the framework depends on those two answers, not on the pledge itself.

Sources