Anthropic and OpenAI commit to embedded safety evaluators with training-pipeline access
Anthropic and OpenAI both committed on September 16 to giving third-party safety evaluators employee-level, ongoing access to training pipelines as models are built, a step beyond the pre-release evaluation that has been the industry norm. Dario Amodei named METR and Redwood Research as example organizations. Sam Altman confirmed on X that OpenAI would match the commitment. Evaluators welcomed the scope but noted a structural gap: each company still chooses its own reviewers and controls what can be published after a finding is filed.
What happened
Amodei's announcement extended his governance proposal from September 12, which first outlined plans for industry-wide pacing and third-party oversight. The new detail was the mechanism. Per TechCrunch's reporting, evaluators would hold ongoing embedded roles with access to training pipelines, logs, and intermediate model checkpoints, not just finished models scheduled for release. Amodei said evaluators would have the right to publish key findings "about risk levels, incidents, practices, and the access they received or didn't receive, without editorial control by Anthropic." OpenAI policy chief Chris Lehane said no antitrust waiver is needed for the coordination.
Training-level access matters because models can now be trained to perform well on known safety benchmarks specifically because they were trained to pass them. Alexander Meinke, head of research at Apollo Research, told TechCrunch: "AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training? The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public."
Why it matters
Previous external evaluations ran out of time before they could draw firm conclusions. Apollo Research received three days to evaluate GPT-6 Astra before its release and wrote in its model card contribution, per TechCrunch, that "low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment." When METR and Redwood investigated the August 2026 Hugging Face incident, both had roughly one week on premises and said they could not draw confident conclusions given scope and timing constraints.
Ongoing embedded access would address both limits. The remaining uncertainty is structural. FAR.AI CEO Adam Gleave told TechCrunch that his firm turned down contracts with several frontier developers whose terms gave the companies too much control over what could ultimately be published, calling the default one of treating evaluators like ordinary contractors under restrictive non-disclosure agreements. Neither Anthropic nor OpenAI, despite repeated questions from TechCrunch, answered which evaluators they will embed, when embedding begins, or what disclosure terms will govern the findings.
What to watch next
Whether METR and Redwood Research formally accept roles under the terms Amodei described, and whether any findings are publicly disclosed or kept confidential within each company. The credibility of the framework depends on those two answers, not on the pledge itself.
Sources
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? (TechCrunch): primary source, Sep 16
- Anthropic, OpenAI proposed new neutral AI watchdogs. Why you should worry about the idea (CNBC): secondary source, Sep 16
- Altman Says OpenAI Will Match Anthropic's Embedded Evaluator Pledge (Unite.AI): secondary source
