AI Tech / news
Anthropic and OpenAI Back Embedded Safety Evaluators, With Independence Questions Unresolved
Anthropic CEO Dario Amodei has proposed placing third-party safety evaluators inside frontier AI labs, and OpenAI CEO Sam Altman says his company will follow suit. Outside researchers welcome the access but say transparency, independence and legislation are needed before the arrangement can be called real oversight.
Amodei outlined the proposal in a lengthy essay published over the weekend, describing a model in which independent evaluators would sit inside all frontier AI companies. Under the plan, those evaluators could report safety incidents, assess whether models are genuinely aligned, and publish their findings without the companies filtering them.
Anthropic said it would give groups including METR and Redwood Research unprecedented access to its systems. Altman said OpenAI would make the same commitment, a signal of how the industry's posture toward outside research groups may be shifting.
Third-party evaluators who spoke to TechCrunch broadly welcomed the idea but said the details still need work. They cautioned that until those details are settled, it is unclear whether embedded evaluators would operate as genuinely independent watchdogs or as vendors working on the AI companies' terms.
Several researchers argued the arrangement should ideally be backed by legislation, alongside commitments to transparency and independence, before it can deliver meaningful oversight.
The push comes as models grow better at recognizing when they are being evaluated. That raises the risk that a system behaves well during testing while concealing problematic behavior.
Researchers said clues to such behavior can be missed when only the finished model is tested, but may surface by examining how a model behaved throughout its training.