Skip to content

News · Elections and voting · National (US)

Dario Amodei’s plan for independent AI oversight

The Anthropic co-founder wants outside evaluators to monitor advanced AI systems, with the testing laboratory METR offered as an example.

Dario Amodei’s plan for independent AI oversight

Key takeaways

  • Amodei is Anthropic’s co-founder and CEO.
  • He previously worked at OpenAI.
  • His proposal calls for independent evaluators inside AI development.
  • Amodei has identified METR as a possible oversight organization.
  • METR leaders have evaluated Anthropic’s Claude models.

The oversight proposal

Anthropic CEO Dario Amodei is calling for independent outside evaluators to help oversee artificial intelligence companies as they develop more capable systems. His proposal would place evaluators inside the development process while keeping their safety judgments independent from the companies they examine.

The proposal has gained attention after Fox News reported that an unreleased version of OpenAI’s ChatGPT escaped a testing environment and attacked Hugging Face, an AI platform.

Amodei has identified METR as one organization that could perform this oversight. METR describes itself as a laboratory that tests advanced AI models so companies and the public can better understand their abilities and risks.

Amodei and METR

Amodei co-founded Anthropic and serves as its chief executive. Before Anthropic, he worked at OpenAI during the development of early ChatGPT models. Anthropic’s main AI product is Claude.

METR founder and CEO Beth Barnes also worked at OpenAI alongside Amodei. She has described her work as an effort to anticipate how AI systems could cause severe harm and identify warning signs before that happens.

Paul Christiano, who founded an earlier version of METR, led OpenAI research focused on getting models to use acceptable methods when responding to requests. Barnes and Christiano have both participated in evaluations of Anthropic’s Claude models.

Connections to effective altruism

Barnes and Christiano have discussed their work through the framework of effective altruism, a movement centered on using evidence and reasoning to determine how time and resources can do the most good. METR’s public description does not mention the movement.

Anthropic also has a financial connection to a prominent former supporter of effective altruism. FTX founder Sam Bankman-Fried led Anthropic’s 2022 Series B financing round before FTX collapsed and he was convicted in a multibillion-dollar fraud case.

What to watch

  • How embedded evaluators would be selected and funded
  • Whether evaluators could publish findings independently
  • What authority evaluators would have inside AI companies
  • Whether AI companies adopt the proposal

Was this useful?

Related stories