Anthropic CEO Dario Amodei believes artificial intelligence (AI) companies need the help of independent outside watchdogs to safeguard their developing technologies. But one of the organizations he believes could do the job, METR, is closely connected to the same AI safety community that has ties to Anthropic's start. Many of its leading figures have ties to a movement called "effective altruism (EA)" — the belief that evidence and careful reasoning can help maximize the good that companies and people can do with their time and resources.
METR describes itself as an AI safety testing laboratory. ANTHROPIC CEO LIKENS AI FIGHT WITH CHINA TO COLD WAR, SEEKS 'DISARMAMENT NEGOTIATIONS' "METR evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose," its website reads. Although its website and public materials don’t make mention of effective altruism, its founders have used it as a framing for their work.
Beth Barnes, METR’s founder and CEO, worked at OpenAI alongside Amodei as the company developed early versions of ChatGPT. At an Effective Altruism Global event, she laid out her vision for how to maintain AI safety. "Our overall plan is — it sort of seems like it would be good if it was someone’s job to look at models and decide if they’re going to kill us, think through the ways that that might happen, anticipate them, figure out what the early warnings would be, that sort of thing," Barnes said.
Similarly, Paul Christiano, who led the research around OpenAI’s efforts to ensure its models followed acceptable strategies for delivering requested results, later founded the first iteration of METR. He too has framed parts of his approach to AI safety as a form of effective altruism. "My suspicion is that it is more important for the ‘effective altruism’ movement to have a fundamentally good product and to generally have our act together than for it to grow more rapidly," Christiano said in a 2014 article.
Figures like Barnes and Christiano provide informal links to Anthropic — a company that received funding from effective altruism’s largest supporters when it emerged as a way of thinking among tech moguls. Barnes and Christiano also both worked on evaluations involving Anthropic models, providing safety checks for Anthropic's flagship AI, Claude. Most notably, Sam Bankman-Fried, the founder of the cryptocurrency exchange FTX, led Anthropic's 2022 Series B financing round before his company collapsed and he was convicted in a multibillion-dollar fraud case.
Before FTX's downfall, Bankman-Fried was one of the highest-profile proponents of the Effective Altruism movement, publicly saying it shaped his approach to earning and giving money. ANTHROPIC'S MORAL COMPASS ARCHITECT SUGGESTED AI OVERCORRECTION COULD ADDRESS HISTORICAL INJUSTICES RUTHLESS EXCLUSIVE: PENTAGON CTO REVEALS REASONING FOR ANTHROPIC REMOVAL Similarly, Skype co-founder Jaan Tallinn led Anthropic's 2021 Series A financing round. Tallinn has been one of the Effective Altruism movement's most prominent supporters, speaking at Effective Altruism Global conferences, helping found the Centre for the Study of Existential Risk and the Future of Life Institute, and donating more than $1 million to the Machine Intelligence Research Institute, an organization focused on AI safety and alignment research.
Extract — continue reading at the source.