Why AI researchers keep building something they think will kill humans
At a Goldman Sachs technology conference in San Francisco, investors viewed the AI boom as a major business opportunity. The article says some Silicon Valley researchers see the same prospects as cause for alarm.
Anthropic alignment science lead Evan Hubinger wrote online that he believes AI could kill everyone, estimating the chance at more than 10% over the next decade. The article also reports growing concern that rapid self-improvement in AI may already be underway or close.
Anthropic researchers face a dilemma: pressing ahead could make AI more dangerous, but stopping could leave other companies or Chinese labs to develop it. Some at Anthropic believe they are better placed to make the technology safe.
Samuel Marks, who leads scalable oversight at Anthropic, set out this reasoning in writing, according to the article. It argues that appeals to safety and protecting humanity help attract researchers motivated by public benefit, even as they work for a commercial AI company.
The article frames these warnings as both sincere safety concerns and a way to justify continued work on powerful AI. It does not describe any specific policy change or next step by Anthropic.
Key points
- Investors see AI as a major commercial opportunity, while some researchers fear its risks.
- Anthropic researchers warn of severe dangers but say rivals advancing unchecked could be worse.
- The article says safety-focused messaging helps recruit researchers and justify continued AI development.