• en
ON NOW

‘AI Could Kill Us All By The End Of The Decade’, Anthropic Researcher Warns After Resignation

 Jacob Coxon resigns from Anthropic, accusing AI labs of racing toward superintelligence despite fears the technology could kill humanity.

Jacob Coxon, a researcher who has worked at both Anthropic and OpenAI, has resigned from Anthropic, accusing the two artificial intelligence companies of pursuing self-improving superintelligence without adequately addressing the risks it poses to humanity.

Coxon announced his resignation in a series of posts on X on Tuesday, September 9, saying his three years of pretraining research at the two companies had convinced him that neither was acting responsibly. His comments were also reported by several outlets following the publication of his post.  

“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon wrote.

He warned against underestimating the capabilities that increasingly advanced AI systems could acquire, arguing that future systems could surpass humans across a wide range of activities and gain significant power and resources.

“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources,” he said.

Coxon further claimed that concerns about AI causing catastrophic harm are shared privately by people working on the technology, despite more measured public statements from industry executives and researchers.

“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt,” he wrote.

He added that some executives and senior researchers publicly soften their warnings while privately expressing similar fears.

“If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately.”

Coxon also drew a distinction between the motivations he believes are driving OpenAI and Anthropic.

“At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.”

According to Coxon, the competitive race between leading AI laboratories has created a situation in which companies could continue developing increasingly powerful systems even while recognising the potential consequences.

He described entering what he called the AI “endgame” as a dangerous gamble that should not be left to decisions made within private companies.

“Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack.”

Coxon said he remains optimistic that greater coordination between AI companies and governments could prevent an uncontrolled global race, but warned that stronger measures may eventually be required.

“I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.”

He also urged researchers working at AI laboratories to reconsider whether they should continue advancing increasingly powerful systems without a sufficiently rigorous understanding of how those systems operate.

“Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?”

Coxon concluded by challenging AI researchers not to accept the industry’s current trajectory simply because they believe the development of superintelligence is inevitable.

“Should you put your head down because ‘it’s happening anyway’ – or take this moment to call for different conditions?”

Coxon’s warning comes despite growing concern within the AI industry over the potential risks associated with increasingly autonomous and self-improving systems. His comments have also drawn support from current Anthropic personnel, including alignment researcher Evan Hubinger, who publicly said he believed AI could potentially kill all humans within the next decade. 

Faridah Abdulkadiri 

Follow us on:

ON NOW