Leading American AI companies are “racing straight to self-improving superintelligence and gambling with our lives,” warns AI researcher Jacob Coxon, who has resigned from Anthropic over safety concerns that he publicized on Sept. 8.
In his X posts, Coxon compared working at OpenAI and Anthropic, noting that staffers at the ChatGPT-maker have not “deeply internalized the civilizational stakes.” He said the stakes were “well-understood” at Anthropic, but the company was “locked in a race to get there first” as they believe “no one else will act responsibly, so they must do it themselves, despite the risk.”
As a result, says Coxon, AI could “kill us all by the end of the decade,” and this is not a “marketing stunt.”
Evan Hubinger, Alignment Science Lead at Anthropic, responded to Coxon on X, agreeing that he and his colleagues “do earnestly believe AI could kill all humans.” He estimated the odds as greater than 10% over the next decade. Hubinger added that “we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
On Monday, Sept. 7, Volker Türk, the UN High Commissioner on Human Rights, speaking at the Human Rights Council in Geneva, raised the prospect of AI systems becoming so powerful that their developers could no longer control them.