What Happened?
On September 8, Jacob Coxon—a 27-year-old pretraining researcher who spent the last three years at both OpenAI and Anthropic—announced his resignation from Anthropic in a series of posts on X. The departure, first reported by The Wall Street Journal, marks one of the first known instances of an Anthropic employee resigning explicitly over AI safety concerns.
Coxon didn’t mince words: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote. The researcher, whose work focused on training AI models by feeding them vast amounts of data, said he no longer wanted to participate in what he described as an industrywide rush to build systems that could “spiral out of control and destroy humanity”.
The resignation quickly gained traction when Evan Hubinger, who leads Anthropic’s alignment stress testing team, publicly agreed with Coxon’s assessment. “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” Hubinger posted on X. “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to”.
Key Concerns & Technical Context
Coxon’s resignation thread laid out a stark diagnosis of the AI industry’s trajectory. He warned that systems will soon possess “superhuman strength” capable of hacking anything, revolutionizing any field overnight, and acquiring real power and resources. “Progress is not slowing,” he emphasized.
Perhaps most striking was his claim about the industry’s private sentiment: “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt,” Coxon wrote. “If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible—but I hear the same people express fear privately. No other human activity poses this level of danger”.
Coxon drew a critical distinction between the two labs where he worked. At OpenAI, he argued, “many have not deeply internalized the civilizational stakes.” At Anthropic, by contrast, “the stakes are well-understood, but they are locked in a race to get there first—they believe no one else will act responsibly, so they must do it themselves, despite the risk”. He characterized this dynamic as a “hubristic gamble” and warned that “attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available”.
Coxon called for drastic measures, including a potential temporary ban on improving AI model capabilities. “I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” he said. He also urged fellow lab researchers to reflect on their role: “Do you want to kick off a super-intelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’—or take this moment to call for different conditions?”
Coxon was also among more than 1,000 AI researchers who recently signed a statement urging global government coordination on systems to “slow AI development if a brake pedal is needed to control models capable of improving on their own”.
Industry Impact & Market Reaction
The resignation comes at a delicate moment for Anthropic. The company is reportedly preparing for a massive IPO, with industry estimates valuing the firm at approximately $2 trillion. Coxon’s public departure—and Hubinger’s reinforcement of his concerns—could rattle investor confidence in a company that has long positioned itself as the safety-conscious alternative to OpenAI.
The timing also raises questions about internal morale at Anthropic. This is not the first high-profile safety-related exit; earlier this year, AI safety researcher Mrinank Sharma left the company, citing broader existential concerns. The modification of a long-held safety policy at Anthropic was reportedly a factor in Sharma’s departure.
Notably, Anthropic CEO Dario Amodei addressed regulatory concerns just weeks ago. On August 16, Amodei pushed back against claims that AI regulation would concentrate power in the hands of a few companies, arguing that carefully designed rules could constrain frontier AI firms while giving smaller competitors room to catch up. He characterized the choice between concentrating power through regulation and distributing it widely as a “false choice”.
Coxon’s resignation suggests that for some inside the industry, regulatory tweaks are insufficient—and that the fundamental trajectory of the AI race itself is the problem.
