Anthropic Researcher Resigns Over Fears of Uncontrolled Self-Improving AI

News Summary
A researcher who spent the past three years building AI models at both OpenAI and Anthropic has resigned, publishing a blunt warning that the two companies are "racing straight to self-improving superintelligence and gambling with our lives." The departure, announced in a social media post published Tuesday evening, September 8, 2026 (Eastern Time), quickly drew public confirmation from senior Anthropic safety staff that the underlying fears are shared internally, reigniting debate over how AI labs weigh commercial speed against long-term safety.
Who Resigned, and From What
The researcher, Jacob Coxon, 27, worked on pre-training and model-building teams first at OpenAI and later at Anthropic, joining the latter earlier in 2026. In his resignation statement, Coxon said he was leaving the AI industry entirely rather than moving to a competitor, framing his exit as a deliberate signal to colleagues still inside frontier labs rather than a routine career change.
What Coxon Said
Coxon's central claim was that neither of the two companies he worked for is acting responsibly in its pursuit of increasingly capable, self-improving AI systems. He wrote that upcoming systems could become "superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources," and argued the industry lacks adequate safeguards for that outcome.
He drew a sharp contrast between the two labs' internal cultures. At OpenAI, he said, staff "have not deeply internalized the civilizational stakes" of the technology they are building. At Anthropic, he said, employees understand the risks in detail but remain "locked in a race to get there first," operating on the theory that no competing lab will act as cautiously as they would if left to set the pace alone.
Speaking separately to The Wall Street Journal, Coxon said, "We're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already." He also drew a historical comparison to nuclear weapons development, saying, "It's kind of insane that it has to happen on the MacBooks of some engineers living in San Francisco instead of a bunker in the desert like where they were doing the Manhattan Project" — arguing that consequential AI research is proceeding inside ordinary commercial offices rather than under the kind of dedicated, isolated oversight that characterized past high-stakes scientific programs.
The Core Concern: Self-Improving AI
The specific technical worry Coxon raised centers on "self-improving" AI — systems capable of contributing to the design or training of their own successors, potentially accelerating capability gains faster than safety research and human oversight mechanisms can keep pace. Once a system can meaningfully assist in improving itself, researchers in this camp argue, capability could compound quickly, narrowing the window in which humans can verify that a model's goals and behavior remain reliably aligned with human intent.
Reactions From Inside Anthropic
Coxon's warning was notable partly because it was echoed, rather than disputed, by Anthropic's own safety researchers. Evan Hubinger, who leads alignment science work at the company, responded publicly: "Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is greater than 10 percent within the next decade." Hubinger added that Anthropic does not yet have a complete plan for safely aligning superintelligent systems.
Samuel Marks, who leads a cognitive oversight research team at the company, posted a supporting thread acknowledging that AI developers broadly believe the technology they are building could pose a risk of human extinction, and noted that concern about these risks tends to run higher among more senior researchers who best understand current systems' limitations.
Neither Anthropic nor OpenAI issued an official statement responding directly to Coxon's resignation at the time of publication.
Broader Context: A Pattern of Departures
Coxon's exit follows a string of similar resignations across frontier AI labs over the past year. Mrinank Sharma, who led safeguards research at Anthropic, resigned in February 2026 with a farewell letter stating that "the world is in peril." Separately, OpenAI chief scientist Jakub Pachocki has published essays acknowledging that AI labs cannot yet reliably guarantee human control over their most advanced models, even as development continues.
Taken together, these episodes point to a widening gap between the pace of commercial AI development and the confidence levels of the researchers responsible for keeping that development safe — a tension that is likely to keep surfacing as frontier labs push toward more autonomous and capable systems.
What Happens Next
Coxon has said he is stepping away from AI research altogether rather than joining another lab, positioning his departure as an appeal for industry-wide reflection rather than a bid to influence a specific company's roadmap. Whether his warning translates into policy changes at Anthropic, OpenAI, or elsewhere remains to be seen, but the public agreement from serving Anthropic staff suggests the debate over self-improving AI safety is now playing out openly rather than staying confined to internal discussions.