An AI safety researcher, Jacob Coxon, employed at Anthropic, resigned from his post on Wednesday and stated in a social media post that the companies building AI technology believe that it can “kill humanity” still they are focused on achieving “self-improving AI” and “superintelligence” and thus are gambling with lives.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
— Jacob Coxon (@hilbertspaess) September 9, 2026
Recursive self-improvement refers to a process where AI models design, code a d train better version of themselves without human oversight. Superintelligence is a condition when AI models and agents surpass humans in every domain. Both are theoretical concepts yet and not proven.
Coxon’s views on AI posing existential threat to humanity has been endorsed by Evan Hubinger, the alignment science head at Anthropic, saying “we really do earnestly believe AI could kill all humans!”
Coxon, a researcher from Cambridge University, who has worked at both OpenAI and Anthropic, took to X and wrote, “ I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. https://t.co/QAIHiFP3QZ
— Evan Hubinger (@EvanHub) September 9, 2026
Coxon claimed that soon AI will evolve into “superhumans systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger,” said Coxon.
His comments about the possibility of AI wiping out humanity was endorsed by Hubinger who is leading alignment efforts at Anthropic.
In a post on X, Hubinger said that the company is not yet on track to solve the challenge of aligning AI systems’ goals with human interests, despite its ongoing efforts.
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger tweeted.
Meanwhile in a post following his warning, Hubinger cited Anthropic’s latest risk report and said the threat posed by present models is low. “To be clear, as we say in our latest Risk Report. I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” Hubinger tweeted.
In its August 2026 “Risk Report”, Anthropic revised its assessment of the risk from misaligned AI behavior from “very low” to “low,” citing increased uncertainty following recent disclosures related to model behavior in cybersecurity evaluations.
Also Read: Explained: Why is Anthropic Calling for a Global Pause on Frontier AI?






