Journalism begins where hype ends

,,

I do not fear computers. I fear the lack of them."

—Isaac Asimov

Anthropic Researcher Quits Warning AI Could Kill Humanity; His Alignment Lead Agrees

Anthropic researcher Jacob Coxon has resigned over concerns that AI companies are racing toward self-improving superintelligence without adequate safeguards. Anthropic alignment lead Evan Hubinger backed Coxon’s warning that AI could kill all humans and said he personally puts the risk above 10% within the next decade.
Illustration of AI warning imagery representing fears over self-improving superintelligence and AI risks to humanity
September 9, 2026 02:18 PM IST | Written by Supriya Singh | Edited by Vaibhav Jha

An AI safety researcher, Jacob Coxon, employed at Anthropic, resigned from his post on Wednesday and stated in a social media post that the companies building AI technology believe that it can “kill humanity” still they are focused on achieving “self-improving AI” and “superintelligence” and thus are gambling with lives.

 

Recursive self-improvement refers to a process where AI models design, code a d train better version of themselves without human oversight. Superintelligence is a condition when AI models and agents surpass humans in every domain. Both are theoretical concepts yet and not proven.

Coxon’s views on AI posing existential threat to humanity has been endorsed by Evan Hubinger, the alignment science head at Anthropic, saying “we really do earnestly believe AI could kill all humans!”

Coxon, a researcher from Cambridge University, who has worked at both OpenAI and Anthropic, took to X and wrote, “ I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.

 

Coxon claimed that soon AI will evolve into “superhumans systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.

“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger,” said Coxon.

His comments about the possibility of AI wiping out humanity was endorsed by Hubinger who is leading alignment efforts at Anthropic.

In a post on X, Hubinger said that the company is not yet on track to solve the challenge of aligning AI systems’ goals with human interests, despite its ongoing efforts.

“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger tweeted.

Meanwhile in a post following his warning, Hubinger cited Anthropic’s latest risk report and said the threat posed by present models is low. “To be clear, as we say in our latest Risk Report. I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” Hubinger tweeted.

In its August 2026 “Risk Report”, Anthropic revised its assessment of the risk from misaligned AI behavior from “very low” to “low,” citing increased uncertainty following recent disclosures related to model behavior in cybersecurity evaluations.

Also Read: Explained: Why is Anthropic Calling for a Global Pause on Frontier AI?

Authors

  • AI FrontPage Reporter Supriya Singh

    Supriya Singh is a Reporter at AI FrontPage covering the AI & Education and AI & Jobs beats. She brings six years of print and digital experience, including three years at The Asian Age, where she reported on higher education, Delhi government, and crime. She is based in Delhi-NCR.

    LinkedIn

  • Vaibhav Jha, editor and co-founder at AI FrontPage

    Vaibhav Jha is an Editor and Co-founder of AI FrontPage. In his decade long career in journalism, Vaibhav has reported for publications including The Indian Express, Hindustan Times, and The New York Times, covering the intersection of technology, policy, and society. Outside work, he’s usually trying to persuade people to watch Anurag Kashyap films.

    LinkedIn