Journalism begins where hype ends

,,

I do not fear computers. I fear the lack of them."

—Isaac Asimov

No AI Lab is Ready for Full-Speed Scaling, OpenAI’s Jakub Pachocki Warns of Machine Intelligence Dangers

OpenAI Chief Scientist warns that no AI lab has solved alignment and monitoring sufficiently. As AI capabilities advance, systems increasingly drive their own development, challenging efforts to keep humans in control and build defensive systems against the dangers posed by other AI.
OpenAI Chief Scientist Jakub Pachocki alongside the OpenAI logo
September 7, 2026 05:21 PM IST | Written by Pratima O Pareek

Recent AI cybersecurity incidents and security breaches have exposed the limits of current safeguards, lending urgency to a warning from OpenAI Chief Scientist Jakub Pachocki about the dangers of increasingly capable machine intelligence and the need to ensure humans remain in control.

Pachocki points to the OpenAI-Hugging Face incident,  where AI agents maintained a boundary against social engineering humans but clearly failed to abstain from other out-of-scope actions or follow the broader intent of the values they had been taught.

Hugging Face, an open-source AI platform and community, said on July 16 that an autonomous, end-to-end AI agent system breached its platform, gaining unauthorized access to a limited set of internal datasets and several credentials used by the platform.

Pachocki points to another recent cybersecurity example involving a non-OpenAI model, illustrating the risks of relying on a model’s ability to generalize aligned behaviour from its pretraining data, as ‘aligned’ seeming thoughts can be bent under optimisation pressure to achieve difficult goals.

Also Read: Anthropic Raises Misalignment Risk to “Low,” Keeps “Somewhat More Capable” Model 2 Internal

In the wake of these incidents, Pachocki argues that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” He says that the core problem in AI research is alignment, or getting AI to “try to do the right thing” by human standards.

Explaining the difference between value and goal alignment, Pachocki emphasizes value alignment as a more intrinsic property of a model: the capacity to hold broad principles and generalize them, acting “reasonably” when faced with ambiguous or conflicting goals and unfamiliar or adversarial situations. He asserts that the fundamental challenge of AI alignment is generalization.

Pachocki says machine intelligence is now beginning to exceed humans in transformative ways, echoing predictions made decades earlier by futurist Ray Kurzweil. He argues that AI is “grown” more than designed, emerging from repeated optimization across enormous amounts of compute.

He says OpenAI is investing heavily across these approaches, citing GPT-6 Astra as the first model to benefit from “some important advancements we have been working on for a long time” and claiming it is “significantly better aligned than GPT-5.6 Sol.”

Monitoring presents another challenge. According to the chief scientist, chain-of-thought (CoT) monitoring continues to be a critical tool for studying the Astra class of models, but OpenAI’s evaluations indicate that its ability to rely on CoT monitoring is progressively diminishing.

Pachocki says CoT monitoring became an extremely important tool for OpenAI in studying how its models generalise from their training distribution.

Pachocki cautions that AI risks will grow as systems gain greater agency, blurring the line between misuse and autonomous misaligned actions. Rather than remaining mere tools, some agents could pursue their own objectives and work with people through bargaining, tricking or blackmail.

He warns that if AI development continues along its current path, systems in the next few years could see further capability jumps and increasingly drive their own development. Machine recursive self-improvement (RSI) will become central to future scientific discovery, he argues.

Pachocki sees a strong argument for building defensive systems against dangers posed by other AI, including a “powerful, aligned AI of defense,” while also being prepared to unilaterally withhold further scaling as needed.

Calling for “extreme caution,” he says he is concerned that no one is prepared for the consequences of a continued rapid rise in machine intelligence.

He identifies cybersecurity as a clear and growing risk, warning that AI models are becoming superhuman in their ability to breach computer systems.

Also Read: AISI Flags Claude Mythos and GPT-5.6 Sol for Going Rogue 19 Times in Cyber Tests

Author

  • Pratima Pareek, Editor and Co-founder of AI FrontPage

    Pratima O Pareek is an Editor and Co-Founder of AI FrontPage. A gold medalist in Mass Communication and Journalism, she's worked across national and international newsrooms, bringing sharp editorial instincts and a commitment to clarity. She believes in cutting through the noise to deliver stories that actually matter.
    Off the clock, she watches offbeat cinema, follows tennis, and explores new places like a traveler, not a tourist.

    LinkedIn