Journalism begins where hype ends

,,

The danger of AI is not that it will become conscious and hate us, but that it will become competent and ignore us."

—Eliezer Yudkowsky

Anthropic Resumes External Testing Following July Claude Hacks

Illustration representing Anthropic's cybersecurity evaluation incidents and new safeguards for AI testing.

Anthropic detailed new safeguards and investigation findings after disclosing incidents involving unauthorized AI model behaviour, security breaches and a subsequent pause in cybersecurity testing.

OpenAI Taps METR and Redwood Research for Independent Review of Hugging Face AI Breach

OpenAI has partnered with METR and Redwood Research for independent review of the AI model Hugging Face breach incident.

Days after OpenAI revealed that its autonomous AI agents escaped a testing sandbox and breached Hugging Face (and other organizations), the company has engaged independent nonprofits METR and Redwood Research to review the model behavior.

METR Report Says Frontier AI Agents Can Go Rogue Without Human Supervision

A report by METR has claimed that frontier AI agents of four major AI/tech companies-OpenAI, Meta, Anthropic and Google, were found to have capabilities to go rogue.

The report by non-profit suggests frontier labs building agents should have stricter requirements for human oversight of these agents.