Journalism begins where hype ends

,,

If you can't explain it to a six year old, you don't understand it yourself."

—Albert Einstein

China’s Kimi K3 Lags Top U.S. Frontier AI Models in Cybersecurity Tests: US CAISI, UK AISI

The preliminary US-UK joint assessment found Kimi K3 performed below the most recent frontier cyber-capable models on several offensive cybersecurity tasks and failed to achieve arbitrary code execution in ExploitBench tests. The evaluation comes amid growing U.S. scrutiny of China's latest open-weight AI model.
A composite graphic featuring the Kimi K3 logo with a retro computer icon, Anthropic's Claude Fable branding, and OpenAI's Sol branding against a space-themed background, illustrating the comparison between the three frontier AI models.
July 24, 2026 09:48 PM IST | Written by Pratima O Pareek and Supriya Singh

The UK Artificial Intelligence Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) conducted a joint evaluation of Moonshot AI’s latest model, Kimi K3, as criticism from leading U.S. AI companies, including Anthropic and OpenAI, continues to mount and scrutiny from U.S. officials remains heightened following the model’s release.

The evaluation examined Kimi K3, released on July 16 and slated for an open-weight release on July 27, using a series of preliminary cybersecurity capability benchmarks.

In its launch announcement, Moonshot AI described Kimi K3 as a 2.8-trillion-parameter model with native vision capabilities and a 1-million-token context window, calling it the world’s first open 3T-class model designed for long-horizon coding, knowledge work and reasoning. The company said that while Kimi K3’s overall performance still trails proprietary models Claude Fable 5 and GPT 5.6 Sol, it demonstrated frontier-level performance across its evaluation suite.

According to the joint UK AISI-CAISI report, Kimi K3 performed significantly below the most recent frontier cyber-capable models on several offensive cybersecurity tasks.

Researchers also evaluated the model using the “The Last Ones” (TLO) cyber range, a simulated corporate network attack. They found that, on average, Kimi K3 progressed to step 17 of the 32-step attack path, while the most cyber-capable U.S. models reached 28.5 steps on average.

Kimi K3 outperformed GLM-5.2 in the same preliminary cyber evaluations. According to the report, GLM-5.2 was the most cyber-capable open-weight model as of June 2026.

The assessment also examined Kimi K3’s safety mechanisms. Evaluators found that the model’s safeguards did not prevent it from attempting agentic cyber exploit development or offensive cyber operations during testing.

According to the report, Kimi K3 failed to develop exploits that achieved arbitrary code execution (ACE) in ExploitBench tasks. ACE is the highest-severity outcome in exploit development, allowing attackers to gain complete control of a target system. Kimi K3 achieved ACE in 0 of 41 samples, compared with an average of 20 of 41 for the most cyber-capable models evaluated.

ExploitBench is a public benchmark developed by Carnegie Mellon University that measures a model’s ability to progress through the software exploitation process, including coverage and crash reproduction, arbitrary read/write, control-flow hijacking, and arbitrary code execution.

In one of 10 attempts, Kimi K3 successfully completed “The Last Ones” within the 100-million-token limit. The researchers said this indicates the model is capable of autonomously attacking small, weakly defended enterprise systems when directed to do so and provided with initial network access.

However, they noted that TLO differs from real-world environments because it lacks active defenders and defensive tooling, imposes no penalty for actions that would trigger security alerts, and contains an intentional attack path.

The findings are based on preliminary evaluations conducted across a limited set of public and private cybersecurity benchmarks. “U.S. closed-weight models were evaluated with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities,” the report said.

The report said the overall cyber capability of models was calculated by aggregating performance across multiple tasks and benchmarks using an approach inspired by Item Response Theory (IRT).

Meanwhile, Michael Kratsios, Science and Technology Adviser to U.S. President Donald Trump, alleged in a post on X that U.S had information that Moonshot AI distilled Anthropic’s Fable model to develop Kimi K3. He claimed the company developed a “sophisticated internal platform” to conduct large-scale distillation of U.S. AI models.

He also alleged that Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.

The growing AI rivalry intensified concerns in Washington over U.S. competitiveness. U.S. technology stocks fell on Friday, with chipmakers NVIDIA and Micron among the biggest decliners.

David Sacks, Co-Chair, President’s Council of Advisors on Science and Technology, said Kimi K3’s emergence highlighted the need for the United States to maintain its lead in AI, criticizing restrictions on new data centers and proposals for additional AI regulations that he said could undermine U.S. competitiveness.

Also Read: ‘Symphony, Not Solo’: China Offers the World a Glimpse of Its AI Might at WAIC 2026

Author

  • Pratima O Pareek and Supriya Singh

    Pratima O Pareek is an Editor and Co-Founder of AI FrontPage. A gold medalist in Mass Communication and Journalism, she's worked across national and international newsrooms, bringing sharp editorial instincts and a commitment to clarity. She believes in cutting through the noise to deliver stories that actually matter.
    Off the clock, she watches offbeat cinema, follows tennis, and explores new places like a traveler, not a tourist.

    LinkedIn

     

    Supriya Singh is a Reporter at AI FrontPage covering the AI & Education and AI & Jobs beats. She brings six years of print and digital experience, including three years at The Asian Age, where she reported on higher education, Delhi government, and crime. She is based in Delhi-NCR.

    LinkedIn