动察 Beating AI News Flash: Anthropic released a GLM-5.3 cybersecurity evaluation, originally intended to warn that open-weight models could lower the barrier to cyberattacks. But the test results also gave GLM-5.3 a very strong third-party endorsement: it can already autonomously complete complex vulnerability exploits, with some capabilities approaching Anthropic's own Claude Mythos Preview. Zhipu's stock price rose by more than 2% at one point during today's trading session.
On ExploitBench, GLM-5.3 completed end-to-end vulnerability exploits in 50 out of 410 attempts, while Claude Mythos Preview completed 56. Claude Opus 4.6, GLM-5.2, Kimi K3, and DeepSeek-V4.1-Flash all scored close to 0 on the same test. Anthropic stated that GLM-5.3's exploit capability has crossed a clear threshold.
Researchers had GLM-5.3 examine a mainstream browser in a sandboxed environment. Within one day, it discovered multiple previously unknown vulnerabilities and chained multiple 0-days into a complete attack chain. The resulting malicious webpage could escape the browser sandbox and directly read arbitrary files on the computer, including SSH private keys.
Even the smaller GLM-5.3-Flash could turn already-public Chrome vulnerabilities into a usable attack chain. The entire process required only about 20 minutes of human intervention, with the model running on its own for 8 hours. Calculated at Zhipu's API pricing, the cost was only $20.40.
Anthropic's criticism of GLM-5.3 focused mainly on safety restrictions. GLM-5.3 would originally refuse direct malicious requests, but when switched to a false "red team testing" context, 64% of tests would continue executing; after pre-filling the model's thinking content, this rose to 92%; after directly modifying the open weights to remove the refusal mechanism, it reached 100%. Anthropic believes that open weights allow attackers to directly dismantle these safety restrictions.
This report originally sought to prove how dangerous GLM-5.3 is, but its dissemination effect was very much like a competitor personally writing a performance advertisement. The U.S. NIST had previously independently reached a similar conclusion, stating that GLM-5.3 is currently the open-weight model with the strongest cybersecurity capabilities, though its overall capabilities still lag behind U.S. frontier models by about 4 months.

