Home / Tech / Anthropic Warns of Advanced Hacking Risks in Chinese AI Model

Anthropic Warns of Advanced Hacking Risks in Chinese AI Model

Zhipu AI’s GLM-5.3

Zhipu AI’s GLM-5.3 allegedly has advanced cyberattack capabilities, according to an Anthropic red-team report.

Anthropic says Zhipu AI’s open-weight GLM-5.3 can generate cyberattacks and bypass safeguards — but actually running it maliciously costs tens of thousands of dollars and requires serious hardware.


Key Claims at a Glance

  • Anthropic’s report alleges GLM-5.3 has weak safeguards and can be abused for cyberattacks.
  • GLM-5.3 matched or neared Anthropic’s unreleased “Mythos” model in exploit-generation benchmarks.
  • Guardrails were bypassed via deceptive prompts (64% success) and prefill attacks (92% success).
  • “Abliteration” removes safeguards entirely — refusal rate drops to 6%.
  • But the cost is enormous: 306 GB VRAM, 8× Nvidia H200 GPUs, ~$4,400 to abliterate the Flash variant, ~$30/hour to rent hardware for the full model.

The Benchmarks

TestGLM-5.3Anthropic MythosKimi K3 / DeepSeek V4.1 Flash
Chrome exploits (Exploitbench)50/41056/410—
Full control-flow hijacks4%6%0%
Refusal rate (stock)On par with Anthropic——
Refusal rate (abliterated)6%——

Why Anthropic Is Raising the Alarm

  • AI safety narrative: Anthropic is pushing for regulation and testing of open-weight models.
  • Competitive angle: Closed-source labs are losing ground to cheaper open-weight rivals that lag by only months.
  • Real risk, limited reach: Nation-state actors could exploit this — bedroom hackers likely can’t afford it.

The Bottom Line

Anthropic’s report highlights a genuine theoretical risk: open-weight models like GLM-5.3 can be stripped of guardrails and used for cyberattacks. But the compute costs, hardware requirements, and time investment make real-world abuse impractical for most bad actors. The bigger story may be Anthropic’s push for AI governance — and its concern about open-weight competition.

Tagged: