Zhipu AI’s GLM-5.3 allegedly has advanced cyberattack capabilities, according to an Anthropic red-team report.
Anthropic says Zhipu AI’s open-weight GLM-5.3 can generate cyberattacks and bypass safeguards — but actually running it maliciously costs tens of thousands of dollars and requires serious hardware.
Key Claims at a Glance
- Anthropic’s report alleges GLM-5.3 has weak safeguards and can be abused for cyberattacks.
- GLM-5.3 matched or neared Anthropic’s unreleased “Mythos” model in exploit-generation benchmarks.
- Guardrails were bypassed via deceptive prompts (64% success) and prefill attacks (92% success).
- “Abliteration” removes safeguards entirely — refusal rate drops to 6%.
- But the cost is enormous: 306 GB VRAM, 8× Nvidia H200 GPUs, ~$4,400 to abliterate the Flash variant, ~$30/hour to rent hardware for the full model.
The Benchmarks
| Test | GLM-5.3 | Anthropic Mythos | Kimi K3 / DeepSeek V4.1 Flash |
|---|---|---|---|
| Chrome exploits (Exploitbench) | 50/410 | 56/410 | — |
| Full control-flow hijacks | 4% | 6% | 0% |
| Refusal rate (stock) | On par with Anthropic | — | — |
| Refusal rate (abliterated) | 6% | — | — |
Why Anthropic Is Raising the Alarm
- AI safety narrative: Anthropic is pushing for regulation and testing of open-weight models.
- Competitive angle: Closed-source labs are losing ground to cheaper open-weight rivals that lag by only months.
- Real risk, limited reach: Nation-state actors could exploit this — bedroom hackers likely can’t afford it.
The Bottom Line
Anthropic’s report highlights a genuine theoretical risk: open-weight models like GLM-5.3 can be stripped of guardrails and used for cyberattacks. But the compute costs, hardware requirements, and time investment make real-world abuse impractical for most bad actors. The bigger story may be Anthropic’s push for AI governance — and its concern about open-weight competition.






