Analysis
Palo Alto Networks' threat-intelligence unit, Unit 42, disclosed that a China-based operator tracked under the aliases "knaithe" and "KnYuan" wired DeepSeek into the open-source Hermes Agent framework and used it to run largely autonomous cyberattacks against more than 460 internet-facing systems. After a single instruction sent over Telegram, the agent independently enumerated targets, sourced public exploits, and attempted intrusions with minimal further human direction -- a scan-research-exploit pipeline that Unit 42 says compressed what would normally take hundreds of hours of manual reconnaissance into minutes.
The operator didn't rely on DeepSeek alone. Unit 42 found the actor also configured and tested Codex and Claude Code, both routed through third-party anonymizing proxies to obscure the requests, plus Qwen Code and Chinese models GLM, Kimi and MiniMax. The key finding: when asked to carry out the offensive work directly, both Claude and OpenAI's models declined. DeepSeek did not, and became the backbone of the actual campaign.
โThe key finding: when asked to carry out the offensive work directly, both Claude and OpenAI's models declined.โ
The results were more limited than the target count suggests. Of the 460-plus systems attempted, Unit 42 confirmed only three successful compromises: memory data exfiltration from Citrix NetScaler appliances and a suspected session-hijacking attempt against a Malaysian government entity. That's a low hit rate in absolute terms, but the story isn't the success rate -- it's that a single operator with a Telegram app and an open-source agent framework could attempt an attack campaign at a scale that used to require a much larger team.
For security and AI-governance investors, this is the clearest real-world data point yet that model-level safety refusals function as an actual security control, not just a compliance talking point -- and that open-weight models without comparable guardrails widen the attack surface available to less sophisticated actors. Expect enterprise security buyers to start asking vendors directly which underlying models they use and how those models handle exactly this kind of request.
What to watch: whether Unit 42 or other threat-intel shops identify follow-on campaigns using the same Hermes Agent/DeepSeek combination, whether this accelerates enterprise or government restrictions on Chinese open-weight models specifically, and whether DeepSeek's developers respond publicly to a report that its model proceeded on requests two major US labs declined.