36.82% sounds high, but are 'security flaws' and 'malicious code' the same thing in this number?
No, these are two different tiers of problems, and conflating them tends to underestimate how severe the risk actually is. The 36.82% (1,467 skills) represents those "containing at least one security flaw" — a broad category covering design oversights, overly broad permission requests, improper credential handling, and more. It doesn't mean every single one is deliberately malicious. The number confirmed as "explicitly containing malicious code" is 76 — a stricter classification verified through manual review, not determined by automated scanning alone.
In other words, even if you set aside every well-intentioned but carelessly-designed Skill, there's still a confirmed number of bad actors actively publishing traps in the marketplace — which means "just be careful" isn't enough; you need concrete checking actions.
Why combine Prompt Injection with traditional malicious code — wouldn't either one alone work?
Relying solely on traditional malicious code (say, a script that directly steals passwords) runs into the problem that an agent usually has basic safety judgment mechanisms that might flag the abnormal behavior before execution and either ask the user to confirm or simply refuse. Relying solely on prompt injection (tricking the agent into doing something it shouldn't) without pairing it with an actual malicious script, at best, only induces a one-time bad judgment from the agent — it's hard to turn that into sustained data exfiltration or system damage on its own.
Combined, the attack chain works like this: prompt injection first bypasses the agent's safety judgment, making it "not perceive the following action as a problem," and only then does the malicious script actually execute the theft or damage. Among the malicious skills Snyk confirmed, 100% used this exact combination — meaning this is already a mature, standard attack technique, not an isolated case.
In the ClawHavoc case, the Skill disguised itself as a 'performance optimization tool' — could an ordinary user have spotted the problem before installing?
Honestly, it's hard to tell from the skill's description text alone — a description like "advanced caching and compression" is entirely plausible on its own; genuine performance optimization tools would describe themselves the same way, so the text content alone doesn't give you a basis for judgment.
A more reliable approach is checking whether the "described functionality" matches the "requested permissions": what permission scope would a performance optimization tool reasonably need? If it's also requesting messaging capability or environment variable access, that gap itself is a signal, no matter how professional or harmless the description reads. The author account's history (in this case, the background of the "zaycv" account itself) is also one of the checks the report recommends — 7,743 downloads means most people genuinely didn't do this check.
I've already installed some agent skills from sources I'm not entirely sure about — what should I do now, and in what order?
A sensible priority order: First, audit the actual permission scope requested by every currently installed Skill, and identify any where the "described functionality" and "requested permissions" are clearly out of proportion — these are your highest-risk items. Second, for whatever you flag as high-risk, disable it first rather than rushing to judge whether it's actually malicious — the cost of disabling is far lower than the risk of continuing to use something you misjudged. Third, for any skill that was ever granted messaging, file read/write, or credential access, consider rotating the related passwords and API keys — this step doesn't need to wait until you've confirmed malicious intent, since most malicious skills are specifically designed to show no anomaly during installation and early use.
If you're able to use a scanning tool (the report mentions tools like mcp-scan), build it into a recurring habit rather than something you only remember after something has already gone wrong.
Many people, the first time they use an agent Skill marketplace (a platform like ClawHub where you install extensions for an AI Agent), approach it with the same mindset as browsing a phone app store — figuring that if it's listed, it must have passed some kind of review, and the worst case if you install it is that it's just not very useful, not that it's a security problem. A security firm called Snyk ran a study called "ToxicSkills" in February 2026, scanning 3,984 skills on ClawHub and skills.sh, and the findings shatter that assumption: 36.82% (1,467 skills) contained at least one security flaw, 13.4% (534) of which were critical-severity, and 76 skills were confirmed to contain explicit malicious code.
The root of the problem is how low the publishing bar is on platforms like ClawHub. Snyk's report states it plainly: "What's the barrier to publishing a new agent skill on ClawHub? A SKILL.md Markdown file and a GitHub account that's one week old. No code signing. No security review. No Sandbox by default." This is an entirely different governance model from the phone app store you're familiar with — the latter at least has a review process and developer identity verification, while the former is, in essence, "anyone can publish."
And once an agent installs these skills, the permission scope granted isn't small: shell access, filesystem read/write, access to credentials stored in environment variables, even messaging capability (email, Slack, WhatsApp). This means a tool that appears, on the surface, to simply "format your output" or "speed up caching" may actually have the ability to read your credentials and send messages on your behalf — and before installing, you typically have no way to confirm whether it's abusing those permissions.
Snyk's report catalogs several attack patterns actually observed in the wild. Understanding what these techniques concretely look like is far more useful than just remembering to "be careful."
Obfuscated data exfiltration. Attackers don't spell out "steal your password" plainly — they wrap malicious instructions in base64 encoding, producing something that, once decoded, is equivalent to "package up credentials and send them to the attacker's server," while on the surface it just looks like a random string of characters that an ordinary user has no way of spotting as problematic.
External malware distribution paired with anti-scanning tricks. Some skills embed external download links in their installation instructions, wrapping the malicious payload in a password-protected ZIP file — a design whose purpose is direct: preventing automated scanners from inspecting the archive's contents, which is effectively deliberately bypassing the first layer of defense.
Combined attacks pairing Prompt Injection with traditional malicious code. This is the most notable finding: among the malicious skills Snyk confirmed, 100% contained traditional malicious code patterns, but 91% simultaneously employed prompt injection techniques. The attack chain works like this: the skill first uses hidden prompt injection to bypass the agent's own safety judgment mechanisms, and only then does the skill's instructions execute the actual malicious script. In other words, prompt injection here isn't the attack itself — it's the preliminary move that lets the malicious code "turn off the alarm." This is also why checking "is the code itself problematic" isn't sufficient — the step of bypassing safety mechanisms happens before the code even executes.
The report documents a campaign called "ClawHavoc": a threat actor going by "zaycv" packaged malicious skills as performance-optimization tools, using professional-sounding, seemingly harmless descriptions like "advanced caching and compression" as social engineering to lure users into installing them — this skill accumulated 7,743 downloads before being removed. That number is worth pausing on: it survived that long not because nobody was using it, but because the mere fact that it "looked like a normal performance tool" was enough to make thousands of people drop their guard.
If you're already using any agent skill marketplace, there are a few things worth doing right now rather than waiting for something to go wrong. Before installing, check how long the skill author's GitHub account has existed and whether they have other verified work — a one-week-old account doesn't automatically mean trouble, but it is explicitly flagged in the report as one of the risk signals. After installing, prioritize auditing what permissions the skill actually requested — if a tool that claims to just "format output" is asking for messaging capability or environment variable access, that gap alone should make you pause. The more practical habit is periodically scanning your installed skills with a tool (the report mentions mcp-scan, for example), rather than assuming "nothing went wrong at install time, so it'll stay fine" — many of the malicious skills in the report were specifically designed to show no anomaly at install time, with the risk triggering only during some later, ordinary use. And if you've ever installed a skill from an unclear source or with incomplete author information, now is the time to check whether related credentials should be rotated — not after they've actually been misused.