AUTOMATA.SALE Системы привлечения заказов
Связаться

NVIDIA's SkillSpector: I Scanned All 41 of My AI Agent Skills, and Here Is What It Actually Found

This is an adapted English version. Original (Russian): automata.sale/blog/ai-automation/skillspector-bezopasnost-ai-skilllov/

NVIDIA's SkillSpector: I Scanned All 41 of My AI Agent Skills, and Here Is What It Actually Found

Skills have become the main way to extend AI agents. In Claude Code, Codex CLI and Gemini CLI, a skill is a directory of instructions and scripts that the agent reads as its operating system. You install one with a single command, and along with it comes someone else’s code with access to your files, repositories, your customer database and your API keys.

NVIDIA treated this as a classic supply chain problem and open-sourced SkillSpector, a security scanner that answers one question: “is this safe to install?”. I installed it and ran it against all 41 of my own skills. Here is what it actually found. Spoiler: nearly everything loud turned out to be false positives, but the scanner did catch one real issue.

Why a skill scanner at all

The baseline numbers from the research NVIDIA references in its README: out of a 42,447-skill dataset, 31,132 were analyzed, and 26.1% contained vulnerabilities while 5.2% showed likely malicious intent. Skills with executable scripts are 2.12x more likely to be vulnerable.

For a business adopting agentic automation, this is a new attack surface that few people think about yet. We learned long ago to check packages in package.json for known CVEs and to avoid piping random scripts from the internet into a shell. But a skill is a hybrid: half documentation for the model (prompt injection, memory poisoning, hidden instructions to ignore refusals), half code (data exfiltration, privilege escalation, supply chain tricks). A regular antivirus sees the second half and is blind to the first.

SkillSpector covers both. It’s a static analyzer with 71 patterns across 17 categories: prompt injection, data exfiltration, privilege escalation, system prompt leakage, memory poisoning, excessive agency, MCP rug-pull behavior (a server changing its tool descriptions after you installed it), YARA malware signatures, and AST-based dangerous code analysis with taint tracking. On top of that, it checks dependencies against known CVEs via OSV.dev.

Know the limits: the scanner never executes the skill. It’s defense in depth, not a sandbox. It won’t catch behavior that only shows up at runtime, and it can’t read images or obfuscated bytecode. But it catches the things no human ever looks at during a casual install.

What a check looks like

Install via uv, one command, nothing polluting your system:

uv tool install git+https://github.com/NVIDIA/skillspector.git

Check a third-party skill before installing:

skillspector scan https://github.com/user/some-skill --no-llm

The scanner outputs a risk score from 0 to 100 with LOW / CAUTION / HIGH / CRITICAL bands and, most importantly, an exit code: 0 means safe to install, 1 means risk above 50, 2 means scan error. That’s a stable contract for CI and install scripts: checking a skill becomes as routine as verifying a package signature. Reports come out as JSON, Markdown or SARIF.

The --no-llm flag keeps all analysis local; the only thing leaving the machine is dependency names for the CVE lookup. The optional LLM mode (semantic “what did the author mean” evaluation) sends file contents to your configured provider, so for client work we simply never enable it.

The experiment: 41 skills under the scanner

My ~/.agents/skills/ directory holds 41 skills, from SEO tooling to CRM automations. Some are handwritten, some grew out of working sessions. A perfect calibration set: I know for a fact there’s no malware in there.

The result: 11 skills scored the maximum 100/CRITICAL with a DO_NOT_INSTALL verdict. Had these been third-party skills, I wouldn’t have installed them. Here’s what triggered the findings, and this is the most instructive part.

A security skill matches against itself. senior-security is a skill that hunts for leaked secrets. Its own regexes for .env files and tokens were, quite correctly, recognized as “credential access”: 20 Privilege Escalation findings. A false positive by definition, because a detection tool contains signatures of what it detects. Same story with code-reviewer, whose reference material includes examples of dangerous code.

A local dev tool reads as SSRF. The impeccable skill with live browser preview does fetch('http://localhost...'), which produced 12 Server-Side Request Forgery findings. For a developer tool that’s normal behavior; it would be alarming in a skill that has no business talking to localhost.

A YARA rule fires on YAML. md-slides tripped the “MCP metadata poisoning” signature on… its YAML frontmatter containing tools: and description: fields. Classic signature analysis: a combination of innocent traits looks suspicious.

An unpinned version is a fair remark. senior-qa got 20 “MCP server referenced without pinned version” findings for npx playwright without a version pin. Formally not an attack, but it is a rug-pull risk: what npx pulls today and what it pulls tomorrow can differ. That’s hygiene worth fixing.

What the scanner caught for real

One finding was not a false positive. Seven skills carried __pycache__ directories with compiled .pyc files. Python creates those automatically when scripts run, and they had leaked into the repositories by accident.

This is exactly what the scanner warns about: standard analysis skips compiled bytecode, which makes it a classic place to hide a payload, since a malicious .pyc next to a clean .py stays invisible to both humans and some tools. I deleted the directories (they regenerate on the next run), and things got cleaner.

The takeaway cuts both ways. On one hand, don’t panic at the sight of CRITICAL: read the evidence (file, line, matched pattern) rather than the verdict. On the other, even on skills I wrote myself the scanner found something worth cleaning. On third-party skills from the internet, NVIDIA’s data puts the density of real problems in the percent range, and checking is far cheaper than dealing with the aftermath.

Making it part of the process

I’ve settled on three rules, which are also the answer to “so what do I do with this”:

Third-party skills go through the scanner. Always. One command before install, exit code as the gate. If the score is above 50, I read the JSON report with the evidence and decide consciously. Takes seconds.

My own skills get periodic audits. Run the scanner over the whole directory monthly: it catches both junk like __pycache__ and drift, such as the day a skill quietly gains a curl call with a token in its arguments. That’s a real exfiltration pattern.

False positives go through a baseline, not through disabling. The scanner has a suppression mechanism for accepted false positives (a baseline file). Suppress specific findings with a comment on why they’re fine. Turning the scanner off entirely is not the way.

If you’re bringing AI agents into business processes, especially with access to a CRM, a customer database or production servers, let’s talk about an architecture that is secure by default: from restricted mandates for agents to checking the entire supply chain. Ping me on Telegram.


📞 +7 (906) 311-77-69 · ✉ hello@automata.sale · 💬 Telegram: @automatasale · 🌐 automata.sale

ИП Урядов Евгений Евгеньевич · ИНН 645112058391 · ОГРНИП 312645301900058 Working across Russia, Belarus and Kazakhstan