← Back to all projects

SkillSpector

Reads an agent skill before you install it and reports what it would be able to do to your machine and your data

OfficialApache-2.0
Stars
15k
Forks
1.3k
Open issues
81
Last commit
26 Aug 2026

What is SkillSpector?

Installing a skill or an MCP server is closer to installing a browser extension than to adding a library: it arrives as instructions and scripts that an agent will follow with your credentials and your files in reach, and almost nobody reads them first. SkillSpector reads them. It scans a repository, archive or directory against a catalogue of patterns covering prompt injection, data being sent somewhere it should not go, privilege escalation, poisoned tool descriptions and supply-chain risk, then returns findings with a score and a recommendation. NVIDIA cites research finding vulnerabilities in 26.1% of skills examined and likely malicious intent in 5.2%. A pattern scanner reports suspicion rather than proof, so expect to review findings and record the ones you have accepted.

What can you do with SkillSpector?

  • Answer one question: is this safe to install — Point it at a repository, a URL, an archive or a directory and it returns findings with a severity, a score out of a hundred and a recommendation.
  • Look for the failures specific to agents — The pattern catalogue covers prompt injection, data exfiltration, excessive agency, memory poisoning and tool descriptions written to mislead — not generic code smells.
  • Check the dependencies too — Known vulnerabilities in what a skill pulls in are looked up against a public advisory database, with an offline fallback when there is no network.
  • Add a second, slower opinion — An optional stage sends the material to a model for a semantic read, which catches intent that a pattern match cannot express.
  • Accept what you have already reviewed — Findings you have judged acceptable can be recorded in a baseline, so a re-scan shows only what is new instead of the same list again.
  • Gate an install from a pipeline — Reports come out as JSON, Markdown or SARIF, so a build can fail on a finding and the result lands in the same place as other security checks.

Before you choose SkillSpector

  • A pattern scanner reports suspicion, not proof. Expect findings on skills that are fine, which is why the tool ships a way to record the ones you have accepted — and a clean report is not a guarantee.
  • The semantic stage that catches intent needs a model, and therefore a key and a bill; the fast static pass runs on its own but sees only what a pattern can describe.

Frequently asked questions

Is SkillSpector free for commercial use?

SkillSpector is released under the Apache-2.0 licence — OSI-approved open source, which permits commercial use.

How can SkillSpector be deployed?

SkillSpector is available as Runs locally / Self-hosted.

Documentation

Reproduced from the NVIDIA/SkillSpector README, published under Apache-2.0. Read the original ↗

SkillSpector

Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks before installing agent skills.

OpenSSF Scorecard

Overview

AI agent skills (used by Claude Code, Codex CLI, Gemini CLI, etc.) execute with implicit trust and minimal vetting. Research shows that 26.1% of skills contain vulnerabilities and 5.2% show likely malicious intent.

SkillSpector helps you answer: “Is this skill safe to install?”

SkillSpector is part of the NVIDIA Verified Skills pipeline, which scans, evaluates, and signs agent skills before publication. Skills that pass are published to the NVIDIA skills catalog.

Documentation

Features

  • Multi-format input: Scan Git repos, URLs, zip files, directories, or single files
  • 71 vulnerability patterns across 17 categories: prompt injection, data exfiltration, privilege escalation, supply chain, excessive agency, output handling, system prompt leakage, memory poisoning, tool misuse, rogue agent, anti-refusal, trigger abuse, dangerous code (AST), taint tracking, YARA signatures, MCP least privilege, and MCP tool poisoning
  • Two-stage analysis: Fast static analysis + optional LLM semantic evaluation
  • Live vulnerability lookups: SC4 queries OSV.dev for real-time CVE data with automatic offline fallback
  • Multiple output formats: Terminal, JSON, Markdown, and SARIF reports
  • Risk scoring: 0-100 score with severity labels and clear recommendations
  • Baseline / false-positive suppression: Accept known findings via a glob-rule or fingerprint baseline so re-scans surface only new issues (docs)

Quick Start

Installation

Open-source software notice: This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

Create and activate a virtual environment first (all make targets assume the venv is active). Use uv or pip; the Makefile uses uv if available, otherwise pip.

Quick install with uv (CLI-only):

uv tool install git+https://github.com/NVIDIA/skillspector.git
# Update later: uv tool update skillspector

If you plan to run skillspector mcp, install the MCP extra at install time:

uv tool install 'skillspector[mcp] @ git+https://github.com/NVIDIA/skillspector.git'

From source:

# Clone the repository
git clone https://github.com/NVIDIA/skillspector.git
cd skillspector

# Create and activate virtual environment
uv venv .venv && source .venv/bin/activate
# or: python3 -m venv .venv && source .venv/bin/activate

# Install for production use
make install

# Or install with development dependencies
make install-dev

Docker (no Python required)

Run SkillSpector without installing Python by building it locally from the included Dockerfile. The image is based on the Docker Official Python 3.12-slim-bookworm image.

Build the image:

make docker-build
# or: docker build -t skillspector .

Scan a local directory by mounting your current directory into /scan, the container’s working directory:

docker run --rm -v "$PWD:/scan" skillspector scan ./my-skill/ --no-llm

Scan with LLM analysis by passing credentials with a local .env file:

cat > .env <<'EOF'
SKILLSPECTOR_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
EOF
docker run --rm \
  -v "$PWD:/scan" \
  --env-file .env \
  skillspector scan ./my-skill/

Or pass credentials directly from your shell environment:

docker run --rm \
  -v "$PWD:/scan" \
  -e SKILLSPECTOR_PROVIDER=anthropic \
  -e ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" \
  skillspector scan ./my-skill/

Write a report to the host filesystem by writing to the mounted directory:

docker run --rm \
  -v "$PWD:/scan" \
  skillspector scan ./my-skill/ --no-llm --format json --output report.json

Optional alias for repeated static scans:

alias skillspector-docker='docker run --rm -v "$PWD:/scan" skillspector'
skillspector-docker scan ./my-skill/ --no-llm

Basic Usage

# Scan a local skill directory
skillspector scan ./my-skill/

# Scan a single SKILL.md file
skillspector scan ./SKILL.md

# Scan a Git repository
skillspector scan https://github.com/user/my-skill

# Scan a zip file
skillspector scan ./my-skill.zip

Size limits

SkillSpector enforces two independent caps on remote and archive inputs to bound the impact of oversized downloads and zip bombs:

  • Per-ingest cap: INGEST_MAX_BYTES (100 MiB) — applied to streamed URL downloads, total uncompressed size of zip archives, and post-clone disk usage of Git repos.
  • Zip member cap: INGEST_MAX_ZIP_MEMBERS (10,000) — caps the number of entries in a single zip.

Note that the per-file 1 MB analysis cap (MAX_FILE_BYTES) is a separate, downstream limit: it bounds what individual analyzers will read out of an already-ingested directory. The ingest caps above bound how much content can land on disk in the first place. A breach of either ingest cap fails closed with an IngestLimitExceededError.

Output Formats

# Terminal output (default) - pretty formatted
skillspector scan ./my-skill/

# JSON output - machine readable
skillspector scan ./my-skill/ --format json --output report.json

# Markdown output - for documentation
skillspector scan ./my-skill/ --format markdown --output report.md

# SARIF output - for CI/CD integration and IDE tooling
skillspector scan ./my-skill/ --format sarif --output report.sarif

Batch Scanning

Scan entire directories of skills in parallel from contrib/batch_scan/:

python -m contrib.batch_scan.batch_scan ./my-skills/ --no-llm
python -m contrib.batch_scan.batch_scan ./my-skills/ --workers 20 -f json -o report.json
python -m contrib.batch_scan.batch_scan ./tests/fixtures/ -f terminal --workers 20

Supports multilingual detection (zh/ja/ko) and terminal/JSON/Markdown output.

For LLM scans with higher concurrency, configure multiple API keys following .env.example — the pool improves throughput and resilience, provided the keys don’t share an account-level rate limit.

See the contrib guide for details.

Note on LLM support: The default configuration targets DeepSeek as the cheapest public option. DeepSeek-Chat is expected to sunset, and the contributor does not have hardware to test against local models. The batch scanner was originally tested with OpenAI-compatible endpoints — DeepSeek’s lack of structured-output support required manual JSON-parsing patches. If you can contribute a more universal backend (Ollama, vLLM, or a different provider), PRs are very welcome.

Suppressing False Positives (baseline)

Suppress known/accepted findings so the risk score reflects only un-triaged issues and re-scans surface only new findings. See the suppression guide for the full reference.

# Accept all current findings into a baseline (run once), then commit it.
skillspector baseline ./my-skill/ -o .skillspector-baseline.yaml

# Scan against the baseline — only NEW findings are reported and scored.
skillspector scan ./my-skill/ --baseline .skillspector-baseline.yaml

# Review what was suppressed (still excluded from the score).
skillspector scan ./my-skill/ --baseline .skillspector-baseline.yaml --show-suppressed

A baseline can also use drift-tolerant glob rules (by rule id, file path, or message) — see .skillspector-baseline.example.yaml. Exact fingerprint baselines are evidence-bound: changing the scanned source or SkillSpector version keeps the finding active until it is reviewed again. When a selected baseline or baseline output is stored inside the skill directory, SkillSpector excludes that exact file from content analysis so its suppression text cannot create findings or enter regenerated fingerprints; sibling files remain in normal scan scope.

LLM Analysis

For the best results, configure an OpenAI-compatible LLM endpoint for semantic analysis. Pick a provider with SKILLSPECTOR_PROVIDER; hosted providers ship bundled default models, while CLI providers fall back to the local runtime’s default model unless SKILLSPECTOR_MODEL is set. SkillSpector also works against local OpenAI-compatible servers (Ollama, vLLM, llama.cpp) and managed inference gateways.

Provider (SKILLSPECTOR_PROVIDER)Credential env varEndpointDefault model
openaiOPENAI_API_KEY (+ optional OPENAI_BASE_URL)api.openai.com (or any OpenAI-compatible URL)gpt-5.4
anthropicANTHROPIC_API_KEYapi.anthropic.comclaude-opus-4-6
anthropic_proxyANTHROPIC_PROXY_API_KEY + ANTHROPIC_PROXY_ENDPOINT_URLAny Vertex-style raw-predict proxyclaude-sonnet-4-6
bedrockAWS_PROFILE (optional) + AWS_REGION — SigV4 via boto3AWS Bedrock Runtimeus.anthropic.claude-sonnet-4-6-20250915-v1:0
nv_buildNVIDIA_INFERENCE_KEYbuild.nvidia.comdeepseek-ai/deepseek-v4-flash
claude_cli(none — uses local CLI auth)local claude binarylocal Claude runtime fallback, or SKILLSPECTOR_MODEL
codex_cli(none — uses local CLI auth)local codex binarylocal Codex runtime fallback, or SKILLSPECTOR_MODEL
# Stock OpenAI
export SKILLSPECTOR_PROVIDER=openai
export OPENAI_API_KEY=sk-...
skillspector scan ./my-skill/

# Anthropic
export SKILLSPECTOR_PROVIDER=anthropic
export ANTHROPIC_API_KEY=sk-ant-...
skillspector scan ./my-skill/

# Anthropic via Vertex-style proxy (corporate gateways, GCP Vertex AI)
export SKILLSPECTOR_PROVIDER=anthropic_proxy
export ANTHROPIC_PROXY_ENDPOINT_URL=https://my-gateway.example.com/models/claude-sonnet-4-6:streamRawPredict
export ANTHROPIC_PROXY_API_KEY=your-bearer-token
export SKILLSPECTOR_MODEL=claude-sonnet-4-6
skillspector scan ./my-skill/

# AWS Bedrock (Claude via SigV4)
export SKILLSPECTOR_PROVIDER=bedrock
# Optional: select an AWS named profile. When unset, the standard
# boto3 credential chain (env vars, instance metadata, SSO, etc.) resolves.
# export AWS_PROFILE=my-profile
export AWS_REGION=us-west-2  # default if unset
# Default model: us.anthropic.claude-sonnet-4-6-20250915-v1:0
# Override with any Bedrock model ID, cross-region inference-profile
# ID, or your own application-inference-profile ARN:
# export SKILLSPECTOR_MODEL=us.anthropic.claude-opus-4-6-20250915-v1:0
skillspector scan ./my-skill/

# NVIDIA build.nvidia.com
export SKILLSPECTOR_PROVIDER=nv_build
export NVIDIA_INFERENCE_KEY=nvapi-...
skillspector scan ./my-skill/

# Local Claude CLI — no API key; uses your existing `claude auth login` session
# Requires: claude CLI installed and authenticated (claude auth login)
export SKILLSPECTOR_PROVIDER=claude_cli
# Uses the local Claude CLI runtime fallback unless SKILLSPECTOR_MODEL is set.
# export SKILLSPECTOR_MODEL=claude-sonnet-4-6
skillspector scan ./my-skill/

# Local Codex CLI — no API key; uses your existing `codex login` session
# Requires: codex CLI installed and authenticated
export SKILLSPECTOR_PROVIDER=codex_cli
skillspector scan ./my-skill/

# Local Ollama or any OpenAI-compatible endpoint
export SKILLSPECTOR_PROVIDER=openai
export OPENAI_API_KEY=ollama
export OPENAI_BASE_URL=http://localhost:11434/v1
export SKILLSPECTOR_MODEL=llama3.1:8b
skillspector scan ./my-skill/

# Override the provider's default model
export SKILLSPECTOR_MODEL=gpt-5.2
skillspector scan ./my-skill/

# Skip LLM analysis (faster, static analysis only)
skillspector scan ./my-skill/ --no-llm

MCP Server

Run SkillSpector as a Model Context Protocol server so any MCP-capable agent (Claude Code, Codex CLI, Gemini CLI) or remote runtime can call scanning as a tool and gate skill/MCP installs on the result — turning SkillSpector into a runtime guardrail instead of an out-of-band audit step.

skillspector mcp requires skillspector[mcp].

# Install, or reinstall if you already used the CLI-only path
uv tool install --force 'skillspector[mcp] @ git+https://github.com/NVIDIA/skillspector.git'

# FastMCP stdio transport for local CLI agents
skillspector mcp

# streamable HTTP/SSE transport for remote / A2A callers
skillspector mcp --transport http --host 127.0.0.1 --port 8000

The stdio transport is the current FastMCP path for local CLI agents, and the initialize hang reported in issue #199 still applies there.

The server exposes a single tool:

  • scan_skill(target, use_llm=true, output_format="json") — scans a Git URL, file URL, .zip, .md file, or directory and returns a structured verdict: risk_score (0-100), severity, recommendation, safe_to_install, and findings. It also reports llm_used / scan_mode so a low score from a static-only scan is never mistaken for a clean full scan.

Register it with Claude Code via:

claude mcp add skillspector -- skillspector mcp

Security — HTTP transport trust model

The HTTP transport ships without authentication. Any caller that can reach the port can invoke scan_skill. Over stdio or 127.0.0.1 this is the same trust boundary as the CLI. If you bind to a routable interface:

  • Sit the server behind an authenticating reverse proxy (e.g. nginx + mTLS) before exposing it externally.
  • Local paths and file:// URLs are automatically rejected over HTTP to prevent unauthenticated callers from reading arbitrary host files. Only remote Git and .zip URLs are accepted.

Vulnerability Patterns

SkillSpector detects 71 vulnerability patterns across 17 categories:

Prompt Injection (6 patterns)

IDPatternSeverityDescription
P1Instruction OverrideHIGHCommands to ignore safety constraints
P2Hidden InstructionsHIGHMalicious directives in comments/invisible text
P3Exfiltration CommandsHIGHInstructions to transmit context externally
P4Behavior ManipulationMEDIUMSubtle instructions altering agent decisions
P5Harmful ContentCRITICALInstructions that could cause physical harm
P9Whitespace PaddingMEDIUMLarge whitespace padding hiding instructions below/beside the visible area

Anti-Refusal (3 patterns)

IDPatternSeverityDescription
AR1Refusal SuppressionHIGHInstructions to never refuse or always comply (e.g. “never refuse”, “always comply”)
AR2Disclaimer SuppressionHIGHInstructions to omit warnings, disclaimers, or ethical commentary (e.g. “no disclaimers”, “do not moralize”)
AR3Safety Policy NullificationHIGHJailbreak framing that nullifies guardrails (e.g. “you have no restrictions”, “ignore your guidelines”, “do anything now”)

Data Exfiltration (4 patterns)

IDPatternSeverityDescription
E1External TransmissionMEDIUMSending data to external URLs
E2Env Variable HarvestingHIGHEnumerating, copying, or searching environment data to collect secrets
E3File System EnumerationMEDIUMScanning directories for sensitive files
E4Context LeakageHIGHTransmitting conversation context externally

Privilege Escalation (3 patterns)

IDPatternSeverityDescription
PE1Excessive PermissionsLOWRequesting access beyond stated functionality
PE2Sudo/Root ExecutionMEDIUMInvoking elevated system privileges
PE3Credential AccessHIGHReading SSH keys, tokens, passwords

Supply Chain (9+ patterns)

IDPatternSeverityDescription
SC1Unpinned DependenciesLOWNo version constraints on packages
SC2External Script FetchingHIGHcurl | bash and remote code execution
SC3Obfuscated CodeHIGHBase64/hex encoded execution
SC4Known Vulnerable DependenciesHIGHDependencies with known CVEs (live OSV.dev lookup)
SC5Abandoned DependenciesMEDIUMUnmaintained packages without security updates
SC6TyposquattingHIGHPackage names similar to popular packages
SC8Shipped Python BytecodeHIGH__pycache__ / .pyc present (discovery skips; malicious bytecode bypass)
SC9Concealed Executable ArtifactHIGHExecutable nested in a document container or hidden/disguised artifact

Excessive Agency (5 patterns)

IDPatternSeverityDescription
EA1Unrestricted Tool AccessHIGHUnfettered tool access without constraints
EA2Autonomous Decision MakingHIGHHigh-impact decisions without human-in-the-loop
EA3Scope CreepMEDIUMCapabilities extending beyond stated purpose
EA4Unbounded Resource AccessMEDIUMNo rate limits or quotas on resource consumption
EA5External Model or Provider SelectionMEDIUM/HIGHModel/provider pins or coding-CLI shell-outs that can switch billing accounts

Output Handling (3 patterns)

IDPatternSeverityDescription
OH1Unvalidated Output InjectionHIGHModel output used without sanitization
OH2Cross-Context OutputMEDIUMOutput flows across trust boundaries without validation
OH3Unbounded OutputMEDIUMNo limits on output size or generation rate

System Prompt Leakage (3 patterns)

IDPatternSeverityDescription
P6Direct LeakageHIGHInstructions that expose system prompts or internal rules
P7Indirect ExtractionMEDIUMExtraction via rephrasing, translation, or side-channels
P8Tool-Based ExfiltrationHIGHSystem prompts exfiltrated via file writes or network requests

Memory Poisoning (3 patterns)

IDPatternSeverityDescription
MP1Persistent Context InjectionHIGHContent designed to persist across interactions
MP2Context Window StuffingMEDIUMFiller content displacing safety constraints
MP3Memory ManipulationHIGHTampering with agent memory or stored state

Tool Misuse (3 patterns)

IDPatternSeverityDescription
TM1Tool Parameter AbuseHIGHCrafted parameters for unintended behavior (shell=True, —force)
TM2Chaining AbuseHIGHTool chains that bypass individual safety checks
TM3Unsafe DefaultsMEDIUMOverly permissive defaults (disabled TLS, no auth)

Rogue Agent (2 patterns)

IDPatternSeverityDescription
RA1Self-ModificationCRITICALModifying own code or configuration at runtime
RA2Session PersistenceHIGHUnauthorized persistence via cron jobs or startup scripts

Trigger Abuse (3 patterns)

IDPatternSeverityDescription
TR1Overly Broad TriggerMEDIUMTrigger patterns matching common words
TR2Shadow Command TriggerHIGHTriggers that shadow built-in commands or other skills
TR3Keyword Baiting TriggerMEDIUMGeneric triggers designed to maximize activation

Behavioral AST (9 patterns)

IDPatternSeverityDescription
AST1exec() CallCRITICALDirect exec() enabling arbitrary code execution
AST2eval() CallHIGHDirect eval() evaluating arbitrary expressions
AST3Dynamic ImportHIGH__import__() loading arbitrary modules at runtime
AST4subprocess CallHIGHExternal command execution via subprocess
AST5os.system / exec-familyHIGHShell commands via os module
AST6compile() CallMEDIUMCode object creation from strings
AST7Dynamic getattr()MEDIUMArbitrary attribute access with non-literal names
AST8Dangerous Execution ChainCRITICALexec/eval combined with dynamic source (network, encoded data)
AST9Reflective getattr() SinkHIGHReflective exec via getattr(os,'system') / getattr(builtins,'exec') that evades AST1/AST5

Taint Tracking (5 patterns)

IDPatternSeverityDescription
TT1Direct Taint FlowHIGHData flows directly from a source to a sink without sanitization
TT2Variable-Mediated Taint FlowMEDIUMData flows from source to sink through intermediate variables
TT3Credential Exfiltration ChainCRITICALCredentials (env vars, secrets) flow to network output sinks
TT4File Read to Network ExfiltrationHIGHFile contents flow to network output sinks
TT5External Input to Code ExecutionCRITICALNetwork or user input flows to exec/eval/subprocess sinks

YARA Signatures (4 patterns)

IDPatternSeverityDescription
YR1Malware MatchCRITICALYARA rule match for known malware signatures
YR2Webshell MatchCRITICALYARA rule match for webshell patterns
YR3Cryptominer MatchHIGHYARA rule match for crypto mining indicators
YR4Hack Tool / Exploit MatchHIGHYARA rule match for hack tools or exploit code

MCP Least Privilege (4 patterns)

IDPatternSeverityDescription
LP1Underdeclared CapabilityHIGHCode uses capabilities not listed in declared permissions
LP2Wildcard PermissionMEDIUMPermission list contains wildcards (*, all, full, any)
LP3Missing Permission DeclarationMEDIUMNo permissions field but code has detectable capabilities
LP4Overdeclared PermissionLOWPermission declared but no corresponding code capability found

This README has been shortened. The full version is on GitHub. Read the original ↗