NVIDIA/SkillSpectorPublic

Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, security risks, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before you install them.

AI summary: A specialized security scanner that detects vulnerabilities and malicious patterns in AI agent skills.

Stars
19.3K
+66 today
Forks
1.7K
Watchers
75
Open issues
54
Open PRs
78
Contributors
~80
Commits
546
Branches
41

PythonApache-2.0Created Mar 21, 2026Last push todayLatest release v2.12.0+1.1K stars this week+3.5K this month

Quick answers

What is SkillSpector?
A specialized security scanner that detects vulnerabilities and malicious patterns in AI agent skills.
What does SkillSpector do?
SkillSpector operates as a highly specialized security tool engineered to rigorously scan and evaluate AI agent skills before they are allowed to be installed. It actively protects systems utilizing tools like Claude Code or Codex by identifying hidden vulnerabilities, prompt injections, and severe data exfiltration risks. The scanner deeply analyzes the skills for any underlying malicious intent or complex supply-chain risks, explicitly ensuring they are safe to execute within trusted, secure environments. It functions as a core, non-negotiable component of the NVIDIA Verified Skills pipeline, which comprehensively vets skills before they are published to the official catalog.
Who is SkillSpector for?
SkillSpector is built specifically for security engineers, AI developers, and enterprise IT teams actively managing autonomous agents. It necessitates familiarity with command-line tools and a deep understanding of AI agent architectures.
How do I get started with SkillSpector?
pip install skillspector
How popular is SkillSpector on GitHub?
NVIDIA/SkillSpector has 19,325 stars and 1,689 forks on GitHub, and gained 1,054 stars in the last 7 days.
What license does SkillSpector use?
NVIDIA/SkillSpector is released under the Apache-2.0 license.

Star history

since Jul 28, 2026
05K10K15KJul 2026Aug 2026Sep 2026Oct 2026
19.3K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 3 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 1 commit2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 1 commit2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 1 commit2026-06-04: 1 commit2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 2 commits2026-06-10: 4 commits2026-06-11: 0 commits2026-06-12: 4 commits2026-06-13: 5 commits2026-06-14: 6 commits2026-06-15: 9 commits2026-06-16: 3 commits2026-06-17: 1 commit2026-06-18: 5 commits2026-06-19: 10 commits2026-06-20: 10 commits2026-06-21: 10 commits2026-06-22: 9 commits2026-06-23: 27 commits2026-06-24: 16 commits2026-06-25: 3 commits2026-06-26: 4 commits2026-06-27: 2 commits2026-06-28: 5 commits2026-06-29: 4 commits2026-06-30: 4 commits2026-07-01: 2 commits2026-07-02: 0 commits2026-07-03: 3 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 1 commit2026-07-07: 1 commit2026-07-08: 0 commits2026-07-09: 2 commits2026-07-10: 4 commits2026-07-11: 5 commits2026-07-12: 0 commits2026-07-13: 1 commit2026-07-14: 1 commit2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 2 commits2026-07-18: 1 commit2026-07-19: 0 commits2026-07-20: 3 commits2026-07-21: 2 commits2026-07-22: 4 commits2026-07-23: 1 commit2026-07-24: 1 commit2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 1 commit2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 2 commits2026-07-31: 7 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 2 commits2026-08-04: 7 commits2026-08-05: 2 commits2026-08-06: 1 commit2026-08-07: 0 commits2026-08-08: 1 commit2026-08-09: 0 commits2026-08-10: 1 commit2026-08-11: 2 commits2026-08-12: 6 commits2026-08-13: 1 commit2026-08-14: 3 commits2026-08-15: 5 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 2 commits2026-08-19: 0 commits2026-08-20: 8 commits2026-08-21: 16 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 6 commits2026-08-25: 5 commits2026-08-26: 5 commits2026-08-27: 5 commits2026-08-28: 3 commits2026-08-29: 0 commits2026-08-30: 2 commits2026-08-31: 9 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 1 commit2026-09-07: 1 commit2026-09-08: 0 commits2026-09-09: 3 commits2026-09-10: 9 commits2026-09-11: 2 commits2026-09-12: 7 commits2026-09-13: 3 commits2026-09-14: 1 commit2026-09-15: 4 commits2026-09-16: 16 commits2026-09-17: 14 commits2026-09-18: 7 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 8 commits2026-09-22: 13 commits2026-09-23: 2 commits2026-09-24: 2 commits2026-09-25: 0 commits2026-09-26: 0 commits
379 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    19,325 stars

  • Rising fast

    +1,054 stars this week

  • Actively maintained

    Pushed within 48 hours

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    10 trending appearances

What SkillSpector does

SkillSpector operates as a highly specialized security tool engineered to rigorously scan and evaluate AI agent skills before they are allowed to be installed. It actively protects systems utilizing tools like Claude Code or Codex by identifying hidden vulnerabilities, prompt injections, and severe data exfiltration risks. The scanner deeply analyzes the skills for any underlying malicious intent or complex supply-chain risks, explicitly ensuring they are safe to execute within trusted, secure environments. It functions as a core, non-negotiable component of the NVIDIA Verified Skills pipeline, which comprehensively vets skills before they are published to the official catalog.

SkillSpector is built specifically for security engineers, AI developers, and enterprise IT teams actively managing autonomous agents. It necessitates familiarity with command-line tools and a deep understanding of AI agent architectures.

  • Vulnerability detection: Rigorously scans agent skills for dozens of known security flaws and complex malicious behavior patterns.
  • Prompt injection defense: Accurately identifies operational vectors where skills might be highly susceptible to targeted prompt injection attacks.
  • Data exfiltration prevention: Deeply analyzes skill behavior to reliably detect and flag potential unauthorized data transfer mechanisms.
  • Supply-chain security: Comprehensively vets skills to actively mitigate the inherent risks associated with untrusted third-party dependencies.
  • NVIDIA pipeline integration: Functions seamlessly as the primary core scanning engine for the broader NVIDIA Verified Skills ecosystem.

Where teams use it

Safe skill installation

Developers can execute SkillSpector to definitively verify that a community-contributed agent skill is safe before adding it to their workflow.

Enterprise security compliance

Information security teams can mandate routine SkillSpector scans to explicitly ensure all AI tools meet strict organizational safety standards.

Skill marketplace vetting

Platform maintainers can seamlessly integrate the scanner into their CI/CD pipelines to automatically vet all newly submitted skills.

Vulnerability research

Dedicated security researchers can utilize the tool to systematically analyze the broader prevalence of vulnerabilities in the AI agent ecosystem.

Getting started: pip install skillspector

README

main branch

SkillSpector

Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks before installing agent skills.

Python 3.12+ License: Apache 2.0

Overview

AI agent skills (used by Claude Code, Codex CLI, Gemini CLI, etc.) execute with implicit trust and minimal vetting. In the 31,132-skill analyzed subset of the research dataset, 26.1% of skills contain vulnerabilities and 5.2% show likely malicious intent.

SkillSpector helps you answer: "Is this skill safe to install?"

SkillSpector is part of the NVIDIA Verified Skills pipeline, which scans, evaluates, and signs agent skills before publication. Skills that pass are published to the NVIDIA skills catalog.

Documentation

  • Scan agent skills before installation — Hosted guide: when to scan, how to read a report, and how to gate installs.
  • Development guide — Architecture, package layout, and how to extend the analyzer pipeline.
  • Analysis resource bounds — Fail-closed bundle, parser, nested-artifact, ledger, and finding ceilings.
  • Pi extension — Install SkillSpector as a Pi tool for scanning skills from inside agent sessions.
  • OpenCode extension — Install SkillSpector as an OpenCode tool and /skillspector command for scanning skills from inside agent sessions.

Features

  • Multi-format input: Scan Git repos, URLs, zip files, directories, or single files
  • 71 vulnerability patterns across 17 categories: prompt injection, data exfiltration, privilege escalation, supply chain, excessive agency, output handling, system prompt leakage, memory poisoning, tool misuse, rogue agent, anti-refusal, trigger abuse, dangerous code (AST), taint tracking, YARA signatures, MCP least privilege, and MCP tool poisoning
  • Two-stage analysis: Fast static analysis + optional LLM semantic evaluation
  • Live vulnerability lookups: SC4 queries OSV.dev for real-time CVE data with automatic offline fallback
  • Multiple output formats: Terminal, JSON, Markdown, and SARIF reports
  • Risk scoring: 0-100 score with severity labels and clear recommendations
  • Baseline / false-positive suppression: Accept known findings via a glob-rule or fingerprint baseline so re-scans surface only new issues (docs)

Quick Start

Installation

Open-source software notice: This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

Create and activate a virtual environment first (all make targets assume the venv is active). Use uv or pip; the Makefile uses uv if available, otherwise pip.

Quick install with uv (CLI-only):

uv tool install git+https://github.com/NVIDIA/skillspector.git
# Update later: uv tool update skillspector

If you plan to run skillspector mcp, install the MCP extra at install time:

uv tool install 'skillspector[mcp] @ git+https://github.com/NVIDIA/skillspector.git'

From source:

# Clone the repository
git clone https://github.com/NVIDIA/skillspector.git
cd skillspector

# Create and activate virtual environment
uv venv .venv && source .venv/bin/activate
# or: python3 -m venv .venv && source .venv/bin/activate

# Install for production use
make install

# Or install with development dependencies
make install-dev

Docker (no Python required)

Run SkillSpector without installing Python by building it locally from the included Dockerfile. The image is based on the Docker Official Python 3.12-slim-bookworm image.

Build the image:

make docker-build
# or: docker build -t skillspector .

Scan a local directory by mounting your current directory into /scan, the container's working directory:

docker run --rm -v "$PWD:/scan" skillspector scan ./my-skill/ --no-llm

Scan with LLM analysis by passing credentials with a local .env file:

cat > .env <<'EOF'
SKILLSPECTOR_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
EOF
docker run --rm \
  -v "$PWD:/scan" \
  --env-file .env \
  skillspector scan ./my-skill/

Or pass credentials directly from your shell environment:

docker run --rm \
  -v "$PWD:/scan" \
  -e SKILLSPECTOR_PROVIDER=anthropic \
  -e ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY" \
  skillspector scan ./my-skill/

Write a report to the host filesystem by writing to the mounted directory:

docker run --rm \
  -v "$PWD:/scan" \
  skillspector scan ./my-skill/ --no-llm --format json --output report.json

Optional alias for repeated static scans:

alias skillspector-docker='docker run --rm -v "$PWD:/scan" skillspector'
skillspector-docker scan ./my-skill/ --no-llm

Basic Usage

# Scan a local skill directory
skillspector scan ./my-skill/

# Scan a single SKILL.md file
skillspector scan ./SKILL.md

# Scan a Git repository
skillspector scan https://github.com/user/my-skill

# Scan a zip file
skillspector scan ./my-skill.zip
Size limits

SkillSpector enforces two independent caps on remote and archive inputs to bound the impact of oversized downloads and zip bombs:

  • Per-ingest cap: INGEST_MAX_BYTES (100 MiB) — applied to streamed URL downloads, total uncompressed size of zip archives, and post-clone disk usage of Git repos.
  • Zip member cap: INGEST_MAX_ZIP_MEMBERS (10,000) — caps the number of entries in a single zip.

Note that the per-file 1 MB analysis cap (MAX_FILE_BYTES) is a separate, downstream limit: it bounds what individual analyzers will read out of an already-ingested directory. The ingest caps above bound how much content can land on disk in the first place. A breach of either ingest cap fails closed with an IngestLimitExceededError.

Output Formats

# Terminal output (default) - pretty formatted
skillspector scan ./my-skill/

# JSON output - machine readable
skillspector scan ./my-skill/ --format json --output report.json

# Markdown output - for documentation
skillspector scan ./my-skill/ --format markdown --output report.md

# SARIF output - for CI/CD integration and IDE tooling
skillspector scan ./my-skill/ --format sarif --output report.sarif

Batch Scanning

Scan entire directories of skills in parallel from contrib/batch_scan/:

python -m contrib.batch_scan.batch_scan ./my-skills/ --no-llm
python -m contrib.batch_scan.batch_scan ./my-skills/ --workers 20 -f json -o report.json
python -m contrib.batch_scan.batch_scan ./tests/fixtures/ -f terminal --workers 20

Supports multilingual detection (zh/ja/ko) and terminal/JSON/Markdown output.

For LLM scans with higher concurrency, configure multiple API keys following .env.example — the pool improves throughput and resilience, provided the keys don't share an account-level rate limit.

See the contrib guide for details.

Note on LLM support: The default configuration targets DeepSeek as the cheapest public option. DeepSeek-Chat is expected to sunset, and the contributor does not have hardware to test against local models. The batch scanner was originally tested with OpenAI-compatible endpoints — DeepSeek's lack of structured-output support required manual JSON-parsing patches. If you can contribute a more universal backend (Ollama, vLLM, or a different provider), PRs are very welcome.

Suppressing False Positives (baseline)

Suppress known/accepted findings so the risk score reflects only un-triaged issues and re-scans surface only new findings. See the suppression guide for the full reference.

# Accept all current findings into a baseline (run once), then commit it.
skillspector baseline ./my-skill/ -o .skillspector-baseline.yaml

# Scan against the baseline — only NEW findings are reported and scored.
skillspector scan ./my-skill/ --baseline .skillspector-baseline.yaml

# Review what was suppressed (still excluded from the score).
skillspector scan ./my-skill/ --baseline .skillspector-baseline.yaml --show-suppressed

A baseline can also use drift-tolerant glob rules (by rule id, file path, or message) — see .skillspector-baseline.example.yaml. Exact fingerprint baselines are evidence-bound: changing the scanned source or SkillSpector version keeps the finding active until it is reviewed again. When a selected baseline or baseline output is stored inside the skill directory, SkillSpector excludes that exact file from content analysis so its suppression text cannot create findings or enter regenerated fingerprints; sibling files remain in normal scan scope.

LLM Analysis

For the best results, configure an OpenAI-compatible LLM endpoint for semantic analysis. Pick a provider with SKILLSPECTOR_PROVIDER; hosted providers ship bundled default models, while CLI providers fall back to the local runtime's default model unless SKILLSPECTOR_MODEL is set. SkillSpector also works against local OpenAI-compatible servers (Ollama, vLLM, llama.cpp) and managed inference gateways.

Provider (SKILLSPECTOR_PROVIDER) Credential env var Endpoint Default model
openai OPENAI_API_KEY (+ optional OPENAI_BASE_URL) api.openai.com (or any OpenAI-compatible URL) gpt-5.4
anthropic ANTHROPIC_API_KEY api.anthropic.com claude-opus-4-6
anthropic_proxy ANTHROPIC_PROXY_API_KEY + ANTHROPIC_PROXY_ENDPOINT_URL Any Vertex-style raw-predict proxy claude-sonnet-4-6
bedrock AWS_PROFILE (optional) + AWS_REGION — SigV4 via boto3 AWS Bedrock Runtime us.anthropic.claude-sonnet-4-6-20250915-v1:0
nv_build NVIDIA_INFERENCE_KEY build.nvidia.com z-ai/glm-5.2
ollama (none) OLLAMA_BASE_URL (default http://localhost:11434/v1) llama3.1:8b
azure_openai AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT Azure OpenAI Service gpt-4o (deployment defaults to the model label)
openai_compatible SKILLSPECTOR_COMPAT_API_KEY + SKILLSPECTOR_COMPAT_BASE_URL Any OpenAI-compatible endpoint llama-3.1-70b-versatile
claude_cli (none — uses local CLI auth) local claude binary local Claude runtime fallback, or SKILLSPECTOR_MODEL
codex_cli (none — uses local CLI auth) local codex binary local Codex runtime fallback, or SKILLSPECTOR_MODEL
gemini_cli (none — uses local CLI auth) local gemini binary local Gemini runtime fallback, or SKILLSPECTOR_MODEL
opencode_cli (none — uses local CLI auth) local opencode 1.18.31 binary local OpenCode runtime fallback, or SKILLSPECTOR_MODEL
# Stock OpenAI
export SKILLSPECTOR_PROVIDER=openai
export OPENAI_API_KEY=sk-...
skillspector scan ./my-skill/

# Anthropic
export SKILLSPECTOR_PROVIDER=anthropic
export ANTHROPIC_API_KEY=sk-ant-...
skillspector scan ./my-skill/

# Anthropic via Vertex-style proxy (corporate gateways, GCP Vertex AI)
export SKILLSPECTOR_PROVIDER=anthropic_proxy
export ANTHROPIC_PROXY_ENDPOINT_URL=https://my-gateway.example.com/models/claude-sonnet-4-6:streamRawPredict
export ANTHROPIC_PROXY_API_KEY=your-bearer-token
export SKILLSPECTOR_MODEL=claude-sonnet-4-6
skillspector scan ./my-skill/

# AWS Bedrock (Claude via SigV4)
export SKILLSPECTOR_PROVIDER=bedrock
# Optional: select an AWS named profile. When unset, the standard
# boto3 credential chain (env vars, instance metadata, SSO, etc.) resolves.
# export AWS_PROFILE=my-profile
export AWS_REGION=us-west-2  # default if unset
# Default model: us.anthropic.claude-sonnet-4-6-20250915-v1:0
# Override with any Bedrock model ID, cross-region inference-profile
# ID, or your own application-inference-profile ARN:
# export SKILLSPECTOR_MODEL=us.anthropic.claude-opus-4-6-20250915-v1:0
skillspector scan ./my-skill/

# NVIDIA build.nvidia.com
export SKILLSPECTOR_PROVIDER=nv_build
export NVIDIA_INFERENCE_KEY=nvapi-...
skillspector scan ./my-skill/

# Local Claude CLI — no API key; uses your existing `claude auth login` session
# Requires: claude CLI installed and authenticated (claude auth login)
export SKILLSPECTOR_PROVIDER=claude_cli
# Uses the local Claude CLI runtime fallback unless SKILLSPECTOR_MODEL is set.
# export SKILLSPECTOR_MODEL=claude-sonnet-4-6
skillspector scan ./my-skill/

# Local Codex CLI — no API key; uses your existing `codex login` session
# Requires: codex CLI installed and authenticated
export SKILLSPECTOR_PROVIDER=codex_cli
skillspector scan ./my-skill/

# Gemini (via OpenAI compatibility layer)
export SKILLSPECTOR_PROVIDER=openai
export OPENAI_API_KEY="YOUR_GEMINI_API_KEY"
export OPENAI_BASE_URL="https://generativelanguage.googleapis.com/v1beta/openai/"
export SKILLSPECTOR_MODEL=gemini-3.5-flash
skillspector scan ./my-skill/

# Local Ollama — no API key
export SKILLSPECTOR_PROVIDER=ollama
# export OLLAMA_BASE_URL=http://localhost:11434/v1  # shown default
export SKILLSPECTOR_MODEL=llama3.1:8b
skillspector scan ./my-skill/

# Azure OpenAI
export SKILLSPECTOR_PROVIDER=azure_openai
export AZURE_OPENAI_API_KEY=...
export AZURE_OPENAI_ENDPOINT=https://example.openai.azure.com/
export AZURE_OPENAI_DEPLOYMENT=my-deployment
skillspector scan ./my-skill/

# Any other OpenAI-compatible endpoint
export SKILLSPECTOR_PROVIDER=openai_compatible
export SKILLSPECTOR_COMPAT_API_KEY=...
export SKILLSPECTOR_COMPAT_BASE_URL=https://api.groq.com/openai/v1
export SKILLSPECTOR_MODEL=llama-3.1-70b-versatile
skillspector scan ./my-skill/

# Override the provider's default model
export SKILLSPECTOR_MODEL=gpt-5.2
skillspector scan ./my-skill/

# Skip LLM analysis (faster, static analysis only)
skillspector scan ./my-skill/ --no-llm

MCP Server

Run SkillSpector as a Model Context Protocol server so any MCP-capable agent (Claude Code, Codex CLI, Gemini CLI) or remote runtime can call scanning as a tool and gate skill/MCP installs on the result — turning SkillSpector into a runtime guardrail instead of an out-of-band audit step.

skillspector mcp requires skillspector[mcp].

# Install, or reinstall if you already used the CLI-only path
uv tool install --force 'skillspector[mcp] @ git+https://github.com/NVIDIA/skillspector.git'

# FastMCP stdio transport for local CLI agents
skillspector mcp

# streamable HTTP/SSE transport for remote / A2A callers
skillspector mcp --transport http --host 127.0.0.1 --port 8000

The stdio transport is the current FastMCP path for local CLI agents, and the initialize hang reported in issue #199 still applies there.

The server exposes a single tool:

  • scan_skill(target, use_llm=true, output_format="json") — scans a Git URL, file URL, .zip, .md file, or directory and returns a structured verdict: risk_score (0-100), severity, recommendation, safe_to_install, and findings. It also reports llm_used / scan_mode so a low score from a static-only scan is never mistaken for a clean full scan.

Register it with Claude Code via:

claude mcp add skillspector -- skillspector mcp

Security — HTTP transport trust model

The HTTP transport ships without authentication. Any caller that can reach the port can invoke scan_skill. Over stdio or 127.0.0.1 this is the same trust boundary as the CLI. If you bind to a routable interface:

  • Sit the server behind an authenticating reverse proxy (e.g. nginx + mTLS) before exposing it externally.
  • Local paths and file:// URLs are automatically rejected over HTTP to prevent unauthenticated callers from reading arbitrary host files. Only remote Git and .zip URLs are accepted.

Vulnerability Patterns

SkillSpector detects 71 vulnerability patterns across 17 categories:

Prompt Injection (6 patterns)

ID Pattern Severity Description
P1 Instruction Override HIGH Commands to ignore safety constraints
P2 Hidden Instructions HIGH Malicious directives in comments/invisible text
P3 Exfiltration Commands HIGH Instructions to transmit context externally
P4 Behavior Manipulation MEDIUM Subtle instructions altering agent decisions
P5 Harmful Content CRITICAL Instructions that could cause physical harm
P9 Whitespace Padding MEDIUM Large whitespace padding hiding instructions below/beside the visible area

Anti-Refusal (3 patterns)

ID Pattern Severity Description
AR1 Refusal Suppression HIGH Instructions to never refuse or always comply (e.g. "never refuse", "always comply")
AR2 Disclaimer Suppression HIGH Instructions to omit warnings, disclaimers, or ethical commentary (e.g. "no disclaimers", "do not moralize")
AR3 Safety Policy Nullification HIGH Jailbreak framing that nullifies guardrails (e.g. "you have no restrictions", "ignore your guidelines", "do anything now")

Data Exfiltration (4 patterns)

ID Pattern Severity Description
E1 External Transmission MEDIUM Sending data to external URLs
E2 Env Variable Harvesting HIGH Enumerating, copying, or searching environment data to collect secrets
E3 File System Enumeration MEDIUM Scanning directories for sensitive files
E4 Context Leakage HIGH Transmitting conversation context externally

Privilege Escalation (3 patterns)

ID Pattern Severity Description
PE1 Excessive Permissions LOW Requesting access beyond stated functionality
PE2 Sudo/Root Execution MEDIUM Invoking elevated system privileges
PE3 Credential Access HIGH Reading SSH keys, tokens, passwords

Supply Chain (10+ patterns)

ID Pattern Severity Description
SC1 Unpinned Dependencies LOW No version constraints on packages
SC2 External Script Fetching HIGH curl | bash and remote code execution
SC3 Obfuscated Code HIGH Base64/hex encoded execution
SC4 Known Vulnerable Dependencies HIGH Dependencies with known CVEs (live OSV.dev lookup)
SC5 Abandoned Dependencies MEDIUM Unmaintained packages without security updates
SC6 Typosquatting HIGH Package names similar to popular packages
SC8 Shipped Python Bytecode HIGH __pycache__ / .pyc present (discovery skips; malicious bytecode bypass)
SC9 Concealed Executable Artifact HIGH Executable nested in a document container or hidden/disguised artifact
SC10 Dependency Source Redirection HIGH Package-manager source added, replaced, or unresolved

Excessive Agency (5 patterns)

ID Pattern Severity Description
EA1 Unrestricted Tool Access HIGH Unfettered tool access without constraints
EA2 Autonomous Decision Making HIGH High-impact decisions without human-in-the-loop
EA3 Scope Creep MEDIUM Capabilities extending beyond stated purpose
EA4 Unbounded Resource Access MEDIUM No rate limits or quotas on resource consumption
EA5 External Model or Provider Selection MEDIUM/HIGH Model/provider pins or coding-CLI shell-outs that can switch billing accounts

Output Handling (3 patterns)

ID Pattern Severity Description
OH1 Unvalidated Output Injection HIGH Model output used without sanitization
OH2 Cross-Context Output MEDIUM Output flows across trust boundaries without validation
OH3 Unbounded Output MEDIUM No limits on output size or generation rate

System Prompt Leakage (3 patterns)

ID Pattern Severity Description
P6 Direct Leakage HIGH Instructions that expose system prompts or internal rules
P7 Indirect Extraction MEDIUM Extraction via rephrasing, translation, or side-channels
P8 Tool-Based Exfiltration HIGH System prompts exfiltrated via file writes or network requests

Memory Poisoning (3 patterns)

ID Pattern Severity Description
MP1 Persistent Context Injection HIGH Content designed to persist across interactions
MP2 Context Window Stuffing MEDIUM Filler content displacing safety constraints
MP3 Memory Manipulation HIGH Tampering with agent memory or stored state

Tool Misuse (3 patterns)

ID Pattern Severity Description
TM1 Tool Parameter Abuse HIGH Crafted parameters for unintended behavior (shell=True, --force)
TM2 Chaining Abuse HIGH Tool chains that bypass individual safety checks
TM3 Unsafe Defaults MEDIUM Overly permissive defaults (disabled TLS, no auth)

Rogue Agent (2 patterns)

ID Pattern Severity Description
RA1 Self-Modification CRITICAL Modifying own code or configuration at runtime
RA2 Session Persistence HIGH Unauthorized persistence via cron jobs or startup scripts

Trigger Abuse (3 patterns)

ID Pattern Severity Description
TR1 Overly Broad Trigger MEDIUM Trigger patterns matching common words
TR2 Shadow Command Trigger HIGH Triggers that shadow built-in commands or other skills
TR3 Keyword Baiting Trigger MEDIUM Generic triggers designed to maximize activation

Behavioral AST (9 patterns)

ID Pattern Severity Description
AST1 exec() Call CRITICAL Direct exec() enabling arbitrary code execution
AST2 eval() Call HIGH Direct eval() evaluating arbitrary expressions
AST3 Dynamic Import HIGH __import__() loading arbitrary modules at runtime
AST4 subprocess Call HIGH External command execution via subprocess
AST5 os.system / exec-family HIGH Shell commands via os module
AST6 compile() Call MEDIUM Code object creation from strings
AST7 Dynamic getattr() MEDIUM Arbitrary attribute access with non-literal names
AST8 Dangerous Execution Chain CRITICAL exec/eval combined with dynamic source (network, encoded data)
AST9 Reflective getattr() Sink HIGH Reflective exec via getattr(os,'system') / getattr(builtins,'exec') that evades AST1/AST5

Taint Tracking (5 patterns)

ID Pattern Severity Description
TT1 Direct Taint Flow HIGH Data flows directly from a source to a sink without sanitization
TT2 Variable-Mediated Taint Flow MEDIUM Data flows from source to sink through intermediate variables
TT3 Credential Exfiltration Chain CRITICAL Credentials (env vars, secrets) flow to network output sinks
TT4 File Read to Network Exfiltration HIGH File contents flow to network output sinks
TT5 External Input to Code Execution CRITICAL Network or user input flows to exec/eval/subprocess sinks

YARA Signatures (4 patterns)

ID Pattern Severity Description
YR1 Malware Match CRITICAL YARA rule match for known malware signatures
YR2 Webshell Match CRITICAL YARA rule match for webshell patterns
YR3 Cryptominer Match HIGH YARA rule match for crypto mining indicators
YR4 Hack Tool / Exploit Match HIGH YARA rule match for hack tools or exploit code

MCP Least Privilege (4 patterns)

ID Pattern Severity Description
LP1 Underdeclared Capability HIGH Code uses capabilities not listed in declared permissions
LP2 Wildcard Permission MEDIUM Permission list contains wildcards (*, all, full, any)
LP3 Missing Permission Declaration MEDIUM No permissions field but code has detectable capabilities
LP4 Overdeclared Permission LOW Permission declared but no corresponding code capability found

MCP Tool Poisoning (4 patterns)

ID Pattern Severity Description
TP1 Hidden Instructions HIGH Hidden directives in metadata (HTML comments, zero-width chars, base64, data URIs)
TP2 Unicode Deception HIGH Homoglyphs, RTL overrides, mixed-script identifiers in tool metadata
TP3 Parameter Description Injection MEDIUM Injection patterns in parameter definitions (overrides, system tokens, malicious defaults)
TP4 Description-Behavior Mismatch MEDIUM Declared tool description does not match actual code behavior (LLM-powered)

All detected patterns are listed in the tables above.

Risk Scoring

Score Calculation

  • CRITICAL issues: +50 points
  • HIGH issues: +25 points
  • MEDIUM issues: +10 points
  • LOW issues: +5 points
  • Executable scripts: 1.3x multiplier

Severity Levels

Score Severity Recommendation
0-20 LOW SAFE
21-50 MEDIUM CAUTION
51-80 HIGH DO NOT INSTALL
81-100 CRITICAL DO NOT INSTALL

Example Output

Terminal Output

 SkillSpector Security Report  v2.0.0

Skill: suspicious-skill
Source: ./suspicious-skill/
Scanned: 2026-01-29 10:30:00 UTC

        Risk Assessment
 Metric          Value
 Score           78/100
 Severity        HIGH
 Recommendation  DO NOT INSTALL

        Components (3)
 File              Type      Lines  Executable
 SKILL.md          markdown    142  No
 scripts/sync.py   python       87  Yes
 requirements.txt  text          3  No

Issues (2)

  HIGH: Env Variable Harvesting (E2)
    Location: scripts/sync.py:23
    Finding: for key, val in os.environ.items():...
    Confidence: 94%
    Explanation: This code collects environment variables containing
    API keys and secrets, then sends them to an external server.

  HIGH: External Transmission (E1)
    Location: scripts/sync.py:45
    Finding: requests.post("https://api.skill.io/env"...
    Confidence: 89%
    Explanation: Data is being sent to an external server. Combined
    with env harvesting above, this indicates credential exfiltration.

Configuration

Environment Variables

Variable Description Required
SKILLSPECTOR_PROVIDER Active LLM provider: openai, anthropic, anthropic_proxy, bedrock, nv_build, ollama, azure_openai, openai_compatible, claude_cli, codex_cli, gemini_cli, or opencode_cli. Hosted providers use bundled model_registry.yaml defaults; CLI providers fall back to the local runtime's default model unless SKILLSPECTOR_MODEL is set. Defaults to nv_build. Optional
NVIDIA_INFERENCE_KEY Credential for the nv_build provider (build.nvidia.com). Required for LLM analysis when SKILLSPECTOR_PROVIDER=nv_build
OPENAI_API_KEY Credential for the OpenAI provider (SKILLSPECTOR_PROVIDER=openai). Also serves as the tier-2 fallback in the credential waterfall when the active provider returns no credentials. Required for LLM analysis when SKILLSPECTOR_PROVIDER=openai
OPENAI_BASE_URL Override the OpenAI endpoint (e.g. point at Ollama). Optional
SKILLSPECTOR_REASONING_EFFORT Optional provider- and model-dependent reasoning-effort setting. Non-empty values are trimmed and passed through unchanged; unset or blank preserves provider-default behavior. Optional
SKILLSPECTOR_OUTPUT_LANGUAGE Short, single-line language label (letters, numbers, spaces, _, or -; maximum 64 characters) for human-readable LLM finding text such as messages, explanations, and remediation. Rule IDs, severity values, paths, code, and other machine-readable values remain unchanged. Unset, blank, or invalid values preserve the default output language. Optional
SKILLSPECTOR_TEMPERATURE Optional sampling temperature from 0 to 1 for hosted providers. Unset or blank preserves the provider default. Lower values can reduce run-to-run variation but do not guarantee identical output. Optional
SKILLSPECTOR_SEED Optional integer sampling seed for OpenAI-compatible and Azure OpenAI providers. Other hosted providers and CLI providers do not receive it. Provider support remains model-dependent. Optional
SKILLSPECTOR_COMPACT_PROMPTS Opt-in compact line numbering in LLM prompts: numbered lines render as L1:, L2: instead of zero-padded L01:, L02:. Accepted truthy values are 1, true, and yes (case-insensitive; surrounding whitespace is trimmed). Unset or any other value keeps the default zero-padded format. Optional
ANTHROPIC_API_KEY Credential for the Anthropic provider (SKILLSPECTOR_PROVIDER=anthropic). Required for LLM analysis when SKILLSPECTOR_PROVIDER=anthropic
ANTHROPIC_BASE_URL Override the native Anthropic endpoint (default: https://api.anthropic.com). Optional
ANTHROPIC_PROXY_ENDPOINT_URL Full endpoint URL for the Anthropic proxy provider (Vertex-style raw-predict). Required when SKILLSPECTOR_PROVIDER=anthropic_proxy
ANTHROPIC_PROXY_API_KEY Bearer token for the Anthropic proxy provider. Required when SKILLSPECTOR_PROVIDER=anthropic_proxy
ANTHROPIC_PROXY_API_VERSION anthropic_version value sent in the request body (default: vertex-2023-10-16). Optional
AWS_PROFILE Named AWS profile for the Bedrock provider — authenticates via SigV4 through boto3. When unset, the standard boto3 credential chain (env vars, instance metadata, SSO, etc.) resolves. Optional (used when SKILLSPECTOR_PROVIDER=bedrock)
AWS_REGION AWS region for the Bedrock Runtime endpoint. Defaults to us-west-2. Optional (used when SKILLSPECTOR_PROVIDER=bedrock)
OLLAMA_BASE_URL Ollama OpenAI-compatible endpoint. Defaults to http://localhost:11434/v1. Optional (used when SKILLSPECTOR_PROVIDER=ollama)
AZURE_OPENAI_API_KEY API key for the Azure OpenAI provider. Required when SKILLSPECTOR_PROVIDER=azure_openai
AZURE_OPENAI_ENDPOINT Azure resource endpoint for the Azure OpenAI provider. Required when SKILLSPECTOR_PROVIDER=azure_openai
AZURE_OPENAI_DEPLOYMENT Azure deployment name. Defaults to the selected model label. Optional
AZURE_OPENAI_API_VERSION Azure OpenAI API version. Defaults to 2024-06-01. Optional
SKILLSPECTOR_COMPAT_API_KEY API key for a generic OpenAI-compatible provider. Required when SKILLSPECTOR_PROVIDER=openai_compatible
SKILLSPECTOR_COMPAT_BASE_URL Base URL for a generic OpenAI-compatible provider. Required when SKILLSPECTOR_PROVIDER=openai_compatible
SKILLSPECTOR_MODEL Override the active provider model. For hosted providers, this replaces the bundled default from the LLM Analysis table. For CLI providers, this is forwarded as --model instead of using the local runtime fallback. Optional
SKILLSPECTOR_MODEL_REGISTRY Override the bundled per-provider YAML registry (src/skillspector/providers/<provider>/model_registry.yaml) with a custom path. Optional
SKILLSPECTOR_LOG_LEVEL Log level: DEBUG, INFO, WARNING, ERROR (default: WARNING). Optional

CLI providers (claude_cli, codex_cli, gemini_cli, opencode_cli): No API key is needed. Authentication is managed entirely by the agent CLI's own login session. SkillSpector never reads or forwards API keys when these providers are active. The subprocess is run with capabilities restricted, and untrusted skill content is delivered only via stdin.

opencode_cli currently fails closed unless the installed OpenCode version is exactly 1.18.31, the version whose configuration precedence and deny-all semantics are verified by this release.

CLI Options

skillspector scan --help

Options:
  -f, --format [terminal|json|markdown|sarif]  Output format [default: terminal]
  -o, --output PATH                            Output file path
  --no-llm                                     Skip LLM analysis (static only)
  --yara-rules-dir PATH                        Extra YARA rules directory
  -b, --baseline PATH                          Suppress findings listed in a baseline
  --show-suppressed                            List baseline-suppressed findings
  -V, --verbose                                Show detailed progress
  --help                                       Show this message and exit

# Generate a baseline of all current findings (see docs/SUPPRESSION.md)
skillspector baseline <path> [-o FILE] [--no-llm] [--reason TEXT]

Integrating SkillSpector

SkillSpector is built to be driven by other tools (CI pipelines, install gates, editor integrations). Its exit code and JSON output are a stable contract.

Exit codes

skillspector scan exits with:

Code Meaning
0 Scan completed, risk_score ≤ 50 (recommendation SAFE or CAUTION), and no enabled strict gate fired
1 Scan completed and either risk_score > 50, --fail-on-findings found an active finding, or --fail-on-incomplete found partial/incomplete analysis
2 Error (bad input, unreadable source, internal failure)

By default, the exit code collapses SAFE and CAUTION into 0. Use --fail-on-findings to gate on any active finding, --fail-on-incomplete to gate on incomplete coverage, or read the JSON recommendation field for custom policy.

Machine-readable output

--format json produces a JSON report; with no --output/-o it is written to stdout:

skillspector scan ./my-skill/ --format json

The top-level shape is (this example shows a full LLM-backed scan; with --no-llm, metadata.llm_requested is false):

{
  "skill": { "name": "...", "source": "...", "scanned_at": "<ISO 8601>" },
  "risk_assessment": { "score": 0, "severity": "LOW", "recommendation": "SAFE" },
  "components": [ { "path": "...", "type": "...", "lines": 0, "executable": false, "size_bytes": 0 } ],
  "issues": [ { "id": "...", "category": "...", "severity": "...", "confidence": 0.0, "location": { "file": "...", "start_line": 0 } } ],
  "metadata": {
    "has_executable_scripts": false,
    "skillspector_version": "...",
    "llm_requested": true,
    "llm_available": true,
    "inference_usage": [
      {
        "node": "semantic_security_discovery",
        "request_kind": "structured_output",
        "provider": "nv_inference",
        "model": "azure/anthropic/claude-opus-4-6",
        "model_source": "provider_response",
        "usage_source": "provider_response",
        "prompt_tokens": 1000,
        "completion_tokens": 100,
        "cached_tokens": 400,
        "cache_write_tokens": 50,
        "total_tokens": 1100
      }
    ]
  }
}
  • risk_assessment.severity ∈ LOW | MEDIUM | HIGH | CRITICAL.
  • risk_assessment.recommendation ∈ SAFE | CAUTION | DO_NOT_INSTALL, mapped from severity: LOW → SAFE, MEDIUM → CAUTION, HIGH/CRITICAL → DO_NOT_INSTALL.
  • metadata.llm_error appears only when LLM analysis was requested but unavailable.
  • AE1 findings use Incomplete referenced artifact analysis. Their source location identifies the reference; evidence identifies the affected target, analyzer reasons, and available bounds. Review the target's completeness ledger when reasons_truncated is true. See referenced-artifact diagnostics and Perl help text for interpretation and corrective actions.
  • metadata.inference_usage contains one sanitized record per LLM response when the provider exposes token counters. It is an empty list when usage is unavailable; SkillSpector never estimates missing tokens. Prompt totals are inclusive of cache reads and writes so downstream pricing can separate those partitions safely. model_source distinguishes an independently identified provider model from the exact requested model used when response identity is absent or ambiguous. SkillSpector does not currently send Anthropic prompt-cache controls, so its scan requests cannot select the separate 5-minute or 1-hour cache-write tiers; TTL-specific response fields are normalized defensively into the aggregate cache-write counter.
  • See Inference usage telemetry for the complete provenance, cache-accounting, privacy, fail-closed ingestion, and downstream pricing contract.
  • The full per-issue shape is defined by Finding.to_dict() in models.py; rely on the fields above and treat any additional fields as best-effort.

For CI/IDE tooling, --format sarif emits SARIF 2.1.0.

Recommended gate mapping

When using SkillSpector as an install gate, map the recommendation to an action:

recommendation Suggested action
SAFE allow
CAUTION prompt / warn the user
DO_NOT_INSTALL block

SkillSpector computes the score band and recommendation; how strict the gate is (e.g. whether CAUTION blocks in CI) is a policy decision for the integrating tool.

Development

Setup

All make targets assume a virtual environment is already created and activated. The Makefile uses uv if available, else pip.

# Clone, create venv, activate, install dev dependencies
git clone https://github.com/NVIDIA/skillspector.git
cd skillspector
uv venv .venv && source .venv/bin/activate
# or: python3 -m venv .venv && source .venv/bin/activate
make install-dev

# Run tests
make test

# Run tests with coverage
make test-cov

# Run linting
make lint

# Format code
make format

How It Works

SkillSpector uses a two-stage detection pipeline:

Stage 1: Static Analysis

  • Fast regex-based pattern matching across 11 static analyzers
  • AST-based behavioral analysis detecting dangerous calls (exec, eval, subprocess, etc.)
  • Live vulnerability lookups via OSV.dev for known CVEs in dependencies
  • Scans all analyzer-eligible files in the skill
  • High recall (catches most issues)
  • Moderate precision (some false positives)

A valid, root-level OpenSSF Model Signing signature (skill.oms.sig) is retained in the component inventory as type oms_signature, but excluded from static and LLM content analysis. OMS bundles necessarily contain long base64-encoded payload, signature, and certificate fields; generic obfuscated-code checks can otherwise misclassify those fields as hidden executable content. The recognizer checks the minimal OMS DSSE/in-toto structure; it does not verify the signature, certificate chain, transparency-log entry, or signer identity. Invalid or unrecognized signature files are scanned normally.

Stage 2: LLM Semantic Analysis (Optional)

  • Evaluates context and intent
  • Filters false positives
  • Provides human-readable explanations
  • Improves precision to ~87%

The LLM prompt includes anti-jailbreak protections to prevent malicious skills from manipulating the analysis.

Live Vulnerability Lookups (SC4)

SC4 uses the OSV.dev API to check dependencies against the full Open Source Vulnerabilities database — covering tens of thousands of advisories across PyPI and npm.

  • No API key required — OSV.dev is free and unauthenticated.
  • Batch queries — all dependencies are checked in a single HTTP call.
  • Automatic fallback — if OSV.dev is unreachable (air-gapped/offline), a small built-in fallback list is used.
  • Caching — results are cached in-memory for 1 hour to avoid redundant API calls during a session.

The tool requires outbound HTTPS access to api.osv.dev for live vulnerability data. When that is not available, findings are limited to the static fallback list.

Trust model and data egress

SkillSpector is defense-in-depth, not a sandbox. Know what it does and does not do before relying on it:

  • It never executes the scanned skill. All analysis is static (regex, Python AST, YARA) plus optional LLM evaluation of file contents — the skill's code is never run.
  • LLM analysis sends analyzer-eligible file contents to the configured provider. When LLM analysis is enabled (the default), file contents are sent to the active SKILLSPECTOR_PROVIDER endpoint. Recognized OMS signature files are excluded. Use --no-llm to keep contents local (static analysis only).
  • SC4 sends dependency names to OSV.dev. The supply-chain check queries OSV.dev with the package names and versions the skill declares, to look up known CVEs. This is fundamental to the check and runs even with --no-llm. It sends dependency coordinates (not file contents), requires no API key, and falls back to a bundled list when OSV.dev is unreachable.
  • It does not sandbox the host. SkillSpector flags risky patterns before you install a skill; it does not contain or isolate a skill you choose to install anyway.

Limitations

  • Non-English content: May miss patterns in other languages
  • Image-based attacks: Cannot analyze text in images
  • Encrypted/binary code: Cannot analyze compiled or encrypted content
  • Runtime behavior: Static analysis only, no dynamic execution
  • Offline SC4: Without network access to api.osv.dev, SC4 uses a small static fallback list

Research Background

Based on research from "Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale" (Liu et al., 2026):

  • Dataset: 42,447 skills from major marketplaces; 31,132 were analyzed for the following rates
  • Vulnerable: 26.1% of the analyzed subset contain at least one vulnerability
  • High-severity: 5.2% of the analyzed subset show likely malicious intent
  • Key finding: Skills with executable scripts are 2.12x more likely to be vulnerable

Python API Integration

from skillspector import graph

# Invoke the LangGraph workflow
result = graph.invoke({
    "input_path": "/path/to/skill",
    "output_format": "json",   # terminal, json, markdown, or sarif
    "use_llm": True,           # False for static-only analysis
})

# Access results
print(f"Risk Score: {result['risk_score']}/100")
print(f"Severity: {result['risk_severity']}")
print(f"Recommendation: {result['risk_recommendation']}")

for finding in result["filtered_findings"]:
    print(f"[{finding['severity']}] {finding['rule_id']}: {finding['message']}")

License

Apache License 2.0 - see LICENSE for details.

Contributing

Contributions are welcome! Please read our contributing guidelines and submit pull requests.

Support

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

17 total
  1. SkillSpector v2.12.0v2.12.0Sep 23, 2026107 downloads

    # SkillSpector v2.12.0 Release status: candidate; publication pending. ## Summary SkillSpector 2.12.0 adds an opt-in CLI gate for any active finding, a configurable static-analysis allowance, OpenCode integrations, Gemini 3.5 Flash registry guidance, interactive scan progress, and TP4 analysis of executable Markdown fences. It also strengthens local input and report handling, expands detection of reflective Python access and shipped bytecode, retries transient provider failures, and distinguishes missing references from ambiguous MCP installation blockers. The release includes discovery, completeness, finding-identity, SARIF output, raw forge-file input, performance, and false-positive fixes described below. The latest additions include sanitized LLM provenance, SC10 dependency-source analysis, GitHub tree-directory inputs, opt-in compact prompt numbering, expanded model-budget metadata, fail-closed recursive reporting, and occurrence-specific JSON/SARIF columns. Further detection and coverage fixes distinguish narrow passive-image cases from active or unknown opaque content without treating uninspected bytes as safe. The candidate also includes the previously pending Markdown

  2. SkillSpector v2.11.2v2.11.2Sep 9, 20263.3K downloads

    # SkillSpector v2.11.2 Released: 2026-09-10 ## Summary SkillSpector 2.11.2 fixes fatal reference-accounting errors and several false static-parser limits triggered by ordinary documentation. This patch also preserves incomplete-analysis reporting when a runtime-selected executable prevents exact command reconstruction. ## Highlights - Complete reference accounting when Markdown labels and destinations identify the same artifact, or when several referenced artifacts appear on one source line. - Avoid false parser limits for simple runtime parameters, inline skill invocations, PowerShell member access, and long quoted prose. - Keep runtime-selected `printf` and wrapper paths marked as partially inspected. ## Added - None. ## Changed - Record reference-coverage completion once per source line. ## Fixed - Deduplicate reference-coverage records for the same source file, line, and target, preventing fatal `unaccounted_work` errors from duplicate Markdown references. - Account for distinct reference targets on the same source line without creating conflicting completion records. - Distinguish simple runtime parameters from command substitutions and complex parameter expansions

  3. SkillSpector v2.11.1v2.11.1Sep 7, 2026304 downloads

    # SkillSpector v2.11.1 Released: 2026-09-07 ## Summary SkillSpector 2.11.1 raises the default aggregate scan deadline from 60 seconds to 600 seconds and makes it configurable, preventing larger valid scans from timing out under the previous one-minute workflow budget. This patch release also includes security-analysis correctness fixes merged since 2.11.0. ## Highlights - Give direct, recursive, transitive, and multi-skill scans a 600-second aggregate workflow deadline by default. - Allow operators to set a positive finite deadline with `SKILLSPECTOR_MAX_WORKFLOW_SECONDS` while retaining the safe default for invalid values. - Preserve security finding classifications during scan-view and report deduplication, and strengthen detection of concealed instructions. ## Added - Add `SKILLSPECTOR_MAX_WORKFLOW_SECONDS` as an optional environment setting for the aggregate workflow deadline. ## Changed - Increase the default end-to-end workflow and transitive traversal deadline from 60 seconds to 600 seconds. - Apply the configured deadline consistently across direct CLI, recursive, transitive, and multi-skill analysis paths. - Enforce `SKILLSPECTOR_MAX_LLM_CONCURRENCY` across all co

  4. SkillSpector v2.11.0v2.11.0Aug 28, 202615.9K downloads

    # SkillSpector v2.11.0 Released: 2026-08-28 ## Summary SkillSpector 2.11.0 expands supply-chain coverage to installed npm dependency versions, adds bounded analysis of bundled lifecycle hooks and project permission grants, and introduces optional LLM sampling controls. It also improves provider guidance and fallback routing, supports secure file traversal in restricted Linux environments, and removes two reported MP3/P6 false positives without weakening directive detection. ## Highlights - Resolve exact direct and transitive npm versions from `package-lock.json` and `npm-shrinkwrap.json` before vulnerability analysis. - Add BH1–BH3 findings for bundled lifecycle hooks, directly proven remote transfer of sensitive content, and broad or ignored project permission modes. - Add `SKILLSPECTOR_TEMPERATURE` and `SKILLSPECTOR_SEED` controls while preserving provider defaults when they are unset. - Avoid MP3 and P6 findings for the reported nominal state-coverage and CSS print-rule descriptions while retaining actionable reset and disclosure directives. ## Added - Parse npm lockfile versions 1, 2, and 3 under the existing dependency-analysis resource bounds, including nested installs

  5. SkillSpector v2.10.0v2.10.0Aug 26, 2026224 downloads

    # SkillSpector v2.10.0 Released: 2026-08-26 ## Summary SkillSpector 2.10.0 expands security coverage across concealed artifacts, referenced skills, structured skill bundles, and external model selection. It also makes incomplete analysis harder to mistake for a clean result, adds localized LLM finding text, and exposes the highest reported issue severity for downstream policy gates. ## Highlights - Inspect hidden files and ZIP-compatible nested artifacts under cumulative safety bounds, with HIGH SC9 findings for concealed executables and provenance-preserving virtual paths. - Add opt-in transitive reference scanning with bounded traversal, source provenance, shared budgets, and fail-closed completeness reporting. - Recognize AISOP/AISP structured skill bundles and render report-only workflow summaries without affecting risk scores. - Add EA5 detection for external model or provider selection, including silent coding-CLI account switches and top-level model pins. - Add `SKILLSPECTOR_OUTPUT_LANGUAGE` for human-readable LLM finding text and `risk_assessment.max_issue_severity` for machine-readable policy gates. ## Added - Add bounded local inspection of hidden and nested ZIP, D

Code frequency

additions and deletions
+36.8K-36.8KWeek of 2026-03-15: +180 linesWeek of 2026-03-15: -2 linesWeek of 2026-03-22: +0 linesWeek of 2026-03-22: -0 linesWeek of 2026-03-29: +0 linesWeek of 2026-03-29: -0 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +0 linesWeek of 2026-04-12: -0 linesWeek of 2026-04-19: +0 linesWeek of 2026-04-19: -0 linesWeek of 2026-04-26: +0 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +26,597 linesWeek of 2026-05-10: -164 linesWeek of 2026-05-17: +25 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +194 linesWeek of 2026-05-31: -71 linesWeek of 2026-06-07: +1,208 linesWeek of 2026-06-07: -510 linesWeek of 2026-06-14: +17,425 linesWeek of 2026-06-14: -5,179 linesWeek of 2026-06-21: +17,445 linesWeek of 2026-06-21: -5,327 linesWeek of 2026-06-28: +1,703 linesWeek of 2026-06-28: -302 linesWeek of 2026-07-05: +9,442 linesWeek of 2026-07-05: -9,213 linesWeek of 2026-07-12: +507 linesWeek of 2026-07-12: -112 linesWeek of 2026-07-19: +1,885 linesWeek of 2026-07-19: -644 linesWeek of 2026-07-26: +9,065 linesWeek of 2026-07-26: -1,576 linesWeek of 2026-08-02: +5,688 linesWeek of 2026-08-02: -434 linesWeek of 2026-08-09: +6,309 linesWeek of 2026-08-09: -340 linesWeek of 2026-08-16: +36,794 linesWeek of 2026-08-16: -3,270 linesWeek of 2026-08-23: +9,725 linesWeek of 2026-08-23: -390 linesWeek of 2026-08-30: +7,846 linesWeek of 2026-08-30: -508 linesWeek of 2026-09-06: +4,824 linesWeek of 2026-09-06: -315 linesWeek of 2026-09-13: +18,501 linesWeek of 2026-09-13: -1,315 linesWeek of 2026-09-20: +19,853 linesWeek of 2026-09-20: -576 linesMar 15, 2026Sep 20, 2026
+195.2K lines added, -30.2K removed over the last year.

Commits per week

last 52 weeks
710Week of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 3 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 1 commitsWeek of 2026-05-17: 1 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 2 commitsWeek of 2026-06-07: 15 commitsWeek of 2026-06-14: 44 commitsWeek of 2026-06-21: 71 commitsWeek of 2026-06-28: 18 commitsWeek of 2026-07-05: 13 commitsWeek of 2026-07-12: 5 commitsWeek of 2026-07-19: 11 commitsWeek of 2026-07-26: 10 commitsWeek of 2026-08-02: 13 commitsWeek of 2026-08-09: 18 commitsWeek of 2026-08-16: 26 commitsWeek of 2026-08-23: 24 commitsWeek of 2026-08-30: 11 commitsWeek of 2026-09-06: 23 commitsWeek of 2026-09-13: 45 commitsWeek of 2026-09-20: 25 commitsSep 28, 2025Sep 20, 2026
379 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 1 commitsSun 2:00 — 1 commitsSun 3:00 — 1 commitsSun 4:00 — 1 commitsSun 5:00 — 0 commitsSun 6:00 — 2 commitsSun 7:00 — 2 commitsSun 8:00 — 3 commitsSun 9:00 — 2 commitsSun 10:00 — 1 commitsSun 11:00 — 3 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 1 commitsSun 15:00 — 1 commitsSun 16:00 — 0 commitsSun 17:00 — 1 commitsSun 18:00 — 2 commitsSun 19:00 — 0 commitsSun 20:00 — 1 commitsSun 21:00 — 0 commitsSun 22:00 — 2 commitsSun 23:00 — 3 commitsMon 0:00 — 0 commitsMon 1:00 — 1 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 1 commitsMon 9:00 — 2 commitsMon 10:00 — 0 commitsMon 11:00 — 11 commitsMon 12:00 — 0 commitsMon 13:00 — 4 commitsMon 14:00 — 5 commitsMon 15:00 — 4 commitsMon 16:00 — 2 commitsMon 17:00 — 9 commitsMon 18:00 — 1 commitsMon 19:00 — 7 commitsMon 20:00 — 2 commitsMon 21:00 — 2 commitsMon 22:00 — 1 commitsMon 23:00 — 6 commitsTue 0:00 — 1 commitsTue 1:00 — 6 commitsTue 2:00 — 8 commitsTue 3:00 — 0 commitsTue 4:00 — 1 commitsTue 5:00 — 0 commitsTue 6:00 — 1 commitsTue 7:00 — 1 commitsTue 8:00 — 2 commitsTue 9:00 — 4 commitsTue 10:00 — 2 commitsTue 11:00 — 3 commitsTue 12:00 — 2 commitsTue 13:00 — 3 commitsTue 14:00 — 3 commitsTue 15:00 — 7 commitsTue 16:00 — 3 commitsTue 17:00 — 6 commitsTue 18:00 — 2 commitsTue 19:00 — 5 commitsTue 20:00 — 1 commitsTue 21:00 — 8 commitsTue 22:00 — 3 commitsTue 23:00 — 1 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 3 commitsWed 3:00 — 1 commitsWed 4:00 — 1 commitsWed 5:00 — 0 commitsWed 6:00 — 8 commitsWed 7:00 — 1 commitsWed 8:00 — 2 commitsWed 9:00 — 0 commitsWed 10:00 — 2 commitsWed 11:00 — 0 commitsWed 12:00 — 6 commitsWed 13:00 — 4 commitsWed 14:00 — 7 commitsWed 15:00 — 6 commitsWed 16:00 — 1 commitsWed 17:00 — 6 commitsWed 18:00 — 1 commitsWed 19:00 — 0 commitsWed 20:00 — 6 commitsWed 21:00 — 1 commitsWed 22:00 — 4 commitsWed 23:00 — 3 commitsThu 0:00 — 3 commitsThu 1:00 — 2 commitsThu 2:00 — 2 commitsThu 3:00 — 1 commitsThu 4:00 — 6 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 3 commitsThu 8:00 — 1 commitsThu 9:00 — 0 commitsThu 10:00 — 5 commitsThu 11:00 — 3 commitsThu 12:00 — 1 commitsThu 13:00 — 1 commitsThu 14:00 — 0 commitsThu 15:00 — 4 commitsThu 16:00 — 5 commitsThu 17:00 — 0 commitsThu 18:00 — 3 commitsThu 19:00 — 3 commitsThu 20:00 — 4 commitsThu 21:00 — 2 commitsThu 22:00 — 3 commitsThu 23:00 — 2 commitsFri 0:00 — 2 commitsFri 1:00 — 7 commitsFri 2:00 — 3 commitsFri 3:00 — 4 commitsFri 4:00 — 0 commitsFri 5:00 — 2 commitsFri 6:00 — 4 commitsFri 7:00 — 5 commitsFri 8:00 — 1 commitsFri 9:00 — 2 commitsFri 10:00 — 8 commitsFri 11:00 — 1 commitsFri 12:00 — 2 commitsFri 13:00 — 3 commitsFri 14:00 — 0 commitsFri 15:00 — 7 commitsFri 16:00 — 2 commitsFri 17:00 — 3 commitsFri 18:00 — 0 commitsFri 19:00 — 1 commitsFri 20:00 — 1 commitsFri 21:00 — 1 commitsFri 22:00 — 5 commitsFri 23:00 — 3 commitsSat 0:00 — 7 commitsSat 1:00 — 4 commitsSat 2:00 — 3 commitsSat 3:00 — 0 commitsSat 4:00 — 2 commitsSat 5:00 — 1 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 4 commitsSat 11:00 — 1 commitsSat 12:00 — 1 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 1 commitsSat 17:00 — 1 commitsSat 18:00 — 0 commitsSat 19:00 — 1 commitsSat 20:00 — 1 commitsSat 21:00 — 8 commitsSat 22:00 — 0 commitsSat 23:00 — 4 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Oct 4, 2026monthly#19+3,495
Oct 3, 2026monthly#19+3,495
Oct 2, 2026monthly#18+3,423
Oct 1, 2026monthly#18+3,388
Sep 30, 2026monthly#16+3,454
Sep 29, 2026monthly#15+3,493
Jun 15, 2026daily#6+33
Jun 14, 2026daily#8+55
Jun 13, 2026daily#5+57
Jun 12, 2026daily#16+23
  • public-apis/public-apis

    A collective list of free APIs

    486.1K stars · Python

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    373.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    285.8K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    272.8K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • n8n-io/n8n

    Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

    206.7K stars · TypeScript