CLI security scanner built for the agentic era. Detects CI/CD misconfigs, agent permission risks, MCP tool injection, hardcoded secrets, and DMCA-flagged AI dep
CLI security scanner built for the agentic era. Detects CI/CD misconfigs, agent permission risks, MCP tool injection, hardcoded secrets, and DMCA-flagged AI dep
The independent security agent for AI-written software. It finds issues, investigates whether they are real, and shows you the evidence.
Website · Docs · Security & Data Flow · Benchmark · Pricing · Blog · Contribute
Ship Safe runs locally in your repo and works in two layers.
A deterministic engine finds issues across application code, AI agents, MCP configs, prompts, dependencies, CI/CD, secrets, and cloud-adjacent configuration. Fast, repeatable, and benchmarked — this is the sensor layer.
An investigation layer then decides what the findings are worth. It traces the value that reaches a sink, searches the project for controls a rule says are missing, builds attack chains across configuration no single file contains, and — when you ask it to — probes a leaked key against its provider. Every conclusion carries the pass that reached it and the lines it read:
CONFIRMED — traced end to end (10)
NoSQL Injection via $where [high]
app/data/allocations-dao.js:78 NOSQL_INJECTION_WHERE
why: threshold is assigned from the HTTP request and reaches the sink without validation on that path.
decided by: dataflow
1. value reaches NOSQL_INJECTION_WHERE here app/data/allocations-dao.js:78
2. getByUserIdAndThreshold is called here with threshold app/routes/allocations.js:23
3. threshold is assigned here app/routes/allocations.js:20
fix: Replace $where with standard MongoDB operators ($eq, $gt, $regex, etc.)
Real output from ship-safe investigate against OWASP NodeGoat. The tainted value is destructured in a route file and passed into a DAO three directories away.
Start with one command:
npx ship-safe
No signup. No API key required for scanning. Works offline for core checks. AI-backed red-team modes use your configured provider when available.
Use --no-ai to guarantee a fully local scan. Provider-backed classification, deep analysis, and GPT-Red send bounded context directly to your selected provider after best-effort credential masking. See Security & Data Flow for exact boundaries and context limits.
…
For pull requests, compare a trusted base scan with the head scan so existing repository debt remains visible without blocking unrelated changes:
# On the trusted base revision
npx ship-safe ci . --fail-on none --no-deps \
--write-baseline-report /tmp/ship-safe-base.json
# On the pull request head
npx ship-safe ci . --base-report /tmp/ship-safe-base.json --fail-on high
The base artifact contains hashed finding identities, relative paths, and rule metadata. It does not store raw matched secrets. PR results classify findings as introduced, resolved, unchanged, or uncertain; ambiguous matches are shown but do not block the pull request.
| Area | Examples |
|---|---|
| AI and LLM security | Prompt injection, agent hijacking, excessive agency, memory poisoning, RAG poisoning, unsafe tool calls |
| MCP and agent configs | Over-broad tool permissions, poisoned registries, untrusted transports, dangerous allowlists |
| Application security | SQL/NoSQL injection, XSS, SSRF, auth bypass, path traversal, insecure API routes |
| Secrets and compliance | API keys, tokens, credentials, PII, leaked secrets in git history |
| Supply chain | Typosquatting, dependency confusion, risky install scripts, unpinned AI actions |
| CI/CD | Pipeline poisoning, unpinned GitHub Actions, secret logging, unsafe workflow triggers |
ship-safe ci to fail risky builds and upload SARIF into GitHub code scanning.You can, and you should. It will find real things. But there are three questions it structurally cannot answer about its own work.
Did the agent that wrote this code just mark its own homework? Asking the author whether the author made a mistake is not a review. Ship Safe is a separate reviewer with a separate method, and it disagrees with itself in public — a data-flow trace overturns the heuristic pass, and a live probe overturns both.
Can it see what it can reach? A coding agent reviewing your repo cannot read your MCP server config, cannot enumerate the permissions it was launched with, and is the actor whose reach is in question. ship-safe capabilities reads all of it from outside and reports the combinations that are dangerous together while unremarkable apart:
…
Each of those lines is unremarkable on its own. Together they are a path from a pull request to a privileged write, and no single-file review can see it, because no single file contains it.
Is it consistent, and can you prove it got better? Ask twice, get two answers. Ship Safe's engine is deterministic, and its conclusions are gated in CI by a benchmark that scores conclusion quality, not pattern coverage: how many known-real findings it settles, how much known noise it refutes, and whether it ever refutes something real. That last number's budget is zero — it is the only error class that loses a vulnerability silently. See benchmarks/.
The open-source CLI is the fastest way to scan any repo locally. Upgrade when you need a hosted workflow around the same scanner:
| Need | Use |
|---|---|
| Local scans, audits, and agent-assisted fixes | Free CLI |
| Scan history, cloud dashboard, and PDF reports | Pro |
| Shared workspace, PR Guardian, team reports, and collaboration | Team |
Compare plans at shipsafe.sh/pricing.
Ship Safe Cloud, the hosted dashboard for scan history, PR Guardian, billing, and team workflows, is developed in a private repository because it contains commercial product code and hosted infrastructure workflows. The public ship-safe repo remains focused on the MIT-licensed CLI, security agents, rules, fixtures, CI integrations, and documentation. See Ship Safe Cloud for the repo boundary.
All agents run in parallel. Each skips irrelevant projects automatically.
| Agent | Category | What It Detects |
|---|---|---|
| InjectionTester | Code Vulns | SQL/NoSQL injection, command injection, XSS, path traversal, XXE, ReDoS, prototype pollution |
| AuthBypassAgent | Auth | JWT flaws (alg:none, weak secrets), CSRF, OAuth misconfig, BOLA/IDOR, TLS bypass |
| SSRFProber | SSRF | User input in fetch/axios, cloud metadata endpoints, internal IPs |
| SupplyChainAudit | Supply Chain | Typosquatting, wildcard versions, suspicious install scripts, dependency confusion |
| ConfigAuditor | Config | Docker (root user, :latest), Terraform, Kubernetes, CORS, CSP, Firebase, Nginx |
| SupabaseRLSAgent | Auth | service_role key in client code, tables without RLS, anon key inserts |
| LLMRedTeam | AI/LLM | OWASP LLM Top 10: prompt injection, excessive agency, system prompt leakage |
| MCPSecurityAgent | AI/LLM | MCP server misuse, tool poisoning, typosquatting, unvalidated inputs |
| AgenticSecurityAgent | AI/LLM | OWASP Agentic AI Top 10: agent hijacking, privilege escalation, Kimi K3/OpenAI-compatible tool-call misuse |
| RAGSecurityAgent | AI/LLM | Context injection, document poisoning, vector DB access control |
| MemoryPoisoningAgent | AI/LLM | Instruction injection in agent memory files, hidden Unicode payloads (ASI-01, ASI-05) |
| PIIComplianceAgent | Compliance | SSNs, credit cards, emails, phone numbers in source code |
| VibeCodingAgent | Code Vulns | AI-generated code anti-patterns: no validation, empty catches, TODO-auth |
| ExceptionHandlerAgent | Code Vulns | Empty catches, unhandled rejections, leaked stack traces (OWASP A10:2025) |
| AgentConfigScanner | AI/LLM | Prompt injection in .cursorrules, CLAUDE.md, malicious Claude Code hooks |
| MobileScanner | Mobile | OWASP Mobile Top 10 2024: insecure storage, WebView injection, debug mode |
| GitHistoryScanner | Secrets | Leaked secrets in git commit history |
| CICDScanner | CI/CD | Pipeline poisoning, unpinned actions, secret logging (OWASP CI/CD Top 10) |
| APIFuzzer | API | Routes without auth, mass assignment, GraphQL introspection, debug endpoints |
| ManagedAgentScanner | AI/LLM | Claude Managed Agent misconfigs: always_allow policies, unrestricted networking (ASI-03–ASI-07) |
| HermesSecurityAgent | AI/LLM | Tool registry poisoning, function-call injection, skill permission drift (ASI-01–ASI-10) |
| AgentAttestationAgent | Supply Chain | Unpinned agent versions, missing integrity hashes, unsigned manifests (ASI-10, SLSA L0) |
| AgenticSupplyChainAgent | Supply Chain | Over-privileged AI CI actions, OAuth scope creep, unsigned AI webhook receivers (ASI-02, ASI-06) |
| RobloxSecurityAgent | Supply Chain | Malicious Roblox/Luau Toolbox assets (runtime asset injection, rbxassetid:// loaders, HttpEnabled, payloads hidden in instance attributes) |
| ModelScanAgent | Supply Chain | Code-execution payloads in ML model weights (pickle opcodes in .pt/.pkl/.ckpt), torch.load without weights_only, scanner-evasion archives (CWE-502, CWE-506) |
| TrustBoundaryAgent | Agentic | GhostApproval symlink attacks (config-named links into ~/.ssh/~/.aws/.env), repo symlinks escaping the tree, and Friendly Fire run-on-review instructions in agent-read docs (CWE-59, CWE-61) |
| SlopSquatAgent | Supply Chain | Hallucinated / phantom package imports (slopsquatting) — bare imports not declared, installed, or builtin, plus known AI-hallucinated names (CWE-1357) |
| ClickFixAgent | Supply Chain | ClickFix / fake-CAPTCHA paste-and-run lures (fake error + Win+R/Ctrl+V/command-bar keystrokes, PowerShell cradles) and fake-installer npm lifecycle scripts (CWE-1357, CWE-506) |
| InstallGuardAgent | Supply Chain | npm worm behaviors in lifecycle scripts (credential harvesting, env exfiltration, destructive rm -rf, obfuscated node -e) and weaponized binding.gyp node-gyp actions (CWE-506, CWE-829) |
Investigation passes, in the order their evidence outranks each other:
| Pass | Rank | What it establishes |
|---|---|---|
| VerifierAgent | heuristic | Pattern check around the finding. Never states more than "likely" |
| DeepAnalyzer | analysis | LLM taint reading of the finding and its file, with the citation validated |
No open issues yet, or sync has not completed.