# Claude Code security review skills, tested on 147 known vulnerabilities > We tested Claude Code with no skill, three security skills, /security-review and Supercov with Jev on 147 known vulnerabilities: accuracy, time and cost. Web page: https://supercov.com/blog/claude-code-security-review-skills · Published 2026-10-01 All Supercov pages for agents: https://supercov.com/llms.txt · Full text: https://supercov.com/llms-full.txt Three popular skills, the built-in /security-review and Supercov with Jev, compared on accuracy, time and cost. We asked Claude Code to review six deliberately vulnerable apps for security issues, 147 known vulnerabilities in all. It did each review six ways: on its own, with three popular security skills, with the built-in `/security-review`, and with the Supercov plugin, which hands the analysis to Jev. The skills didn't make Claude Code better at finding vulnerabilities. The best of them scored about the same as no skill. Supercov's plugin matched that accuracy at a fraction of the cost: 25 seconds and 17 cents a review, against about a minute and 55 to 65 cents for the skills. Jev's own part took about 4 seconds and a cent. ## Results | Option | F1 | Found | Right | Time | Cost | | --- | ---: | ---: | ---: | ---: | ---: | | [Sentry security-review skill](https://github.com/getsentry/skills/tree/main/skills/security-review) | 0.58 | 48% | 74% | 81 s | $0.65 | | **[Supercov plugin, with Jev](https://github.com/supercorp-ai/supercov/tree/main/plugins/supercov)** | **0.55** | **52%** | **59%** | **25 s** | **$0.17** | | No skill | 0.55 | 76% | 44% | 57 s | $0.58 | | [OWASP security-check skill](https://github.com/sergiodxa/agent-skills/tree/main/skills/owasp-security-check) | 0.55 | 74% | 43% | 62 s | $0.58 | | [everything-claude-code security-review skill](https://github.com/affaan-m/ECC/tree/main/skills/security-review) | 0.51 | 73% | 39% | 55 s | $0.55 | | `/security-review`, built in | 0.36 | 24% | 70% | 298 s | $2.81 | **Found** is the share of the 147 known vulnerabilities a review reported in the right place. **Right** is the share of its findings that were real. **F1** balances the two; higher is better. Time is the median per app. Cost is the API list price of one review, including Jev for Supercov. ![F1 against cost per review. Supercov with Jev scores 0.55 at $0.17 in 25 seconds. The skills and no skill score 0.51 to 0.58 at $0.55 to $0.65 in about a minute. The built-in /security-review scores 0.36 at $2.81 in 298 seconds.](https://supercov.com/assets/security-skills-cost-accuracy-1360.d385885e2d.webp) *Accuracy against cost per review, with the median time per app.* ## How Jev does security analysis for a cent [Jev](https://supercov.com/docs/jev.md) is the model behind Supercov's security checks. It doesn't write a review. It answers twelve fixed questions about each file, such as whether request data reaches a database query or whether a route checks who is calling, and gives each answer a confidence. Supercov turns the answers into findings with a file, a class of weakness and, where it can, a line. Answering questions instead of writing prose makes Jev fast and cheap. In this test it assessed each app in about 4 seconds, for about a cent; Jev charges only for what it reads. Most of the plugin's 25 seconds and 17 cents is Claude reading the report and writing it up. Run on its own, `npx supercov security` is just the Jev part. The questions are the same every run, so results stay steady. On these six apps, Jev's report this time scored about the same as in our [72-repository benchmark](https://supercov.com/docs/security-benchmark.md) in September: F1 0.57 against 0.58. ## Skills trade coverage for confidence Every option caught most of the injection and cross-site scripting bugs. They differed on issues that take judgment: | Found, of labelled | Sentry | Supercov | No skill | OWASP | ECC | Built in | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | Injection and XSS, 21 | 20 | 19 | 21 | 21 | 21 | 17 | | Missing authorization, 20 | 8 | 6 | 15 | 14 | 15 | 3 | | Secrets in source, 15 | 7 | 9 | 12 | 10 | 7 | 4 | | Insecure configuration, 14 | 4 | 9 | 11 | 11 | 10 | 2 | Sentry's skill and `/security-review` report only what they are confident about. Their findings are more often right, but they skip most of these judgment calls. ## /security-review is built for pull requests `/security-review` checks the changes on your branch and holds back anything it isn't sure is exploitable. That keeps pull request reviews quiet, but it doesn't make an audit. On one app it found command injection, template injection and path traversal, then dropped them all because the routes needed an internal token. Claude Code sometimes runs `/security-review` on its own when you ask for a security review, so check which one ran. ## Which one to use - **On every change and pull request:** Supercov's plugin, at 25 seconds and 17 cents a review. `/supercov:security patch` checks only what the change introduced. - **In CI:** `npx supercov security` runs Jev alone, in seconds, for about a cent per app. - **For a full audit:** Claude Code with no skill finds the most. About half of its findings are false alarms, so check each one. - **For fewer false alarms:** Sentry's skill. Its findings were right 74% of the time. ## How we ran it - Six small apps from the test half of [RealVuln](https://github.com/kolega-ai/Real-Vuln-Benchmark): three Python, two JavaScript and one TypeScript. Four are well-known training apps; two were written by coding agents for the benchmark. - Claude Code 2.1.280 with `claude-sonnet-5-5`, one run per option and app. Each run started from a fresh copy, with no other settings or plugins, and could not edit files. - Each skill came from a pinned commit and was called by its slash command with the same request. `/security-review` and `/supercov:security` take no request, so they ran on their own. - In the same session, Claude then listed what it had reported as JSON, so every option was scored the same way. Jev's time is its step inside the plugin's session; its cost comes from the tokens it read. - A finding counts when it names the right file, a line within 10 of the labelled one, and an accepted CWE or one from the same [RealVuln family](https://github.com/kolega-ai/Real-Vuln-Benchmark/blob/main/config/cwe-families.json). A finding for a whole file counts once. A real issue that RealVuln didn't label counts as wrong. - One run is noisy. Rerunning the same setup moved one app's F1 from 0.59 to 0.73, so treat differences of a few hundredths as ties. - Trail of Bits publishes many security skills, but none is a general review for web apps, so we left them out. For comparison, RealVuln's own Claude Code run on these apps, with Sonnet 5 and a longer audit prompt, scored 0.57. [Download every review, finding and script](https://supercov.com/downloads/supercov-claude-code-security-skills.zip), or see [the results and prompts as JSON](https://supercov.com/downloads/claude-code-security-skills-2026-10-01.json). ## Try it Add the plugin in Claude Code: ```text /plugin marketplace add supercorp-ai/supercov /plugin install supercov@supercov ``` Then run `/supercov:security`, or `npx supercov security` in any terminal. Jev needs a `TYPESAFE_API_KEY` from [TypeSafe AI](https://typesafe.ai); [security](https://supercov.com/docs/security.md) covers setup and what each check looks for.