Claude Code security review skills, tested on 147 known vulnerabilities
Three popular skills, the built-in /security-review and Supercov with Jev, compared on accuracy, time and cost.

We asked Claude Code to review six deliberately vulnerable apps for security
issues, 147 known vulnerabilities in all. It did each review six ways: on its
own, with three popular security skills, with the built-in /security-review,
and with the Supercov plugin, which hands the analysis to Jev.
The skills didn’t make Claude Code better at finding vulnerabilities. The best of them scored about the same as no skill. Supercov’s plugin matched that accuracy at a fraction of the cost: 25 seconds and 17 cents a review, against about a minute and 55 to 65 cents for the skills. Jev’s own part took about 4 seconds and a cent.
Results
| Option | F1 | Found | Right | Time | Cost |
|---|---|---|---|---|---|
| Sentry security-review skill | 0.58 | 48% | 74% | 81 s | $0.65 |
| Supercov plugin, with Jev | 0.55 | 52% | 59% | 25 s | $0.17 |
| No skill | 0.55 | 76% | 44% | 57 s | $0.58 |
| OWASP security-check skill | 0.55 | 74% | 43% | 62 s | $0.58 |
| everything-claude-code security-review skill | 0.51 | 73% | 39% | 55 s | $0.55 |
/security-review, built in | 0.36 | 24% | 70% | 298 s | $2.81 |
Found is the share of the 147 known vulnerabilities a review reported in the right place. Right is the share of its findings that were real. F1 balances the two; higher is better. Time is the median per app. Cost is the API list price of one review, including Jev for Supercov.

How Jev does security analysis for a cent
Jev is the model behind Supercov’s security checks. It doesn’t write a review. It answers twelve fixed questions about each file, such as whether request data reaches a database query or whether a route checks who is calling, and gives each answer a confidence. Supercov turns the answers into findings with a file, a class of weakness and, where it can, a line.
Answering questions instead of writing prose makes Jev fast and cheap. In this
test it assessed each app in about 4 seconds, for about a cent; Jev charges only
for what it reads. Most of the plugin’s 25 seconds and 17 cents is Claude
reading the report and writing it up. Run on its own, npx supercov security
is just the Jev part.
The questions are the same every run, so results stay steady. On these six apps, Jev’s report this time scored about the same as in our 72-repository benchmark in September: F1 0.57 against 0.58.
Skills trade coverage for confidence
Every option caught most of the injection and cross-site scripting bugs. They differed on issues that take judgment:
| Found, of labelled | Sentry | Supercov | No skill | OWASP | ECC | Built in |
|---|---|---|---|---|---|---|
| Injection and XSS, 21 | 20 | 19 | 21 | 21 | 21 | 17 |
| Missing authorization, 20 | 8 | 6 | 15 | 14 | 15 | 3 |
| Secrets in source, 15 | 7 | 9 | 12 | 10 | 7 | 4 |
| Insecure configuration, 14 | 4 | 9 | 11 | 11 | 10 | 2 |
Sentry’s skill and /security-review report only what they are confident
about. Their findings are more often right, but they skip most of these
judgment calls.
/security-review is built for pull requests
/security-review checks the changes on your branch and holds back anything it
isn’t sure is exploitable. That keeps pull request reviews quiet, but it doesn’t
make an audit. On one app it found command injection, template injection and
path traversal, then dropped them all because the routes needed an internal
token.
Claude Code sometimes runs /security-review on its own when you ask for a
security review, so check which one ran.
Which one to use
- On every change and pull request: Supercov’s plugin, at 25 seconds and 17
cents a review.
/supercov:security patchchecks only what the change introduced. - In CI:
npx supercov securityruns Jev alone, in seconds, for about a cent per app. - For a full audit: Claude Code with no skill finds the most. About half of its findings are false alarms, so check each one.
- For fewer false alarms: Sentry’s skill. Its findings were right 74% of the time.
How we ran it
- Six small apps from the test half of RealVuln: three Python, two JavaScript and one TypeScript. Four are well-known training apps; two were written by coding agents for the benchmark.
- Claude Code 2.1.280 with
claude-sonnet-5-5, one run per option and app. Each run started from a fresh copy, with no other settings or plugins, and could not edit files. - Each skill came from a pinned commit and was called by its slash command with
the same request.
/security-reviewand/supercov:securitytake no request, so they ran on their own. - In the same session, Claude then listed what it had reported as JSON, so every option was scored the same way. Jev’s time is its step inside the plugin’s session; its cost comes from the tokens it read.
- A finding counts when it names the right file, a line within 10 of the labelled one, and an accepted CWE or one from the same RealVuln family. A finding for a whole file counts once. A real issue that RealVuln didn’t label counts as wrong.
- One run is noisy. Rerunning the same setup moved one app’s F1 from 0.59 to 0.73, so treat differences of a few hundredths as ties.
- Trail of Bits publishes many security skills, but none is a general review for web apps, so we left them out.
For comparison, RealVuln’s own Claude Code run on these apps, with Sonnet 5 and a longer audit prompt, scored 0.57.
Download every review, finding and script, or see the results and prompts as JSON.
Try it
Add the plugin in Claude Code:
/plugin marketplace add supercorp-ai/supercov
/plugin install supercov@supercov
Then run /supercov:security, or npx supercov security in any terminal. Jev
needs a TYPESAFE_API_KEY from TypeSafe AI;
security covers setup and what each check looks for.