SupercovBlog

Research

Claude Code security review skills, tested on 147 known vulnerabilities

Three popular skills, the built-in /security-review and Supercov with Jev, compared on accuracy, time and cost.

4 seconds. 1 cent. A brushed-aluminum stopwatch with a violet arc over a third of its dial.

We asked Claude Code to review six deliberately vulnerable apps for security issues, 147 known vulnerabilities in all. It did each review six ways: on its own, with three popular security skills, with the built-in /security-review, and with the Supercov plugin, which hands the analysis to Jev.

The skills didn’t make Claude Code better at finding vulnerabilities. The best of them scored about the same as no skill. Supercov’s plugin matched that accuracy at a fraction of the cost: 25 seconds and 17 cents a review, against about a minute and 55 to 65 cents for the skills. Jev’s own part took about 4 seconds and a cent.

Results

OptionF1FoundRightTimeCost
Sentry security-review skill0.5848%74%81 s$0.65
Supercov plugin, with Jev0.5552%59%25 s$0.17
No skill0.5576%44%57 s$0.58
OWASP security-check skill0.5574%43%62 s$0.58
everything-claude-code security-review skill0.5173%39%55 s$0.55
/security-review, built in0.3624%70%298 s$2.81

Found is the share of the 147 known vulnerabilities a review reported in the right place. Right is the share of its findings that were real. F1 balances the two; higher is better. Time is the median per app. Cost is the API list price of one review, including Jev for Supercov.

F1 against cost per review. Supercov with Jev scores 0.55 at $0.17 in 25 seconds. The skills and no skill score 0.51 to 0.58 at $0.55 to $0.65 in about a minute. The built-in /security-review scores 0.36 at $2.81 in 298 seconds.
Accuracy against cost per review, with the median time per app.

How Jev does security analysis for a cent

Jev is the model behind Supercov’s security checks. It doesn’t write a review. It answers twelve fixed questions about each file, such as whether request data reaches a database query or whether a route checks who is calling, and gives each answer a confidence. Supercov turns the answers into findings with a file, a class of weakness and, where it can, a line.

Answering questions instead of writing prose makes Jev fast and cheap. In this test it assessed each app in about 4 seconds, for about a cent; Jev charges only for what it reads. Most of the plugin’s 25 seconds and 17 cents is Claude reading the report and writing it up. Run on its own, npx supercov security is just the Jev part.

The questions are the same every run, so results stay steady. On these six apps, Jev’s report this time scored about the same as in our 72-repository benchmark in September: F1 0.57 against 0.58.

Skills trade coverage for confidence

Every option caught most of the injection and cross-site scripting bugs. They differed on issues that take judgment:

Found, of labelledSentrySupercovNo skillOWASPECCBuilt in
Injection and XSS, 21201921212117
Missing authorization, 20861514153
Secrets in source, 1579121074
Insecure configuration, 14491111102

Sentry’s skill and /security-review report only what they are confident about. Their findings are more often right, but they skip most of these judgment calls.

/security-review is built for pull requests

/security-review checks the changes on your branch and holds back anything it isn’t sure is exploitable. That keeps pull request reviews quiet, but it doesn’t make an audit. On one app it found command injection, template injection and path traversal, then dropped them all because the routes needed an internal token.

Claude Code sometimes runs /security-review on its own when you ask for a security review, so check which one ran.

Which one to use

  • On every change and pull request: Supercov’s plugin, at 25 seconds and 17 cents a review. /supercov:security patch checks only what the change introduced.
  • In CI: npx supercov security runs Jev alone, in seconds, for about a cent per app.
  • For a full audit: Claude Code with no skill finds the most. About half of its findings are false alarms, so check each one.
  • For fewer false alarms: Sentry’s skill. Its findings were right 74% of the time.

How we ran it

  • Six small apps from the test half of RealVuln: three Python, two JavaScript and one TypeScript. Four are well-known training apps; two were written by coding agents for the benchmark.
  • Claude Code 2.1.280 with claude-sonnet-5-5, one run per option and app. Each run started from a fresh copy, with no other settings or plugins, and could not edit files.
  • Each skill came from a pinned commit and was called by its slash command with the same request. /security-review and /supercov:security take no request, so they ran on their own.
  • In the same session, Claude then listed what it had reported as JSON, so every option was scored the same way. Jev’s time is its step inside the plugin’s session; its cost comes from the tokens it read.
  • A finding counts when it names the right file, a line within 10 of the labelled one, and an accepted CWE or one from the same RealVuln family. A finding for a whole file counts once. A real issue that RealVuln didn’t label counts as wrong.
  • One run is noisy. Rerunning the same setup moved one app’s F1 from 0.59 to 0.73, so treat differences of a few hundredths as ties.
  • Trail of Bits publishes many security skills, but none is a general review for web apps, so we left them out.

For comparison, RealVuln’s own Claude Code run on these apps, with Sonnet 5 and a longer audit prompt, scored 0.57.

Download every review, finding and script, or see the results and prompts as JSON.

Try it

Add the plugin in Claude Code:

/plugin marketplace add supercorp-ai/supercov
/plugin install supercov@supercov

Then run /supercov:security, or npx supercov security in any terminal. Jev needs a TYPESAFE_API_KEY from TypeSafe AI; security covers setup and what each check looks for.