Dashboard

Can AI Find Security Vulnerabilities in Your Code?

Vendors published hard numbers on AI vulnerability discovery this week. The honest reading: strong on pattern-shaped classes, weak on business logic, and useless without a human triage step.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
3 September 20261 min read

Yes, with caveats that matter. As of September 2026 the published figures put frontier models at roughly 47% pass@1 on automated patching benchmarks and at genuine, verified discoveries in production codebases. That is far better than the sceptics expected two years ago and far worse than a replacement for a security review. The useful question is not whether AI finds vulnerabilities, it is which classes it finds, at what false positive cost, and what you have to do around it.

What the vendors published this week

Two releases in two days gave us unusually concrete numbers, and both came with third-party attribution rather than pure self-reporting.

Claim

Figure

Attributed to

CWE-Bench patching, Gemini 3.8 Flash Cyber

47.2% pass@1

Google

Correct patches vs leading commercial models

2.6 times more

Chrome Security team

Vulnerability recall, at lower cost

7.5% to 9.7% higher, 2.3x to 5.2x cheaper

Wiz

Time to find a critical vulnerability

Under 2 hours

Google Cloud Vulnerability Research

Cyber safeguard false positives

60% fewer than the previous generation

Anthropic

Sources: Google's Gemini 3.8 Flash and 3.8 Flash Cyber announcement and Anthropic's Claude Fable 5.1 and Mythos 5.1 release, both published this week. Note what is missing from every one of those rows: a false positive rate on real code. Recall without precision is easy to buy.

The classes AI is good at

Pattern-shaped vulnerabilities. The model has seen ten thousand examples of the shape, and your code either matches it or does not.

  • Injection of all kinds: SQL built by string concatenation, shell commands assembled from user input, template injection.

  • Missing authorisation checks on an endpoint that its siblings all have. Models are unusually good at noticing the odd one out in a set.

  • Unsafe deserialisation, path traversal, and the standard file-handling mistakes.

  • Secrets and credentials sitting in source, config or test fixtures.

  • Dependency issues where a known-vulnerable version is pinned, though a scanner does this faster and cheaper.

The classes it is bad at

Anything that requires knowing what your application is supposed to mean.

  • Business logic flaws. A discount that can be applied twice, a refund path that does not decrement inventory. There is nothing syntactically wrong with the code and the model has no way to know the intent.

  • Multi-step exploitation chains, where each individual step is legitimate and only the sequence is dangerous.

  • Race conditions and anything where timing rather than structure is the bug.

  • Vulnerabilities that live in the gap between two services, where neither codebase alone contains the flaw.

This maps onto the long-horizon reliability problem: chaining is exactly where models lose the thread, which we cover in what a long-horizon task is.

The false positive tax

Anthropic's 60% reduction in false positives is a bigger deal than it sounds, because false positives are what kill AI security tooling in practice. A scanner that reports 40 issues of which 3 are real does not save anyone time. It moves the work from finding bugs to dismissing reports, which is less interesting and just as slow.

Before you adopt anything here, run it on a codebase whose bugs you already know. Count real findings, count noise, and decide from that ratio rather than from a benchmark score. If the noise ratio is worse than one real finding in five reports, your reviewers will start rubber-stamping, and you have made things worse.

A workflow that actually works

  1. Scope the ask. Point the model at one module and one vulnerability class per pass. Broad requests produce broad, shallow, noisy output.

  2. Give it the threat model. Tell it which inputs are untrusted, which endpoints are public, and what the auth boundary is. Most false positives come from the model guessing at trust boundaries.

  3. Demand a proof sketch. Require, for every finding, the specific input that triggers it and the path from entry point to sink. Findings without a path are usually wrong and are cheap to discard.

  4. Triage with a human. Every time. This is not a gate you can automate away yet, and the vendors making the strongest claims are the ones gating the capability behind verification programmes.

  5. Fix separately. Have the model propose a patch in its own pass, and review that patch like any other change.

Practical companions: how to review AI-generated code before you ship it covers the review discipline, and how to catch an AI coding agent introducing a vulnerability covers the mirror-image problem of your own agent creating the bug.

The access question

The most capable versions of this are not generally available. Google's Flash Cyber goes to trusted defenders through its Fairwind Program, aimed at government bodies, critical infrastructure and software maintainers. Anthropic's Mythos 5.1 is behind a Cyber Verification Program. If you are a small team, you are using the general-availability model with tighter refusals, which is materially less capable at offensive-shaped tasks by design.

That is a reasonable trade, but it means benchmark numbers from the gated variants are not the numbers you will get. Plan against the general model. Broader context on the risk side is in our AI risks pillar, what red teaming means in AI, and whether AI agents are a cybersecurity risk.

Frequently asked questions

Can AI replace a penetration test?

No. A penetration test covers configuration, infrastructure, chained exploitation and human factors. Model-based code review covers code. They overlap less than the marketing suggests.

Is it safe to paste proprietary code into a model to scan it?

It depends entirely on the provider's retention terms. Check whether the endpoint you are using has zero data retention before you paste anything you would not email to a stranger.

What does pass@1 mean on a patching benchmark?

The share of problems the model solves correctly on its first attempt, with no retries. It is the strict version of the metric, which is why 47.2% is a respectable number rather than a weak one.

Should I run this on every pull request?

Only if your false positive rate is low enough that people still read the output. Start on a schedule, measure the noise, and move to per-PR when the signal justifies it.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.