FREE FOR MAINTAINERS · DELIVERED PRIVATELY

Help secure the code the world builds on.

WinnowSec puts AI agents to work for open-source maintainers. Every finding arrives with a reachable path, a blind re-check by a different model and a person's sign-off.

THE PROBLEM

Attackers can guess. Maintainers have to check.

An attacker's agent can try thousands of attacks; a miss costs it nothing, and one hit is enough. A maintainer has to check every report by hand, including the false ones, and each check takes time.

Attacker's agent0 attempts

Automated. Each red flash is one attempt; misses cost nothing, and one hit is enough.

Still probing, but the files it needs were already fixed.

Maintainer's inbox0 reports checked
0 waiting to be checked

By hand. Reports arrive faster than one person can check them, false alarms included.

Checked first. Only verified findings arrive, each with a path and a patch.

Illustrative. Switch to see what reaches the maintainer when the checking is done first: the real bug gets fixed before the attacker lands a hit.

THE GATES

Overwhelmed by your scanners' false positives? We can help.

Every scanner hit passes four gates before it reaches you. Anything that fails one is listed in the report, with the reason.

pipeline · illustrative flow RUNNING
01 SCAN
every hit
Rule-based SAST and pattern sets mark candidate lines. Nothing is reported yet.
nothing dropped silently
02 REACHABILITY
a real path?
From attacker-controlled input to this code, with nothing in between that stops it.
unreachable: listed as dismissed
03 VALIDATE
30
Re-reads every cited line, then confirms, downgrades or dismisses.
confirmed · vtm benchmark
04 BLIND RE-READ
27
A second agent gets the file and line, never the argument.
upheld · 2 refuted · 1 unreadable

VTM, a deliberately vulnerable Django app: 30 confirmed at validation; 27 upheld on blind re-read, 2 refuted, 1 judge answer unreadable (a tool failure, not an abstention). Reviewer and judge were the same model, so read 27 as optimistic; reviews sent to maintainers use a different judge model. One target, and an unsubtle one. The particle flow is illustrative.

NO SCANNERS YET?

Most open-source projects have none set up. We host open-source scanners, including OpenGrep, OSV-Scanner and DefectDojo, run them on your code and review what they find. Nothing to install on your side.

In a study of over a million npm and PyPI packages, 94–97% had no automated dependency updates. Zahan et al., 2023

THE DEFENDER'S WINDOW

Defenders can still go first.

Agents that find vulnerabilities are available to defenders now, before attackers use them at scale. OpenAI calls that gap the defender's window, and says it is closing.

What agents can find What attackers use at scale Review by hand With agent review
Illustrative, after the chart in OpenAI's The Defense Factory.

One fix travels

A fix merged upstream can protect every project built on it.

Both sides can read it

Open-source code is public to attackers too. The advantage is going first.

The decision stays with you

Every finding carries its evidence, tied to the file, line and commit. You decide what to fix.

WHY NOW

Four signals from the field.

What each number means for the people who maintain open source.

planted bugs found54
of those, patched43
real bugs, not planted18
18

Agents can already find real bugs

In DARPA's AI Cyber Challenge, autonomous systems analysed more than 54 million lines of code from open-source projects. Besides the bugs planted for the contest, they found 18 real ones, which are being reported to the projects' maintainers.

DARPA, AI Cyber Challenge final, Aug 2025

real issuevaries by toolfalse alarm
40–60%

About half of scanner alerts are false alarms

Vendors of mainstream static-analysis tools report that 40 to 60 of every 100 alerts turn out to be false. Someone still has to check each one.

Vendor-reported range, mainstream SAST tools

0 of 7

Unchecked AI reports cost maintainers time

In the final week of curl's bug bounty, 7 reports came in and none described a real vulnerability. The project ended the program in January 2026 to remove the incentive for AI-generated reports.

The Register, Jan 2026

1 weekend

Attackers already use agent swarms

In July 2026, an autonomous agent framework broke into Hugging Face over a weekend. It started from a malicious dataset and ran thousands of actions from a swarm of short-lived sandboxes. The response used AI too: an open-weight model analysed 17,000+ attack events in hours.

Hugging Face incident report, Jul 2026

WHO WE ARE

Security and applied AI researchers, for a safer digital world.

We are a team of cybersecurity and applied AI researchers, committed to making the digital world a safer place. We run every review ourselves, read every report before it goes out, and publish what we learn, including whether agents, deterministic tools and independent review really cut false positives.

START

Stronger foundations for the agent era.

It starts with one repository: free for public GitHub repositories, delivered privately. We publish what we learn along the way, including what went wrong.

  1. 01

    Send the link

    One public GitHub URL, by email.

  2. 02

    Confirm you maintain it

    A draft security advisory with our account added. About two minutes.

  3. 03

    Get the report

    Inside that advisory, read by a person, with a checked patch for each confirmed finding, or the reason there is none.

Next on the roadmap: private repositories and reviews on demand.

METHOD

How a review runs.

Deterministic tools gather context. Four agent stages investigate it, and the runner sends back whatever they left unread or unanswered. A blind re-check and a person stand between the pipeline and delivery.

  1. MapPinned commit, code graph, routes and access controls.
  2. DiscoverOpenGrep first, then a checklist, domain by domain.
  3. ValidateA reachable path, a blind re-check, a sandbox when needed.
  4. OwnerDelivered only to a verified maintainer, in a private advisory.
  5. FixA checked patch for each confirmed finding, or the reason there is none.

Today, one turn per request, for public repositories. The goal: a turn whenever you need one, private code included.

EACH STAGE IN DETAIL
BEFORE ANY MODELdeterministic, runs once, same result every time
01

Checkout

Cloned at a pinned commit. Every finding links back to that exact commit.

02

Code graph

Symbols, calls and imports from the syntax tree, with zero model tokens. The agents query it to navigate.

03

OpenGrep

Rule-based SAST, run once in a sandbox and queried by every step. A scan with errors is reported as incomplete, never as clean.

YOUR SCANNERS, OR OURSnothing to set up on your side
SASTSemgrepSASTSonarQubeSCAOSV-Scanner SECRETSGitleaksCONTAINERSTrivyIaCCheckovSASTCodeQL
DefectDojomerges and de-duplicates
WinnowSec AIeSCRa verdict per finding

Already run scanners? Send the consolidated DefectDojo report. No stack yet? WinnowSec AIeSCR runs Semgrep, SonarQube and OSV-Scanner itself, OSV against an offline database, and builds the same report. Either way, each finding gets its own verdict: confirmed, rejected or inconclusive.

in userecommended next
AGENTSone agent per step; a step's output must parse before the next one starts
04

Recon

Behaviour, stack, routes and access controls.

05

Checks

The checklist, domain by domain. Every scanner hit ends as a finding or a dismissal with a reason.

06

Sweep

The runner lists the code no step opened and sends it to bounded readers, more of them for a larger repository. What they flag goes to Validate like any other candidate.

07

Validate

Re-reads every cited line, then confirms, downgrades or dismisses. Any scanner hit or candidate it leaves without a verdict is asked again; what still has none is marked open.

08

Report

Counts computed in code. Permalinks built by the runner. Every file you receive shows the same findings, and the runner checks that it does.

Deterministic tools the agents call:route mapauth-guard matrixdeployment inventorygit-history secret searchpattern scannerscitation checkcode-graph queriesscanner hit index
PER CONFIRMED FINDINGone bounded conversation each; the runner checks what it returns
09

Attack path

Traced from entry point to impact. The runner checks every hop against the code and reads the severity off a fixed table; the model never sets it, and nothing is removed.

10

Patch

A diff checked against your commit with git apply --check. When the finding was reproduced, the patch is re-run in the sandbox: the exploit, the build, a normal request, your tests. A committed secret gets rotation steps instead of a diff.

BEFORE IT REACHES YOU
11

Blind re-check

A different model gets the file and line of each finding, never the argument. It runs separately, started by hand.

12

A person reads it

Including whether each step actually read code before it answered.

EXAMPLE: ONE HIT THROUGH THE GATES
example · taskManager/views.py (abridged) SCANNING
@login_requireddef ping(request):    host = request.POST['ip']    cmd = "ping -c 1 " + host    subprocess.Popen(cmd, shell=True)    return render(request, 'ok.html')
CLASSCommand injection · CWE-78
SOURCEPOST parameter ip, unvalidated
SINKsubprocess.Popen(shell=True)
GUARDAuthentication only. No allow-list.
RE-READSecond agent rebuilt the path from source alone.
✓ CONFIRMED · CRITICAL
THE RULES

How the agents are kept honest.

A reachable path

A dangerous call is reported only if attacker-controlled input can reach it with nothing in the way.

Re-read blind

A different model gets the file and line, never the argument, and has to rebuild the finding from the code.

14 =14

Counted in code

Totals and citations are computed by the runner. No model is asked to count.

Gates set in advance

Pass marks are set in advance and results are recorded regardless, including a change that shipped as the default without passing its gate.

A person signs off

Every report is read by a person before it is sent.

Named models

Swapping the model doesn't touch the skills or the orchestration: it runs on open-weight models or hosted ones. The test runs on the Results page used a hosted frontier model. Every result we publish names the model behind it.

WHEN A STEP COUNTS AS DONE

A pipeline can finish
without doing the review.

It happened in this project: the scanner had flagged real code, yet the report said “no security findings”.

WHAT THE SCANNER SAW
hits found

Opengrep had flagged the repository. The code under review was real.

WHAT THE REPORT SAID
“No security findings
were identified.”
risk: low · exit 0
output cut off at the provider capFAILED
“I've now verified all the key paths”FAILED
one newline, finish_reason: stopFAILED
prose where JSON was declaredFAILED
a single raw JSON objectCOMPLETED

A step now counts as done only if the next step can read its output.

An agent's answer had been cut off midway. It wasn't empty and carried no error marker, so it counted as finished. Now every output has to pass the next step's parser. A parser can't tell whether any code was actually read, so a person checks that before anything is sent.

ISOLATION

Everything that touches your code runs inside limits.

Some findings are confirmed by reproducing them, and that happens only in the reproduction sandbox.

boundreview runnerSAST scanreproduction sandboxcode graph
memory●●●●
swap disabled●●●●
cpu●●●—
process count●●●●
single-file size●●●—
core dumps · /dev/shm●●●—
log volume capped●●●—
killed on timeout●●●●
seccomp · AppArmor●●●n/a
caps dropped · no-new-privs●●●n/a
read-only filesystem · checkout●●copyn/a

— not bounded on that component · n/a does not apply to it · copy: the sandbox works on a disposable copy of the checkout, which it may change. Bounds are read back from the live cgroup for the review runner, the reproduction sandbox and the code graph, because this project has shipped limits that were present in writing and absent in effect; the scan container's are checked in its command line. Those tests are run by hand for now.

WHAT YOU GET

From scanner hits to findings you can act on.

Static analysis covers every file and gives the same answer twice, but it can't tell which hits are real. Agents investigate each one; you get the verdict, the reasoning and the evidence to check it.

WITHOUT IT
scanner
youevery line

A location for every hit, with no verdict and no fix. Thousands of lines on a large codebase.

WITH IT
scanner
agents
+ checklist
SQL injection · path + patchCONFIRMED
Command injection · path + patchCONFIRMED
Path traversal · input never reaches itDISMISSED
youshort list

A verdict, the reasoning and a patch for each finding that survives. Dismissed hits are listed with the reason.

SCANNER OUTPUT
[HIGH] SQL Injection   UserDAO.java:142
[HIGH] SQL Injection   UserDAO.java:158
[HIGH] SQL Injection   UserDAO.java:174
[HIGH] Path Traversal  FileCtrl.java:88
[MED]  Weak Hash       CryptoUtil.java:31
… and thousands more

Where, and nothing else.

ONE FINDING IN A REVIEW · example
CONFIRMEDCRITICAL · CWE-89
SQL injection · UserDAO.java:142
PATHparam filter (LoginController:88) → UserService.search() → UserDAO.query(). No sanitisation on the way. Public endpoint.
WHY IT'S REALThe value is concatenated into the query. The check at :131 only tests its length.
- String q = "SELECT … WHERE name='" + filter + "'";
+ PreparedStatement ps = con.prepareStatement(
+     "SELECT … WHERE name = ?");
+ ps.setString(1, filter);
CHECKSPatch applies to the reviewed commit (git apply --check) · permalink to the exact line · re-checked blind
severity with an exploit scenario evidence you can audit patch a proposed diff dismissals every ruled-out hit, and why one set report, page, patches, SARIF and a developer edition, checked to agree
LIMITS

Known limits.

What nothing points to

Reviews start from scanner hits and a fixed checklist. A vulnerability neither leads to can be missed.

Large codebases, read in part

On the four largest test repositories the review opened 1 to 13% of the code, and two of them got no confirmed finding from it. The report states the share it read. The reading budget now grows with the repository; that change is not yet measured.

Authorization and business logic

The weakest area. Whether an access rule is wrong depends on what the application is meant to allow, which the code rarely says.

Cryptography and policy weaknesses

Published studies find agent triage least reliable here, Sifting the Noise included.

Text written to steer a model

Code under review can contain comments aimed at the agents. A person reads every report before it is sent.

Patches are proposals

A patch that applies cleanly can still fail to build, break tests or miss the bug. It is re-run only when the finding was reproduced in the sandbox; review it like any other contribution.

Model choice

The test runs on the Results page used a hosted frontier model. A self-hosted open-weight model keeps your code on our machines, with less raw capability. We tell you which one reviewed your code.

RESULTS

Public summaries, listed only with consent.

A review is a map of open vulnerabilities, with file names and line numbers. The public summary shows none of that.

EXyour-org/your-projectEXAMPLE
Goreviewed 2026-09-02@ a1b2c3d
0 critical2 high 5 med3 low
Example. Real listings appear only after the maintainer agrees in writing.
01

The first review goes here.

Request a review

Public, if the maintainer agrees

  • Repository name and language
  • Commit and date reviewed
  • Counts by severity

Never public

  • File names, line numbers, symbols
  • Evidence, exploit paths, suggested patches
  • Anything, before the maintainer has read the report
  • Anything, without a written yes. No reply means no.
TEST RUNS

Fourteen public repositories, before the first request.

Seven were written for security training, so their flaws are public already, and they are named below. Seven are production projects whose maintainers have not seen the results: they appear only in totals, and their findings go to the maintainers first. All ran in September 2026 on a hosted frontier model.

VTM · a Django training app with 32 known vulnerabilities
11scanners and the review
6scanners only
6review only: two missing access checks, broken object access in the API, a guessable reset token, a script inside a template, uploads served as live pages
9neither

Scanners: Semgrep, SonarQube and OpenGrep. Reached means a scanner row near the vulnerability whose rule describes the same weakness; 3 of their 17 were only partial. The review confirmed 17 of the 32, and 17 of its 19 scored confirmations match a known vulnerability. The known set is ours, revised four times, twice by an auditor who saw no review output. One run, one target, an unsubtle one.

4,995
scanner rows, each with its own verdict
Semgrep, SonarQube and OSV-Scanner across the 14 repositories: 55% rejected with a written reason, 40% confirmed, 6% left open and marked as open. How many rejections are wrong is not yet measured.
14 REPOSITORIES
43 / 88
review findings no scanner pointed at
Of the 88 findings the review confirmed, 43 sit where no scanner row or OpenGrep hit falls within five lines. That is the half these scanners did not report.
14 REPOSITORIES
1–13%
of the code read on the four largest
Small applications were read almost whole; on the largest, the review opened about 1%. The report states the share it read, and the reading budget now grows with the repository. That change is not yet measured.
THE WEAK PART
built for trainingstackscanner rowsrejectedreview confirmedno scanner nearbycode read
VTMDjango13216%22799%
NodeGoatNode.js35067%11878%
RailsGoatRails12550%16236%
SKEA DjangoDjango50%1186%
SKEA NodeNode.js7382%11100%
SKEA RailsRails13862%2037%
Handoutstemplates55248%00100%
7 production projects, combined—3,62055%35241–69%

Scanner rows: Semgrep, SonarQube and OSV-Scanner, each adjudicated on its own. No scanner nearby: no scanner row or OpenGrep hit within five lines of the finding. Code read: the share of the code the review opened, measured by the runner from its reads, not reported by the model.

TARGETS

The bar we set for ourselves.

Targets, not results: goals the project set before its first review. Each is published here when it's measured, pass or fail. The first readings come from VTM alone.

≤ 20%
false positives after validation
Few enough false alarms that maintainers keep reading the reports. VTM: 2 of 19 scored confirmations match no known vulnerability, 11%.
FIRST READING · 1 TARGET · MET
≈ 0%
real critical issues dismissed
A filter that hides real vulnerabilities is worse than no filter. VTM: the review dismissed 1 of 10 known critical issues and downgraded another; scanner rows confirmed both, so they reached the report. A third was missed by everything.
FIRST READING · NOT MET BY THE REVIEW ALONE
≥ 70%
patches that apply and pass the project's tests
A patch that breaks the build is one more problem to triage. The sandbox re-run that would measure this is built; no test run has used it yet.
NOT YET MEASURED

Maintain a public repository?

REQUEST

Start with your repository.

Free reviews for public GitHub repositories you maintain. Send a link, not code. We confirm you maintain it before any review begins, and the report comes back privately, inside a draft security advisory.

new request
TO—
REPO
SUBJECT
Open in mail app Copied

Prepared in your browser. Nothing is sent until you send it from your own mail app.

01

Send the link

One public GitHub URL. Don't send code, attachments or access tokens.

02

Confirm you maintain it

Open a draft security advisory and add our account to it. Only admins can create one, and it stays private.

03

The review runs

At a pinned commit. A different model re-checks each finding blind; a person reads the report before it's sent.

04

Private delivery

The report is posted inside that draft advisory. It isn't public.

STEP 02, IN DETAIL

Open a draft advisory and add us to it

About two minutes. You need admin access to the repository.

  1. Open the draft form. In your repository: Security and quality tab, then Advisories, then New draft security advisory.
    Open it for your repository →
  2. Save a draft. Fill in the fields marked with an asterisk. A title like WinnowSec review is enough; nothing is public while it's a draft. Then click Create draft security advisory.
  3. Add us as a collaborator. On the draft, under Collaborators on the right, add our account.
  4. Reply to our email. We check the invitation, and the review starts.

GitHub's own guides: Creating a repository security advisory · Adding a collaborator to an advisory

Not an admin? Commit the one-time token we email you to the default branch as .winnowsec-verify, and delete it once we confirm: the file is public. The report then comes encrypted to the address you wrote from.

✓

Read the public source

Cloned read-only at one commit, and deleted when the review closes.

✓

Reproduce in a sandbox

Only to confirm a finding, inside the bounded container shown on the Method page.

✕

Write to your repository

No commits, branches, issues or pull requests. The only thing we write is inside your private advisory.

✕

Publish without you

A public summary appears only after the maintainer has read the report and agreed in writing.

CLOSED UNREAD

Code you don't maintain

Without proof of ownership, nothing runs. There is no way around step 02.

CLOSED UNREAD

Forks and mirrors

A fork or mirror makes you admin of a copy, not maintainer of the original.

NEXT

Private repositories

Coming with the GitHub App on the roadmap below.

OUT OF SCOPE

Live systems

This reads and reproduces from source. It does not test running services.

WHERE THIS IS GOING

Toward security reviews on demand, for any repository you maintain.

NOW

Email intake

Public GitHub repositories, free. Ownership proved through a draft advisory, and every report read by a person.

NEXT · PLANNED

GitHub App

Private repositories, through an app that can only read file contents. GitHub has no read-only collaborator role on personal repositories, so the app replaces invites entirely.

LATER · PLANNED

Reviews on demand

Request a review of any repository you maintain, whenever you need one, including results from the scanners you already run through DefectDojo. Reviewing a single change instead of the whole repository is already built.

Prefer your own mail client?