FREE FOR MAINTAINERS · DELIVERED PRIVATELY

Help secure the code the world builds on.

WinnowSec puts AI agents to work for open-source maintainers. Every finding arrives with a reachable path, a blind re-check by a different model and a person's sign-off.

THE PROBLEM

Attackers can guess. Maintainers have to check.

An attacker's agent can try thousands of attacks; a miss costs it nothing, and one hit is enough. A maintainer has to check every report by hand, including the false ones, and each check takes time.

Attacker's agent0 attempts

Automated. Each red flash is one attempt; misses cost nothing, and one hit is enough.

Still probing, but the files it needs were already fixed.

Maintainer's inbox0 reports checked
0 waiting to be checked

By hand. Reports arrive faster than one person can check them, false alarms included.

Checked first. Only verified findings arrive, each with a path and a patch.

Illustrative. Switch to see what reaches the maintainer when the checking is done first: the real bug gets fixed before the attacker lands a hit.

THE GATES

Overwhelmed by your scanners' false positives? We can help.

Every scanner hit passes four gates before it reaches you. Anything that fails one is listed in the report, with the reason.

pipeline · illustrative flow RUNNING
01 SCAN
every hit
Rule-based SAST and pattern sets mark candidate lines. Nothing is reported yet.
nothing dropped silently
02 REACHABILITY
a real path?
From attacker-controlled input to this code, with nothing in between that stops it.
unreachable: listed as dismissed
03 VALIDATE
30
Re-reads every cited line, then confirms, downgrades or dismisses.
confirmed · vtm benchmark
04 BLIND RE-READ
27
A second agent gets the file and line, never the argument.
upheld · 2 refuted · 1 unreadable

VTM, a deliberately vulnerable Django app: 30 confirmed at validation; 27 upheld on blind re-read, 2 refuted, 1 judge answer unreadable (a tool failure, not an abstention). Reviewer and judge were the same model, so read 27 as optimistic; reviews sent to maintainers use a different judge model. One target, and an unsubtle one. The particle flow is illustrative.

NO SCANNERS YET?

Most open-source projects have none set up. We host open-source scanners, including OpenGrep, OSV-Scanner and DefectDojo, run them on your code and review what they find. Nothing to install on your side.

In a study of over a million npm and PyPI packages, 94–97% had no automated dependency updates. Zahan et al., 2023

THE DEFENDER'S WINDOW

Defenders can still go first.

Agents that find vulnerabilities are available to defenders now, before attackers use them at scale. OpenAI calls that gap the defender's window, and says it is closing.

What agents can find What attackers use at scale Review by hand With agent review
Illustrative, after the chart in OpenAI's The Defense Factory.

One fix travels

A fix merged upstream can protect every project built on it.

Both sides can read it

Open-source code is public to attackers too. The advantage is going first.

The decision stays with you

Every finding carries its evidence, tied to the file, line and commit. You decide what to fix.

WHY NOW

Four signals from the field.

What each number means for the people who maintain open source.

planted bugs found54
of those, patched43
real bugs, not planted18
18

Agents can already find real bugs

In DARPA's AI Cyber Challenge, autonomous systems analysed more than 54 million lines of code from open-source projects. Besides the bugs planted for the contest, they found 18 real ones, which are being reported to the projects' maintainers.

DARPA, AI Cyber Challenge final, Aug 2025

real issuevaries by toolfalse alarm
40–60%

About half of scanner alerts are false alarms

Vendors of mainstream static-analysis tools report that 40 to 60 of every 100 alerts turn out to be false. Someone still has to check each one.

Vendor-reported range, mainstream SAST tools

0 of 7

Unchecked AI reports cost maintainers time

In the final week of curl's bug bounty, 7 reports came in and none described a real vulnerability. The project ended the program in January 2026 to remove the incentive for AI-generated reports.

The Register, Jan 2026

1 weekend

Attackers already use agent swarms

In July 2026, an autonomous agent framework broke into Hugging Face over a weekend. It started from a malicious dataset and ran thousands of actions from a swarm of short-lived sandboxes. The response used AI too: an open-weight model analysed 17,000+ attack events in hours.

Hugging Face incident report, Jul 2026

WHO WE ARE

Security and applied AI researchers, for a safer digital world.

We are a team of cybersecurity and applied AI researchers, committed to making the digital world a safer place. We run every review ourselves, read every report before it goes out, and publish what we learn, including whether agents, deterministic tools and independent review really cut false positives.

START

Stronger foundations for the agent era.

It starts with one repository: free for public GitHub repositories, delivered privately. We publish what we learn along the way, including what went wrong.

  1. 01

    Send the link

    One public GitHub URL, by email.

  2. 02

    Confirm you maintain it

    A draft security advisory with our account added. About two minutes.

  3. 03

    Get the report

    Inside that advisory, read by a person, with a patch for each finding.

Next on the roadmap: private repositories and reviews on demand.

METHOD

How a review runs.

Deterministic tools gather context. Four agent stages investigate it. A separate blind re-check and a person stand between the pipeline and delivery.

  1. MapPinned commit, code graph, routes and access controls.
  2. DiscoverOpenGrep first, then a checklist, domain by domain.
  3. ValidateA reachable path, a blind re-check, a sandbox when needed.
  4. OwnerDelivered only to a verified maintainer, in a private advisory.
  5. FixA patch for each finding, checked against your commit.

Today, one turn per request, for public repositories. The goal: a turn whenever you need one, private code included.

EACH STAGE IN DETAIL
BEFORE ANY MODELdeterministic, runs once, same result every time
01

Checkout

Cloned at a pinned commit. Every finding links back to that exact commit.

02

Code graph

Symbols, calls and imports from the syntax tree, with zero model tokens. The agents query it to navigate.

03

OpenGrep

Rule-based SAST, run once in a sandbox and queried by every step. A scan with errors is reported as incomplete, never as clean.

YOUR SCANNERS, OR OURSnothing to set up on your side
SASTSemgrepSASTSonarQubeSCAOSV-Scanner SECRETSGitleaksCONTAINERSTrivyIaCCheckovSASTCodeQL
DefectDojomerges and de-duplicates
WinnowSec AIeSCRa verdict per finding

Already run scanners? Send the consolidated DefectDojo report. No stack yet? We host the open-source ones and DefectDojo, run them on your code and merge the results. Either way, WinnowSec AIeSCR gives each finding its own verdict: confirmed, rejected or inconclusive.

in userecommended next
AGENTSone agent per step; a step's output must parse before the next one starts
04

Recon

Behaviour, stack, routes and access controls.

05

Checks

The checklist, domain by domain. Every scanner hit ends as a finding or a dismissal with a reason.

06

Validate

Re-reads every cited line, then confirms, downgrades or dismisses.

07

Report

Counts computed in code. Permalinks built by the runner.

Deterministic tools the agents call:route mapauth-guard matrixdeployment inventorygit-history secret searchpattern scannerscitation check
BEFORE IT REACHES YOU
08

Blind re-check

A different model gets the file and line of each finding, never the argument. It runs separately, started by hand.

09

A person reads it

Including whether each step actually read code before it answered.

EXAMPLE: ONE HIT THROUGH THE GATES
example · taskManager/views.py (abridged) SCANNING
@login_requireddef ping(request):    host = request.POST['ip']    cmd = "ping -c 1 " + host    subprocess.Popen(cmd, shell=True)    return render(request, 'ok.html')
CLASSCommand injection · CWE-78
SOURCEPOST parameter ip, unvalidated
SINKsubprocess.Popen(shell=True)
GUARDAuthentication only. No allow-list.
RE-READSecond agent rebuilt the path from source alone.
✓ CONFIRMED · CRITICAL
THE RULES

How the agents are kept honest.

A reachable path

A dangerous call is reported only if attacker-controlled input can reach it with nothing in the way.

Re-read blind

A different model gets the file and line, never the argument, and has to rebuild the finding from the code.

14 =14

Counted in code

Totals and citations are computed by the runner. No model is asked to count.

Gates set in advance

Pass marks are set in advance and results are recorded regardless, including a change that shipped as the default without passing its gate.

A person signs off

Every report is read by a person before it is sent.

Open-weight models

Models with public weights, sized for a single GPU. Swapping the model doesn't touch the skills or the orchestration.

WHEN A STEP COUNTS AS DONE

A pipeline can finish
without doing the review.

It happened in this project: the scanner had flagged real code, yet the report said “no security findings”.

WHAT THE SCANNER SAW
hits found

Opengrep had flagged the repository. The code under review was real.

WHAT THE REPORT SAID
“No security findings
were identified.”
risk: low · exit 0
output cut off at the provider capFAILED
“I've now verified all the key paths”FAILED
one newline, finish_reason: stopFAILED
prose where JSON was declaredFAILED
a single raw JSON objectCOMPLETED

A step now counts as done only if the next step can read its output.

An agent's answer had been cut off midway. It wasn't empty and carried no error marker, so it counted as finished. Now every output has to pass the next step's parser. A parser can't tell whether any code was actually read, so a person checks that before anything is sent.

ISOLATION

Everything that touches your code runs inside limits.

Some findings are confirmed by reproducing them, and that happens only in the reproduction sandbox.

boundreview runnerSAST scanreproduction sandboxcode graph
memory●●●●
swap disabled●●●●
cpu●●●—
process count●●●●
single-file size●●●—
core dumps · /dev/shm●●●—
log volume capped●●●—
killed on timeout●●●●
seccomp · AppArmor●●●n/a
caps dropped · no-new-privs · read-only●●●n/a

— not bounded on that component · n/a does not apply to it. Bounds are checked against the live cgroup rather than the command line, because this project has shipped limits that were present in writing and absent in effect. Those tests are run by hand for now.

WHAT YOU GET

From scanner hits to findings you can act on.

Static analysis covers every file and gives the same answer twice, but it can't tell which hits are real. Agents investigate each one; you get the verdict, the reasoning and the evidence to check it.

WITHOUT IT
scanner
youevery line

A location for every hit, with no verdict and no fix. Thousands of lines on a large codebase.

WITH IT
scanner
agents
+ checklist
SQL injection · path + patchCONFIRMED
Command injection · path + patchCONFIRMED
Path traversal · input never reaches itDISMISSED
youshort list

A verdict, the reasoning and a patch for each finding that survives. Dismissed hits are listed with the reason.

SCANNER OUTPUT
[HIGH] SQL Injection   UserDAO.java:142
[HIGH] SQL Injection   UserDAO.java:158
[HIGH] SQL Injection   UserDAO.java:174
[HIGH] Path Traversal  FileCtrl.java:88
[MED]  Weak Hash       CryptoUtil.java:31
… and thousands more

Where, and nothing else.

ONE FINDING IN A REVIEW · example
CONFIRMEDCRITICAL · CWE-89
SQL injection · UserDAO.java:142
PATHparam filter (LoginController:88) → UserService.search() → UserDAO.query(). No sanitisation on the way. Public endpoint.
WHY IT'S REALThe value is concatenated into the query. The check at :131 only tests its length.
- String q = "SELECT … WHERE name='" + filter + "'";
+ PreparedStatement ps = con.prepareStatement(
+     "SELECT … WHERE name = ?");
+ ps.setString(1, filter);
CHECKSPatch applies to the reviewed commit (git apply --check) · permalink to the exact line · re-checked blind
severity with an exploit scenario evidence you can audit patch a proposed diff dismissals every ruled-out hit, and why
LIMITS

Known limits.

What nothing points to

Reviews start from scanner hits and a fixed checklist. A vulnerability neither leads to can be missed.

Authorization and business logic

The weakest area. Whether an access rule is wrong depends on what the application is meant to allow, which the code rarely says.

Cryptography and policy weaknesses

Published studies find agent triage least reliable here, Sifting the Noise included.

Text written to steer a model

Code under review can contain comments aimed at the agents. A person reads every report before it is sent.

Patches are proposals

A patch that applies cleanly can still fail to build, break tests or miss the bug. Review it like any other contribution.

Open-weight models

Less raw capability than frontier models, traded for control and predictable cost. If measurement shows that isn't good enough, the choice gets revisited.

RESULTS

Public summaries, listed only with consent.

A review is a map of open vulnerabilities, with file names and line numbers. The public summary shows none of that.

EXyour-org/your-projectEXAMPLE
Goreviewed 2026-09-02@ a1b2c3d
0 critical2 high 5 med3 low
Example. Real listings appear only after the maintainer agrees in writing.
01

The first review goes here.

Request a review

Public, if the maintainer agrees

  • Repository name and language
  • Commit and date reviewed
  • Counts by severity

Never public

  • File names, line numbers, symbols
  • Evidence, exploit paths, suggested patches
  • Anything, before the maintainer has read the report
  • Anything, without a written yes. No reply means no.
TARGETS

The bar we set for ourselves.

Targets, not results: goals the project set before its first review. Each is published here when it's measured, pass or fail.

≤ 20%
false positives after validation
Few enough false alarms that maintainers keep reading the reports.
NOT YET MEASURED
≈ 0%
real critical issues dismissed
A filter that hides real vulnerabilities is worse than no filter. This target can fail the project.
NOT YET MEASURED
≥ 70%
patches that apply and pass the project's tests
A patch that breaks the build is one more problem to triage.
NOT YET MEASURED

Maintain a public repository?

REQUEST

Start with your repository.

Free reviews for public GitHub repositories you maintain. Send a link, not code. We confirm you maintain it before any review begins, and the report comes back privately, inside a draft security advisory.

new request
TO—
REPO
SUBJECT
Open in mail app Copied

Prepared in your browser. Nothing is sent until you send it from your own mail app.

01

Send the link

One public GitHub URL. Don't send code, attachments or access tokens.

02

Confirm you maintain it

Open a draft security advisory and add our account to it. Only admins can create one, and it stays private.

03

The review runs

At a pinned commit. A different model re-checks each finding blind; a person reads the report before it's sent.

04

Private delivery

The report is posted inside that draft advisory. It isn't public.

STEP 02, IN DETAIL

Open a draft advisory and add us to it

About two minutes. You need admin access to the repository.

  1. Open the draft form. In your repository: Security and quality tab, then Advisories, then New draft security advisory.
    Open it for your repository →
  2. Save a draft. Fill in the fields marked with an asterisk. A title like WinnowSec review is enough; nothing is public while it's a draft. Then click Create draft security advisory.
  3. Add us as a collaborator. On the draft, under Collaborators on the right, add our account.
  4. Reply to our email. We check the invitation, and the review starts.

GitHub's own guides: Creating a repository security advisory · Adding a collaborator to an advisory

Not an admin? Commit the one-time token we email you to the default branch as .winnowsec-verify, and delete it once we confirm: the file is public. The report then comes encrypted to the address you wrote from.

✓

Read the public source

Cloned read-only at one commit, and deleted when the review closes.

✓

Reproduce in a sandbox

Only to confirm a finding, inside the bounded container shown on the Method page.

✕

Write to your repository

No commits, branches, issues or pull requests. The only thing we write is inside your private advisory.

✕

Publish without you

A public summary appears only after the maintainer has read the report and agreed in writing.

CLOSED UNREAD

Code you don't maintain

Without proof of ownership, nothing runs. There is no way around step 02.

CLOSED UNREAD

Forks and mirrors

A fork or mirror makes you admin of a copy, not maintainer of the original.

NEXT

Private repositories

Coming with the GitHub App on the roadmap below.

OUT OF SCOPE

Live systems

This reads and reproduces from source. It does not test running services.

WHERE THIS IS GOING

Toward security reviews on demand, for any repository you maintain.

NOW

Email intake

Public GitHub repositories, free. Ownership proved through a draft advisory, and every report read by a person.

NEXT · PLANNED

GitHub App

Private repositories, through an app that can only read file contents. GitHub has no read-only collaborator role on personal repositories, so the app replaces invites entirely.

LATER · PLANNED

Reviews on demand

Request a review of any repository you maintain, whenever you need one, including results from the scanners you already run through DefectDojo.

Prefer your own mail client?