Solutions

Agentic Adversarial Exposure Validation

Your engineers just doubled their output. Your security review didn’t.
Sybil is the automated product security engineering layer for the AI-native SDLC. It discovers what changed, decides what actually needs testing, proves what’s exploitable, and opens a fix back into the pipeline. All of it runs continuously, at the speed your agents are already shipping code.
Code snippet from api/routes/invoices.py showing a change in access control: a line checking if current_user.id equals invoice.owner_id is removed, replaced by a check if current_user.org_id equals invoice.org_id returning serialized invoice, else abort 403. The change extends ownership check from user-level to org-level access. Notes explain this broadens authorization to all org members, tested against cross-org and cross-role access, confirms a low-privilege user can read other org data, and includes a scoped query and regression test ready to merge.
CHANGED FILE TREE
★ api/routes/invoices.py
api/routes/billing.py
models/invoice.py
tests/test_invoices.py
api/routes/auth.py
api/routes/invoices.py · PR #4128
✓ Reviewed by Sybil
87
- if current_user.org_id == invoice.org_id:
88
+ return metadata(invoice)
89
abort(403)
YOU CHANGED THIS
Ownership check widened from user-level to org-level access.
WHY IT MATTERS
Any member of an org can now reach every invoice in it: a new authorization boundary, not just a bigger blast radius.
WE TESTED
Cross-org and cross-role access paths against this endpoint and its three downstream callers.
HERE'S WHAT WAS REAL
Confirmed: a low-privilege user in Org A can enumerate invoice IDs and read Org B's data via a shared billing proxy.
HERE'S THE FIX
PR opened: scoped query + regression test attached, ready to merge.
Trusted by security & engineering teams — including at Carta
Notion logoCursor logoThinking Machines logoCarta logoBaseten logo
10
x
the rate AI coding agents can already out-produce human code review. That's the pace validation now has to match.
(Forrester, Agentic Development Security, 2026)
60%
of organizations Gartner expects to adopt structured exposure validation by 2029.
(Gartner Market Guide for Adversarial Exposure Validation)
~2:1000
AppSec engineers per developer at the median enterprise, and that's before we consider how agentic development is scaling their output.
(Industry AppSec hiring benchmark, 2026)
71
%
of AppSec job postings call for scanning tool skills. Only 24% call for remediation.
(Industry AppSec hiring benchmark, 2026)

It runs itself.

Discover, decide, test, and remediate: the same five steps, automatically, onevery change. Nothing sits in a backlog waiting for someone to get to it.
01

Discover

Map the attack surface on an ongoing basis, including services, endpoints, identities, and the paths between them, not once a quarter.
Screenshot showing a web server response with HTTP 200 OK status and related headers in a black interface.
/admin/polls/question/?q=zap
Expand icon
Response
Request
Response headers
HTTP/1.1 200 OK Server: nginx/1.14.2 Date: Fri, 09 Oct 2020 21:43:04 GMT Content-Type: text/html; charset=utf-8 Content-Length: 5088 Connection: keep-alive Expires: Fri, 09 Oct 2020 21:43:04 GMT Cache-Control: max-age=0, no-cache, no-store, must-revalidate, private Vary: Cookie X-Frame-Options: DENY X-Content-Type-Options: nosniff Set-Cookie: csrftoken=K4tvsNT4IAs8recg8m1bLf

02

Understand what changed

Every deploy is diffed against the last known-good state of the attack surface, not just the codebase.

Attack Plan list showing categories like Admin Access Control & Broken Auth, RBAC & Multi Tenant Authorization, all marked Active.

03

Decide what matters

Changes are risk-ranked so testing effort goes where exposure actually moved, not where a scanner happened to look.

Findings table listing vulnerabilities with severity, CVSS score, status, and test type.

04

Test & validate

Authenticated, application-aware testing that proves exploitability, not another theoretical finding to triage.

Security report showing SQL injection attack evidence on login endpoint with unusual database query patterns.

/api/v1/findings/4471/validate

Expand icon

Response

Request

Response headers

HTTP/1.1 200 OK Server: nginx/1.14.2 Date: Fri, 09 Oct 2020 21:52:03 GMT Content-Type: application/json Content-Length: 866 X-Finding-Status: confirmed X-False-Positive: false Cache-Control: no-store

05

Remediate

A validated finding ships back as a pull request into the existing pipeline, not a ticket in a queue nobody owns.

Timeline showing three completed steps: Discovered, Finding Reported on 8/5/2026, and Closed (Fixed).

/api/v1/findings/4471/remediate

Expand icon

Response

Request

Response headers

HTTP/1.1 200 OK Server: nginx/1.14.2 Date: Fri, 09 Oct 2020 22:03:41 GMT Content-Type: application/json Content-Length: 512 X-Remediation-Status: patched X-Retest-Result: pass Connection: keep-alive

What this is built on

Four commitments, and one we think the industry should adopt.

1

Validation velocity matches development velocity

Engineering output is compounding, increasingly agent-authored. Security validation has to compound with it too: continuous re-testing as code and attack surface change, not an annual engagement.
Industry velocity roughly doubled year over year, 2026
2

Continuous understanding of the attack surface

Know what exists, what changed, and where exposure moved. Most teams are flying without this today.
Only 30% of teams are confident in 90%+ visibility
3

Test and validate, not just scan

Sybil exercises the application and proves exploitability, rather than adding to a queue of theoretical findings someone else has to triage.
Gartner’s own definition of Adversarial Exposure Validation
4

Findings close the loop into engineering

Validation and remediation land inside the sprint, the PR, the pipeline, not a side queue that depends on someone chasing it down.
Detection tooling still outpaces remediation tooling roughly 3 to 1
5
The one we think should be industry standard

Coverage becomes a heat map, not a snapshot

Every validated result updates a live picture of where exposure concentrates and where it’s proven clean, so the next round of testing knows where to look first. That same picture doubles as the defensible record for your board, your auditors, and your next funding round’s security review. We think that’s what “coverage” should mean industry-wide, not an asset count.
Also satisfies the requirement for

We sit inside your CTEM program. We don’t ask you to build a new one.

01

Scoping

You define what the program protects and how success is measured. That stays a human, organizational decision.
Your call
02

Discovery

Sybil builds and maintains continuous visibility into the real attack surface, including what changed since the last deploy.
Sybil
03

Prioritization

Every change is understood against your repos and your own knowledge base: what actually changed, what it touches, and why it’s worth testing.
Sybil
04

Validation

Authenticated, application-aware testing proves which exposures are real. Existing pentest and bounty findings feed in here too.
Sybil
05

Mobilization

Validated findings become remediation PRs routed to the owning team, with retesting to confirm the fix actually holds.
Sybil
If pentesting or bug bounty is the easiest way to get this funded today, that’s fine. Sybil’s offensive security perspective matches that of traditional human-driven efforts like pentesting and bug bounty.
Checkmark icon
Penetration testing
Checkmark icon
Bug bounty / VDP
Checkmark icon
DAST
Checkmark icon
SAST
Checkmark icon
Penetration testing
Checkmark icon
Bug bounty / VDP
Checkmark icon
DAST
Checkmark icon
SAST
Checkmark icon
Penetration testing
Checkmark icon
Bug bounty / VDP
Checkmark icon
DAST
Checkmark icon
SAST
Checkmark icon
Penetration testing
Checkmark icon
Bug bounty / VDP
Checkmark icon
DAST
Checkmark icon
SAST
Checkmark icon
Penetration testing
Checkmark icon
Bug bounty / VDP
Checkmark icon
DAST
Checkmark icon
SAST
Checkmark icon
Penetration testing
Checkmark icon
Bug bounty / VDP
Checkmark icon
DAST
Checkmark icon
SAST
Is this you

Built for teams who need a real test, on their timeline.

AI-authored code, and growing
Outnumbered by engineers and agents
Existing AppSec budget & battle scars
Leader ready to design-partner the metrics
Small security team, big engineering org
Faq

Frequently Asked Questions

What is adversarial exposure validation (AEV)?

Adversarial exposure validation (AEV) is a security practice that tests whether vulnerabilities and misconfigurations can actually be exploited by an attacker. It uses realistic attack scenarios to confirm which exposures are real, so teams fix proven risk instead of theoretical findings. Gartner named the category as part of exposure management. Sybil applies it to live applications and proves exploitability before a finding reaches your team.

Penetration testing is an engagement that attempts to exploit a system at a set point in time. AEV applies the same attacker perspective continuously, re-validating exposure as code and the attack surface change. The difference is cadence and coverage, not finding quality. Automated penetration testing is one way to deliver AEV, and Sybil also shows which parts of the attack surface have been tested.

Breach and attack simulation (BAS) runs predefined attack scenarios to test whether security controls detect or block known techniques. In Gartner's framing, BAS is one of the technologies within adversarial exposure validation, alongside automated penetration testing. BAS shows how well your defenses respond, while penetration-style validation proves which exposures in your applications can be exploited. Sybil focuses on the application side.

Continuous threat exposure management (CTEM) is a Gartner framework with five stages: scoping, discovery, prioritization, validation, and mobilization. AEV is the validation stage, where exposures are tested to confirm they can be exploited. Those results feed prioritization and remediation, so teams address proven exposures first. Sybil works across discovery through mobilization inside an existing CTEM program.

Exposure validation reduces false positives by testing each finding against the live environment before anyone triages it. A finding that cannot be exploited is rejected instead of being added to a queue. This lets security teams spend their time on confirmed vulnerabilities and real risk. Sybil uses a multi-agent validation pipeline, so only exploitable findings reach your team.

Exposure validation should run continuously, or at least whenever code or the attack surface meaningfully changes. Annual or quarterly testing leaves gaps, because modern applications ship new code every day. Continuous validation re-tests what changed and tracks which areas have been covered. Sybil can run on a defined schedule or start automatically based on what changed in the application.

Let’s look at what changed in your last 100 PRs.

Bring us your current AppSec program, not just your attack surface. We want to talk about how product security scales from here: where Sybil already fits today, and where you help us define what’s next.