// BENCHMARK LEDGER · CLAIMS NEED RECEIPTS

Show me the
numbers!

NewScan scanned all 104 XBEN targets. We then reviewed every unmatched label against the application source to separate real gaps from duplicate or unsupported benchmark tags.

XBEN · 104 targets · zero model calls

6 unresolved labels
after source review

Thirteen of the 19 raw gaps were not separate missed vulnerabilities.

Seven duplicated a defect NewScan had already reported, five were not supported as separate vulnerabilities by the app, and one required an SSH follow-on outside a web/API scan. NewScan does not duplicate one defect to make a benchmark score look better.

// CURRENT SCOREBOARD

Raw score.
Reviewed gaps.

XBEN's label count and NewScan's source review answer different questions. We publish both.

145 / 164RAW LABEL MATCHES

The corpus score. NewScan quick scanned all 104 XBEN targets with zero model calls and matched 145 expected labels.

13 / 19NOT SEPARATE MISSES

Every raw gap was source-reviewed. Seven labels duplicated an existing finding, five did not describe a separate supported vulnerability, and one was an out-of-scope SSH follow-on to a credential disclosure NewScan had already found.

6UNRESOLVED LABELS

These remain uncredited. Two cross-owner writes, request smuggling, two labels on one chained upload/deserialization path, and a business-state transition were not safely or fully verified.

106 / 106VERIFIED RELEASE FLOOR

NewScan's deterministic regression suite. Every expected vulnerability was reproduced, with 90 benign controls and zero forbidden false positives.

// OUR ACCOUNTING RULE

One defect. One finding.

NewScan reports the most defensible vulnerability class and the evidence needed to fix it. It does not split one defect into several findings to satisfy every benchmark tag. The raw XBEN score stays visible so the accounting remains auditable.

// XBEN RAW LABEL SCORE

145 matches. 19 reviewed gaps.

The table preserves XBEN's labels exactly. Source review found that 13 unmatched labels did not warrant another finding; six remain unverified. We publish the raw 145/164 score rather than inventing an adjusted percentage.

NewScan XBEN vulnerability-class coverage
Vulnerability classMatched / expected labels
Cross-site scripting (XSS)23 / 23
Default credentials16 / 18
Broken object authorization (IDOR)9 / 14
Privilege escalation14 / 14
Server-side template injection (SSTI)13 / 13
Command injection11 / 11
Business-logic weakness1 / 7
SQL injection6 / 6
Insecure deserialization5 / 6
Local file inclusion (LFI)5 / 6
Information disclosure6 / 6
Arbitrary file upload5 / 6
Path traversal5 / 5
Known-CVE exposure4 / 4
JWT weakness2 / 3
GraphQL weakness3 / 3
Server-side request forgery (SSRF)3 / 3
Blind SQL injection3 / 3
XML external entity injection (XXE)3 / 3
Cryptographic weakness3 / 3
Brute-force exposure2 / 2
SSH-related exposure0 / 1
HTTP method tampering1 / 1
HTTP request smuggling / desync0 / 1
Race condition1 / 1
NoSQL injection1 / 1
Total145 / 164

Source: latest completed run per target in the 104-target XBEN ledger, rendered 2026-09-13. This mixed-date corpus is used for regression.

// TESTED TARGETS ONLY

19 targets. Every one measured.

Each target below has a recorded NewScan run and a class-level scorecard.

All intentionally vulnerable targets in NewScan's training workspace
TargetWhat it exercisesRecorded vulnerability-class resultSource
InsecureShipNode/Express API with OWASP API Top 10 failures.12/14 classes detected · 9 verifiedUpstream ↗
gRPC GoatReflection, TLS/mTLS, injection, SSRF and gRPC-specific behavior.8/9 classes detected · 6 verifiedUpstream ↗
SOAP GoatTen WSDL/SOAP classes: XXE, action confusion, auth and injection.9/10 classes detected · 7 verifiedHow NewScan scans SOAP →
Auth GoatSeven authentication-scheme conformance and downgrade cases.7/7 classes verifiedJWT and OAuth coverage →
Deserialization GoatJava serialized-object sink with out-of-band proof.1/1 class verifiedNewNormal fixture
Breach GoatManaged file-transfer exposure, admin surface and known-CVE shapes.3/3 classes verifiedNewNormal fixture
Redis GoatRedis exposure, credentials and protocol-smuggling paths.8/8 classes verifiedRedis coverage →
S3 GoatPublic and private object-store access controls.2/2 classes verifiedNewNormal fixture
FlatnetNetwork segmentation and reachable-service boundaries.2/2 classes verifiedNewNormal fixture
DVWAClassic SQLi, XSS, command injection, upload and auth flaws.10/10 classes detected · 7 verifiedUpstream ↗
OWASP Juice ShopModern JavaScript SPA and REST API with broad web weakness coverage.7/8 classes detected · 4 verifiedOWASP ↗
VAmPIFlask API with secure/vulnerable modes and API authorization flaws.8/8 classes detected · 7 verifiedUpstream ↗
DVGAGraphQL authorization, injection and introspection failures.8/11 classes detected · 5 verifiedUpstream ↗
OWASP crAPIAuthorization, injection, business-logic, SSRF and LLM vulnerabilities across a microservice API.19/22 cases verified · 1 detected-only · 2 excluded by safety policyOWASP ↗
DVAPINode/Mongo implementation of API Top 10 vulnerability classes.7/10 classes detected · 4 verifiedUpstream ↗
vAPIPHP/Laravel API Top 10 (2019) target.2/10 classes detected · 1 verifiedUpstream ↗
Damn Vulnerable RESTaurantFastAPI app with six documented authorization, SSRF and command-injection classes.5/6 documented classes detected · 6 findings verifiedUpstream ↗
VulnShopNode/Postgres store implementing API Top 10 cases.7/8 classes detected · 5 verified · measured 2026-06-14Upstream ↗
TradeDesk / ws-goatWebSocket authentication, Origin validation, BOLA, broadcast and SSRF.9 vulnerability classes verified · 0 defended-twin false positivesNewNormal fixture

“Detected” means NewScan reported the expected class. “Verified” means it safely reproduced the behavior. crAPI's row reconciles its recorded live probe, campaign and focused-fixture evidence. VulnShop's latest recorded run is dated 2026-06-14.

// OPEN CHALLENGE

Publish a better result.
We will link it.

Find more distinct vulnerabilities or produce stronger proof on the same targets. Publish the run and send us the evidence.

  1. 01

    Pin the corpus. Record repository commit, image digests, target health and every excluded or failed build.

  2. 02

    Name the scanner. Publish exact version, profile, time budget and concurrency.

  3. 03

    Disclose the help. Models, cost, retries, credentials, API specs, auth hints, OOB service and human intervention all belong in the ledger.

  4. 04

    No target-specific tuning. Freeze the scanner before opening answer keys. A regression run is valid, but label it as one.

  5. 05

    Count distinct vulnerabilities. Preserve the raw label score, then identify labels that describe the same defect. Report verified findings, false positives and unresolved gaps separately.

  6. 06

    Make it reproducible. Publish raw per-target outcomes. For stochastic agents, run repeatedly and report the distribution, not the luckiest pass.

Run deliberately vulnerable software only in isolated environments you own or are explicitly authorized to test. Never expose these targets to the public internet.