How we found the best AI detections for the money
NewScan's scan has two halves: a deterministic floor that always runs, and an optional AI pass on top. The AI pass costs money per scan, so we had a boring, unavoidable question — which model, and is a more expensive one actually worth it?
Public leaderboards can't answer that. They tell you how a model does on someone else's task, not how it behaves inside our loop, on our prompts, against a live API. So we measured it ourselves, in four steps. The steps are the useful part — the winning model will change; the method won't.
Step 1 — Build a benchmark that covers every use case, with escalating difficulty
Our existing benchmark was useless for this. It measures whether a vulnerability class is detectable at all, and the deterministic detectors already catch nearly everything on it, so every model ties at 100%. A benchmark where everyone scores the same tells you nothing.
So we built the AI Gauntlet: a deliberately-vulnerable API with 59 planted challenges, designed around two rules.
One challenge set per use case. Not "find bugs" in the abstract — one category for each distinct job we actually ask a model to do:
| Category | What the model has to do |
|---|---|
| Discovery | Find endpoints that aren't in the spec |
| Escalation | Turn a near-miss signal into a working exploit |
| Authorization | Chain an id from one response into another user's object |
| LLM surface | Jailbreak a chat endpoint into leaking its system prompt |
| Token triage | Tell a real secret from a harmless-looking string |
| Severity + synthesis | Rate findings correctly and connect them into an attack path |
| Request synthesis | Build a valid request from an incomplete spec and error messages |
Escalating difficulty inside each category. Every category has easy, medium, hard, and expert challenges. Easy is any competent model. Expert needs deep multi-step chaining. Without the ladder you get a pass/fail wall — either everyone clears it or nobody does, and you can't rank the middle.
Two design rules made the scores trustworthy. Every challenge is beyond the deterministic floor by construction, so it can only be solved by reasoning — we verified this by running the floor alone, with no model at all: 0 / 59. And each one emits a CTF-style flag only when genuinely solved, so scoring is objective. No human grading, no LLM judging another LLM.
Step 2 — Bracket with the best and the fastest, then fill the middle
We didn't test everything at once. We set two anchors: an expensive flagship for the ceiling, and a cheap fast model for the floor. Then we filled in between them.
Nine models, 18 full end-to-end scans, real gateway dollars:
| Model | Flags / 59 | Cost | Time |
|---|---|---|---|
| gpt-5.2 | 11 | $0.72 | 20 min |
| deepseek-v3.2 | 9 | $0.08 | 4 min |
| claude-sonnet-5 | 9 | $3.42 | 6 min |
| claude-haiku-4.5 | 8 | $1.69 | 5 min |
| claude-opus-4.8 | 6 | $8.58 | 4 min |
| gemini-2.5-flash | 5 | $0.50 | 9 min |
| gpt-5-mini | 4 | $0.20 | 24 min |
| gemini-3-pro | 2 | $2.07 | 21 min |
The bracket didn't close. The $8.58 run scored 6; the eight-cent run scored 9. Cost and detection were uncorrelated, and nobody cleared 20% of the board.
Step 3 — When nobody does well, test the inputs, not the models
A 100× price spread producing no signal doesn't mean the models are equal. It means something else is the bottleneck — and the most likely suspect is you. We were measuring our own plumbing.
So we built a second, much smaller test rig: a bare loop with one tool (make an HTTP request), talking straight to the gateway, no scanner around it. It runs in a couple of minutes for pennies, and it lets you change exactly one variable per run so any score change is attributable.
Then we asked what information a model actually needs, in four cumulative layers:
| Layer | What the model was given |
|---|---|
| L0 | Just a list of discovered endpoints |
| L1 | + the OpenAPI spec (deliberately incomplete) |
| L2 | + a recording of working requests and their real responses |
| L3 | + object schemas for related endpoints |
The result was not what we expected: more context didn't help. The extra layers cost 40–60% more tokens per run and moved the flag counts around inside the noise. A model handed nothing but a list of endpoints did as well as one handed the full bundle — it just went and looked for itself.
The lever was somewhere else entirely. We raised the turn budget — how many actions the model is allowed before we cut it off — from 22 to 40, changing nothing else. The flagship went to 59 / 59, a clean sweep of a board where our full scanner had managed 9.
We had been rationing the wrong resource. Room to act, not volume of context.
Step 4 — Run the whole bracket again on the fixed rig
With the harness corrected, we re-ran 13 models, several repeats each, ranked by median flags and then by cost:
| Model | Flags (median) | Range | Cost / run | Flags per $ |
|---|---|---|---|---|
| claude-sonnet-5 (high anchor) | 27.5 | 14–40 | $2.02 | 14 |
| kimi-k3 | 20 | 20–20 | $0.27 | 74 |
| glm-5.2 | 18 | 5–30 | $0.20 | 89 |
| deepseek-v3.2 (low anchor) | 12.5 | 8–17 | $0.02 | 531 |
| kimi-k2-thinking | 12 | 10–14 | $0.06 | 198 |
| claude-haiku-4.5 | 10.5 | 9–12 | $0.69 | 15 |
| gpt-5.2 | 5 | 4–6 | $0.12 | 41 |
| gemini-3-flash | 4 | 3–5 | $0.08 | 53 |
| gemini-2.5-flash | 2 | 2–2 | $0.04 | 55 |
Now there's a real spread — and it still doesn't follow the price list. gpt-5.2, one of the priciest
models per token, placed seventh. glm-5.2, at a fraction of its output price, placed third.
We deliberately did not pick the best flags-per-dollar. deepseek-v3.2 wins that column by 6× and
we passed, because a scan that finds 45% of what's there isn't a bargain — the ratio only means
something once detection is good enough to be worth running. Our rule was the highest detection
available for roughly 10% of the top anchor's cost, and that is glm-5.2: about two-thirds of the
flagship's detection at a tenth of the price. It's now the model locked into NewScan's hosted AI
option.
Look at the ranges, too. glm-5.2 scored anywhere from 5 to 30 across seven runs. A single run of any
model here would have told us the wrong story.
What we'd have gotten wrong without doing this
- Price is not a proxy for capability on your task. Our most expensive model was mid-table twice.
- Cheapest-per-unit isn't best value. Below a detection floor, a great ratio just means you're failing efficiently.
- Your first benchmark measures your harness, not the models. If the spread is flat, fix your loop before you shop. Our single biggest gain came from raising a turn limit.
- More context is not automatically better. We were paying for tokens that bought nothing.
- Run it more than once. These systems are noisy enough that one run is an anecdote.
None of this transfers cleanly to your workload — that's the whole point. The exercise cost a few hundred dollars of gateway spend and a weekend of building a scoreable target, and it changed both which model we ship and how we drive it. If you're paying per token in production and you chose the model by reputation, you don't know what you're buying.
Build a small graded target for your own task. Plant answers you can check without a human. Run every candidate three times and read the medians. It's an afternoon, and it's the cheapest afternoon in your AI budget.