← All posts

Time-based detections are false-positive engines

Jul 25, 2026 false-positives api-security sqli scanning

Somewhere in almost every API scan report there is a finding that reads like this:

Blind SQL injection — POST /api/coupons/validate, parameter code. Payload ' || pg_sleep(5)-- returned in 5.2s versus a baseline of 0.2s.

It looks like proof. A stopwatch is objective, the delta is huge, and the payload is a real SQL sleep function. It is also the single most common false positive in API security scanning, and the reason is simple: you did not measure the database. You measured the internet.

What the clock actually measured

A response time is the sum of everything between your socket and the answer. On a modern API that list is long, and most of it has nothing to do with your payload:

Source of delay Typical size Fires when
Serverless cold start 0.3–8 s the first request to an idle function — i.e. your first payload
Autoscaler adding a pod 2–30 s traffic ramps, which a scanner does by definition
DB connection-pool exhaustion 1–10 s the scan opens more concurrent requests than the pool allows
Garbage-collection pause 0.1–2 s JVM/.NET services under sudden load
Rate limiter or WAF tarpit 1–30 s deliberately, in response to attack-shaped input
Upstream retry with backoff 1–15 s a gateway retries a 5xx your payload caused
Cache miss vs. hit 0.05–3 s your payload is a novel cache key — always
Cross-region call, noisy neighbour, shared CI runner 0.1–5 s whenever

Look at that list again with an attacker's payload in mind. A SQL-sleep string is long, weird, never seen before, and attack-shaped. It is guaranteed to miss the cache. It is unusually likely to trip a WAF that responds by slowing you down. If it causes a 500, an API gateway may retry it twice with backoff before answering. Every one of those produces exactly the signal the detector is looking for — slow response after malicious payload — with no vulnerability anywhere.

The timing detector cannot distinguish "the database slept for me" from "the platform had a bad five seconds". It was never able to. It just reports the delta.

The base-rate problem, with numbers

Take an endpoint with a p50 of 180 ms and a p99 of 6 s — completely ordinary for an API that talks to a third party or occasionally cold-starts. A timing detector with a 5-second threshold fires on anything in that top 1%.

Now run a normal scan: 40 parameters × 12 SQLi payloads × 3 injection points = 1,440 requests. At a 1% tail, roughly 14 requests cross the threshold by chance alone. If even a third of them happen to land on a sleep payload, that is four or five "confirmed blind SQL injections" in a single scan of a perfectly healthy service.

Meanwhile the real thing — an actually injectable parameter — is one finding. The signal-to-noise ratio is upside down, and it gets worse the more thorough the scan is. Timing detections punish coverage.

Why the retest ritual doesn't save it

Every scanner that ships timing detection knows this, so it adds a confirmation step: send the payload again, maybe with a different sleep duration, and check that the delay scales.

That helps with random noise. It does nothing about systematic noise, which is the kind you actually have:

  • A cold start repeats every time the function has been idle — including between your retries.
  • A tarpitting WAF delays every attack-shaped request, and delays the 10-second payload for at least 10 seconds too, so "the delay scaled with the sleep" is satisfied by a WAF that never touched a database.
  • Pool exhaustion persists for as long as the scan keeps up the request rate — which is the whole scan.

And it is expensive. Confirming with 5- and 10-second sleeps across a few dozen candidate parameters means a scan that spends minutes sitting still, doing nothing but waiting, to produce a verdict it still can't defend. That is the worst trade in scanning: slower and wronger.

Why API scanning is where this hurts most

Three things about APIs make timing especially tempting and especially unreliable:

  1. No reflection to key on. A web page echoes your input; a JSON API frequently returns {"valid": false} no matter what you send. With no in-band signal, timing looks like the only remaining channel — so tools reach for it exactly where they have the least corroboration.
  2. DEBUG=False is the norm. Production APIs don't return database error strings, which is the evidence an error-based detector needs. Blind classes dominate.
  3. API infrastructure is timing-hostile by design. Autoscaling, serverless, gateways, retries, rate limits, circuit breakers, and per-tenant throttling are all standard on the platforms APIs run on, and all of them inject multi-second variance.

So the environment with the least reliable clock is the one where the clock gets used most. That's the whole story of why this class of false positive is so common.

What to use instead

Every blind class has at least one oracle that doesn't depend on wall-clock time. These are the four worth building on.

1. Boolean differential. Send a condition that is true and one that is false, and compare the responses to each other and to a benign baseline:

code=abc' OR '1'='1   ->  200, 3 results
code=abc' OR '1'='2   ->  200, 0 results   <- matches the benign baseline
code=abc              ->  200, 0 results

The TRUE case stands out from both the FALSE case and the baseline. Nothing here can be produced by a slow network. Requiring FALSE to match the baseline is what rules out "any weird input changes the response".

2. Error-state differential. You don't need the error text, only the fact that the parser choked:

code=abc'         ->  500      <- broken quote: SQL syntax error
code=abc'--       ->  200      <- balanced injection: parses fine
code=abc!!##      ->  200      <- benign-but-weird control: NOT rejected

That triple is decisive. A 500 on the broken quote and a 200 on both the balanced payload and the junk control means the 500 came from SQL parsing, not from generic input validation. Any one of those three responses alone proves nothing — which is exactly why single-probe detectors misfire.

3. An evaluated-result oracle. For command injection, don't time a sleep; make the target compute something and hand it back:

name=x$(expr 731 \* 419)      ->  "x306289"    <- it ran the command
name=x`id`                    ->  "xuid=0(root)…"

The finding is the product, never the echoed expression. A reflected payload looks identical to a naive detector and completely different to this one.

4. An out-of-band callback. When there is genuinely no in-band channel — blind SSRF, blind XXE, Log4Shell, a jku fetch — make the target contact a listener you control. A callback that arrives is binary evidence: the server made a request it should not have made, and you have the log entry. No callback, no finding, and there is no tail-latency scenario that fabricates one.

What NewScan does

NewScan's blind SQL-injection detector is built on this specifically, and its docstring says so in one line: never uses timing. It runs a matched probe set and records a finding only on one of two false-positive-guarded differentials — the 500-state triple above, or the boolean differential where the FALSE case must match the benign baseline — and then re-confirms it on a repeat before the finding is written. Command injection is confirmed by a cross-platform arithmetic oracle plus a command-output marker. The genuinely blind classes are proven with an out-of-band collaborator, and with no collaborator configured those probes are never sent at all, so an offline scan cannot produce that false positive either.

The point isn't that these oracles are clever. It's that each one produces evidence a reader can check: a request, a response, and a difference that only the vulnerability explains. Every finding carries that evidence into the report, so an assessor can replay it.

Is timing ever legitimate?

Yes — narrowly. HTTP request smuggling detection is genuinely timing-based, because the desync is a timing artifact and there is no other in-band signal. A few DoS-adjacent checks are too.

The rule for those cases isn't "never measure time", it's:

  • Never auto-verify on a clock. Timing evidence justifies a suspected item at most.
  • Always carry a control. A benign payload of similar shape and length must not produce the same delay, in the same run, under the same conditions.
  • Say what you measured. "5.2s vs 0.2s baseline, single sample" is an honest note. "Confirmed blind SQL injection" is not.

A short checklist for evaluating any scanner

  1. Ask what the oracle is for each blind class. If the answer is "response time", ask what the control is.
  2. Ask whether findings are re-confirmed, and whether the confirmation is a different signal or the same one repeated.
  3. Point it at a healthy, cold, autoscaling staging environment and see how many criticals it invents.
  4. Check whether the report includes the request/response pair that proves each finding. If it can't show you, it doesn't know.

A finding you have to re-test by hand isn't a finding. It's homework — and time-based detections assign more of it than anything else in an API scan.