Independent benchmark for whether MCP/source security scanners detect authorization-logic vulnerabilities.
Expanded result: 19 labeled cases, including 11 authorization-logic cases across 10 root-cause classes, 2 control bugs, and 6 clean negatives.
Evidence caveat: semgrep and bandit report 0/11 authorization-logic detections while catching 2/2 control bugs; because two authz cases share a root cause, the stricter class-level reading is 0/10.
authz logic 0/11
class caveat 0/10
controls detected
disposable CI
Read-only AI-app and MCP-style security triage scanner. It is a lead generator, not a precision gate.
Validation: 59 public AI repos, 1410 raw leads, 140 sampled for human adjudication, about 3-4 percent precision.
read only
human confirmed
SARIF receipts
Reference implementation that composes claims, citations, benchmarks, datasets, calibration, scope, and security leads into one receipt.
Absent checks report n/a. Receipts are reproducible, and limitations are stated plainly.
one receipt
transparent rollup
n/a is honest
Shared audit and sealing primitive for the verification toolchain.
Unkeyed seals are treated as integrity checks, not overstated as tamper-proof evidence.
audit chain
canonical JSON
replayable