Public Benchmark v1 · real local CLIs · independently replayed
Measured behavior, not a black-box promise.
CodeRiskTools was exercised on 240 deterministic synthetic changes with frozen source bytes, pinned CodeRiskTools versions, label-blind execution, machine-readable raw results and independent offline replay.
What the audited results demonstrate
Shared secret/config cases scored correctly by CodeRiskTools: 60 true positives, 60 true negatives, 0 false positives and 0 false negatives on the eligible synthetic set.
Intent-policy outcomes matched exactly: every case matched the expected allow/block decision and expected Firewall Rule ID set.
Agent workflow outcomes matched: Scanner + Firewall status, integrity binding and per-engine minimum finding counts reconciled with zero workflow mismatches.
Why this is meaningful: the benchmark invokes the real offline command-line tools, keeps labels outside subprocess inputs, treats runtime errors separately, redacts findings, and recomputes scores from published case-level results. Independent replay produced byte-identical scored outputs.
Public Benchmark v1 · frozen synthetic evidence
Verified benchmark evidence — selected tools
Selected tools are shown in descending predeclared eligible-case coverage for the shared 120-case secret/config track. Accuracy is shown only on each tool’s eligible denominator; NOT_APPLICABLE is never counted as a failure or zero. This is not a complete leaderboard: one omitted full-coverage comparator tied CodeRiskTools at 120/120, so no unique-winner claim is made.
100%
CodeRiskTools 4.0.0-stage17-baselineAI Change Firewall + Secret Scanner Engine · frozen source identity · MIT
120/120 eligible100% shared-track coverage
83%
TruffleHog 3.95.9Truffle Security · AGPL-3.0
100/100 eligible100/120 cases eligible · 20 N/A
N/A
Semgrep CE 1.169.0Semgrep · LGPL-2.1
Not scoredNo official scope-equivalent offline ruleset
Interpretation boundary: this selected evidence display is not a complete leaderboard and is not a broad security, production-efficacy, speed or product-quality ranking. One omitted full-coverage comparator tied CodeRiskTools at 120/120; its complete result and identity remain in the downloadable evidence pack at RESULTS.md, results/shared_scored_results.json and tools.lock.json, and no unique-winner claim is made. Separate CodeRiskTools workflow tracks are not added to the shared score. Read methodology and limitations.
CodeRiskTools standalone synthetic detection track
Audited CodeRiskTools score: 120/120 points across 120 eligible synthetic secret and configuration detection cases: 60 true positives, 60 true negatives, 0 false positives and 0 false negatives in this frozen corpus.
Secret Scanner Engine
Shared synthetic detection cases
AI Change Firewall
Intent-policy workflow cases
AI Change Firewall Agent E2E
Agent workflow cases
| CodeRiskTools track | Eligible cases | Audited synthetic result |
|---|---|---|
| Secret Scanner Engine | 120 | 120/120 · 60 TP · 60 TN · 0 FP · 0 FN |
| AI Change Firewall intent policy | 80 | 80/80 workflow-policy matches |
| AI Change Firewall Agent E2E | 40 | 40/40 workflow-policy matches |
Interpretation: these are separate CodeRiskTools tracks with different scopes. They are not summed into a combined accuracy or superiority score. Results describe this versioned synthetic corpus only and do not establish production efficacy or guarantee that a repository is secure.
Evidence you can inspect
- 240 versioned cases with expected outputs and synthetic provenance.
- Exact tool versions, archive hashes, source lock and commands.
- Normalized redacted raw results plus independently recomputable scored results.
- Zero benchmark runtime errors and zero timeouts in the frozen run.
- Independent replay in an offline network namespace.
Expected reproducibility-pack ZIP SHA-256: 9e3bc04bc2d18f8d003b489d257aa69234abc564cc7df77fa27e018f7d2fb0eb. Copy this value from this page when running the verifier; do not trust a hash supplied only inside the download.
Methodology and denominator rules
- Frozen inputs: 240 deterministic synthetic cases, corpus manifest, source lock and exact tool identities are SHA-256 pinned.
- Label-blind execution: subprocesses receive an opaque execution identity, track name and neutral
input.diff; expected labels and public case IDs remain outside the subprocess boundary. - Fair eligibility: each tool is scored only on declared scope-equivalent cases.
NOT_APPLICABLEis excluded from the denominator;ERRORandTIMEOUTare never clean negatives. - Separate tracks: shared secret/config detection, intent-bound policy and Agent E2E are scored independently. No combined Scanner + Firewall accuracy is reported.
- Independent replay: raw results were rescored offline; all three scored outputs were reproduced byte-for-byte, including a fail-closed E2E mutation probe.
Download the pack for exact case bytes, tool/source locks, commands, raw and scored JSON, full limitations, checksums and verifier.
The honest boundary
This is strong evidence for the exact published synthetic corpus, not a claim of universal or production efficacy. It does not measure production prevalence, long-term alert burden, adversarial adaptation or every repository type. It does not establish broad competitive superiority. CodeRiskTools is not presented as a replacement for SAST, SCA, container security or an enterprise AppSec platform.
The report–diff binding verifies integrity and consistency of exact bytes. It is not a signature, MAC or authenticity proof. Scanner detection, Firewall policy enforcement and Agent E2E remain separate evidence tracks; combined_accuracy_claim=false.