CodeRiskTools

AI Change Firewall + Secret Scanner — Public Benchmark v1

Public Benchmark v1 · real local CLIs · independently replayed

Measured behavior, not a black-box promise.

CodeRiskTools was exercised on 240 deterministic synthetic changes with frozen source bytes, pinned CodeRiskTools versions, label-blind execution, machine-readable raw results and independent offline replay.

What the audited results demonstrate

120/120

Shared secret/config cases scored correctly by CodeRiskTools: 60 true positives, 60 true negatives, 0 false positives and 0 false negatives on the eligible synthetic set.

80/80

Intent-policy outcomes matched exactly: every case matched the expected allow/block decision and expected Firewall Rule ID set.

40/40

Agent workflow outcomes matched: Scanner + Firewall status, integrity binding and per-engine minimum finding counts reconciled with zero workflow mismatches.

Why this is meaningful: the benchmark invokes the real offline command-line tools, keeps labels outside subprocess inputs, treats runtime errors separately, redacts findings, and recomputes scores from published case-level results. Independent replay produced byte-identical scored outputs.

Public Benchmark v1 · frozen synthetic evidence

Verified benchmark evidence — selected tools

Selected tools are shown in descending predeclared eligible-case coverage for the shared 120-case secret/config track. Accuracy is shown only on each tool’s eligible denominator; NOT_APPLICABLE is never counted as a failure or zero. This is not a complete leaderboard: one omitted full-coverage comparator tied CodeRiskTools at 120/120, so no unique-winner claim is made.

100%
CodeRiskTools 4.0.0-stage17-baselineAI Change Firewall + Secret Scanner Engine · frozen source identity · MIT
120/120 eligible100% shared-track coverage

Confusion matrix60 TP · 60 TN · 0 FP · 0 FN
Eligible coverage120 / 120 cases
Source-lock file SHA-256b39965cbb98714cba4b6275d3fd8b7d3bcc41fc8db4dbdb934bc5f997ad6d2ac
Exact source/artifact identity58c610caf28ad0348f5046003a7d2a77dae6bdc0b078959a1235877e72b95f02
Separate intent-policy track80 / 80 exact outcomes
Separate Agent E2E track40 / 40 exact outcomes
83%
TruffleHog 3.95.9Truffle Security · AGPL-3.0
100/100 eligible100/120 cases eligible · 20 N/A

Confusion matrix50 TP · 50 TN · 0 FP · 0 FN
Eligible coverage100 / 120 cases
Outside declared scope20 NOT_APPLICABLE
Other CodeRiskTools tracksNot comparable
N/A
Semgrep CE 1.169.0Semgrep · LGPL-2.1
Not scoredNo official scope-equivalent offline ruleset

Eligible denominator0 cases
Displayed rankUnranked
ReasonNo pinned official ruleset with equivalent declared scope
InterpretationNo loss or zero score assigned

Interpretation boundary: this selected evidence display is not a complete leaderboard and is not a broad security, production-efficacy, speed or product-quality ranking. One omitted full-coverage comparator tied CodeRiskTools at 120/120; its complete result and identity remain in the downloadable evidence pack at RESULTS.md, results/shared_scored_results.json and tools.lock.json, and no unique-winner claim is made. Separate CodeRiskTools workflow tracks are not added to the shared score. Read methodology and limitations.

CodeRiskTools standalone synthetic detection track

Audited CodeRiskTools score: 120/120 points across 120 eligible synthetic secret and configuration detection cases: 60 true positives, 60 true negatives, 0 false positives and 0 false negatives in this frozen corpus.

120/120

Secret Scanner Engine
Shared synthetic detection cases

80/80

AI Change Firewall
Intent-policy workflow cases

40/40

AI Change Firewall Agent E2E
Agent workflow cases

CodeRiskTools track Eligible cases Audited synthetic result
Secret Scanner Engine 120 120/120 · 60 TP · 60 TN · 0 FP · 0 FN
AI Change Firewall intent policy 80 80/80 workflow-policy matches
AI Change Firewall Agent E2E 40 40/40 workflow-policy matches

Interpretation: these are separate CodeRiskTools tracks with different scopes. They are not summed into a combined accuracy or superiority score. Results describe this versioned synthetic corpus only and do not establish production efficacy or guarantee that a repository is secure.

Evidence you can inspect

  • 240 versioned cases with expected outputs and synthetic provenance.
  • Exact tool versions, archive hashes, source lock and commands.
  • Normalized redacted raw results plus independently recomputable scored results.
  • Zero benchmark runtime errors and zero timeouts in the frozen run.
  • Independent replay in an offline network namespace.

Expected reproducibility-pack ZIP SHA-256: 9e3bc04bc2d18f8d003b489d257aa69234abc564cc7df77fa27e018f7d2fb0eb. Copy this value from this page when running the verifier; do not trust a hash supplied only inside the download.

Methodology and denominator rules

  • Frozen inputs: 240 deterministic synthetic cases, corpus manifest, source lock and exact tool identities are SHA-256 pinned.
  • Label-blind execution: subprocesses receive an opaque execution identity, track name and neutral input.diff; expected labels and public case IDs remain outside the subprocess boundary.
  • Fair eligibility: each tool is scored only on declared scope-equivalent cases. NOT_APPLICABLE is excluded from the denominator; ERROR and TIMEOUT are never clean negatives.
  • Separate tracks: shared secret/config detection, intent-bound policy and Agent E2E are scored independently. No combined Scanner + Firewall accuracy is reported.
  • Independent replay: raw results were rescored offline; all three scored outputs were reproduced byte-for-byte, including a fail-closed E2E mutation probe.

Download the pack for exact case bytes, tool/source locks, commands, raw and scored JSON, full limitations, checksums and verifier.

The honest boundary

This is strong evidence for the exact published synthetic corpus, not a claim of universal or production efficacy. It does not measure production prevalence, long-term alert burden, adversarial adaptation or every repository type. It does not establish broad competitive superiority. CodeRiskTools is not presented as a replacement for SAST, SCA, container security or an enterprise AppSec platform.

The report–diff binding verifies integrity and consistency of exact bytes. It is not a signature, MAC or authenticity proof. Scanner detection, Firewall policy enforcement and Agent E2E remain separate evidence tracks; combined_accuracy_claim=false.

Exit mobile version