Skip to content

Current overview of DataFlowBench

Nobody is perfect. No analyzer answers a whole core correctly in any of the 13 languages, and of the 8 in the field only Bifrost makes no decisive mistake at all — bought by declining 392 of the 886 core assertions rather than guessing at them.

Each kernel poses one positive and one negative assertion per semantic template, and each language's kernel is its own population — cores of 28, 31, 34 or 35 templates, never pooled or ranked against each other, and never compared with the smaller cores of an earlier snapshot. Because every pair is balanced, an analyzer that always answers the same way — or answers blindly where it cannot see a construct — scores exactly half on the affected pairs: read correctness against that 50% blind baseline, and read approximation character from each vendor's TPR/FPR split in the details dialog. A blank analyzer on a kernel means no report in this freeze: no extractor, no frontend, or no adapter, which is coverage rather than a score.

DataFlowBench is SlopCop's own benchmark for semantic data flow: we build it to measure where Bifrost stands against 7 reference analyzers, and to keep ourselves honest while doing it. The methodology is deliberately analyzer-neutral — every number here is generated from digest-bound freeze evidence that we cannot edit after the fact, misses and crashes included — but the motivation is not neutral, and we would rather say so.

What v0.8.0 contains

This is the first freeze of the expanded recursive-composition kernels: 1,000 cases across 82 freshly rerun reports produce 4,036 outcomes. The release adds 556 cells and preserves 124 changed outcomes from the previous population as measured evidence, including CodeQL's C/C++ incompleteness and Bifrost's newly decisive JavaScript, TypeScript, C#, and Ruby cells. Incomplete and unsupported outcomes remain distinct.

Below the kernels sit two further populations: the modeling matrix, which asks whether each tool's own model-declaration surface is load-bearing, and the tool-native probes, which ask what each product decides with nothing supplied by us. Beside them sits the latency tier — descriptive per-case timings on a stated machine, kept strictly beside the correctness results. These timings were freshly measured on the current pins, with desktop activity, cache state and sandbox/elevated execution contexts retained; no uncontended-host or timing-parity claim is implied. All populations have separate denominators. There is no combined leaderboard, and benchmark-controlled results are never pooled with, or compared number-to-number against, tool-native ones.

On the direct-flow breadth baseline, Bifrost answered both assertions correctly in 13 of 13 languages.

Frozen case results, across three populations
4,036
Kernel languages, 28, 31, 34 or 35 templates each
13
Languages in the direct-flow breadth baseline
13
Immutable evidence release
v0.8.0

How decisive, and how correct

Overall analyzer decisiveness and correctness Scatter plot for Overall. The horizontal axis is decisive answers divided by 886 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 886 assertions answered Correctness among decisive answers Bifrost CodeQL Joern Semgrep CE OpenTaint Infer FlowDroid Pysa

Overall decisiveness uses all 886 assertions in the 13 kernel corpus as every analyzer's denominator; a language with no analyzer entry therefore remains visible as unanswered coverage.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnanswered
Bifrost 494/886 (55.8%) 494/494 (100.0%) 392
CodeQL 619/886 (69.9%) 499/619 (80.6%) 267
Joern 412/886 (46.5%) 323/412 (78.4%) 474
Semgrep CE 154/886 (17.4%) 132/154 (85.7%) 732
OpenTaint 140/886 (15.8%) 121/140 (86.4%) 746
Infer 194/886 (21.9%) 157/194 (80.9%) 692
FlowDroid 140/886 (15.8%) 122/140 (87.1%) 746
Pysa 70/886 (7.9%) 56/70 (80.0%) 816
c analyzer decisiveness and correctness Scatter plot for c. The horizontal axis is decisive answers divided by 56 assertions. The vertical axis is correct answers divided by decisive answers, from 0% to 100%; a line marks the 50% a blind answerer would score. 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 56 assertions answered Correctness among decisive answers Bifrost CodeQL Semgrep CE Infer

c is its own 56-assertion population. Analyzers without a c kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 42/56 (75.0%) 42/42 (100.0%) 14
CodeQL 0/56 (0.0%) 0/0 (0.0%) 56
Semgrep CE 14/56 (25.0%) 12/14 (85.7%) 42
Infer 56/56 (100.0%) 47/56 (83.9%) 0
cpp analyzer decisiveness and correctness Scatter plot for cpp. The horizontal axis is decisive answers divided by 68 assertions. The vertical axis is correct answers divided by decisive answers, from 0% to 100%; a line marks the 50% a blind answerer would score. 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 68 assertions answered Correctness among decisive answers Bifrost CodeQL Semgrep CE Infer

cpp is its own 68-assertion population. Analyzers without a cpp kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 30/68 (44.1%) 30/30 (100.0%) 38
CodeQL 0/68 (0.0%) 0/0 (0.0%) 68
Semgrep CE 14/68 (20.6%) 12/14 (85.7%) 54
Infer 68/68 (100.0%) 53/68 (77.9%) 0
csharp analyzer decisiveness and correctness Scatter plot for csharp. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL

csharp is its own 70-assertion population. Analyzers without a csharp kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 36/70 (51.4%) 36/36 (100.0%) 34
CodeQL 69/70 (98.6%) 55/69 (79.7%) 1
go analyzer decisiveness and correctness Scatter plot for go. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL Semgrep CE

go is its own 70-assertion population. Analyzers without a go kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 40/70 (57.1%) 40/40 (100.0%) 30
CodeQL 70/70 (100.0%) 54/70 (77.1%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56
java analyzer decisiveness and correctness Scatter plot for java. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL Joern Semgrep CE OpenTaint Infer FlowDroid

java is its own 70-assertion population. Analyzers without a java kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 46/70 (65.7%) 46/46 (100.0%) 24
CodeQL 70/70 (100.0%) 57/70 (81.4%) 0
Joern 70/70 (100.0%) 57/70 (81.4%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56
OpenTaint 70/70 (100.0%) 61/70 (87.1%) 0
Infer 70/70 (100.0%) 57/70 (81.4%) 0
FlowDroid 70/70 (100.0%) 61/70 (87.1%) 0
javascript analyzer decisiveness and correctness Scatter plot for javascript. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL Joern Semgrep CE

javascript is its own 70-assertion population. Analyzers without a javascript kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 40/70 (57.1%) 40/40 (100.0%) 30
CodeQL 70/70 (100.0%) 57/70 (81.4%) 0
Joern 70/70 (100.0%) 55/70 (78.6%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56
kotlin analyzer decisiveness and correctness Scatter plot for kotlin. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL Semgrep CE OpenTaint FlowDroid

kotlin is its own 70-assertion population. Analyzers without a kotlin kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 32/70 (45.7%) 32/32 (100.0%) 38
CodeQL 70/70 (100.0%) 55/70 (78.6%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56
OpenTaint 70/70 (100.0%) 60/70 (85.7%) 0
FlowDroid 70/70 (100.0%) 61/70 (87.1%) 0
php analyzer decisiveness and correctness Scatter plot for php. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost Joern Semgrep CE

php is its own 70-assertion population. Analyzers without a php kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 34/70 (48.6%) 34/34 (100.0%) 36
Joern 70/70 (100.0%) 58/70 (82.9%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56
python analyzer decisiveness and correctness Scatter plot for python. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL Joern Semgrep CE Pysa

python is its own 70-assertion population. Analyzers without a python kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 36/70 (51.4%) 36/36 (100.0%) 34
CodeQL 69/70 (98.6%) 56/69 (81.2%) 1
Joern 70/70 (100.0%) 58/70 (82.9%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56
Pysa 70/70 (100.0%) 56/70 (80.0%) 0
ruby analyzer decisiveness and correctness Scatter plot for ruby. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL Joern Semgrep CE

ruby is its own 70-assertion population. Analyzers without a ruby kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 36/70 (51.4%) 36/36 (100.0%) 34
CodeQL 69/70 (98.6%) 57/69 (82.6%) 1
Joern 70/70 (100.0%) 46/70 (65.7%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56
rust analyzer decisiveness and correctness Scatter plot for rust. The horizontal axis is decisive answers divided by 62 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 62 assertions answered Correctness among decisive answers Bifrost CodeQL Joern Semgrep CE

rust is its own 62-assertion population. Analyzers without a rust kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 42/62 (67.7%) 42/42 (100.0%) 20
CodeQL 62/62 (100.0%) 51/62 (82.3%) 0
Joern 62/62 (100.0%) 49/62 (79.0%) 0
Semgrep CE 14/62 (22.6%) 12/14 (85.7%) 48
scala analyzer decisiveness and correctness Scatter plot for scala. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost

scala is its own 70-assertion population. Analyzers without a scala kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 42/70 (60.0%) 42/42 (100.0%) 28
typescript analyzer decisiveness and correctness Scatter plot for typescript. The horizontal axis is decisive answers divided by 70 assertions. The vertical axis is correct answers divided by decisive answers, from 40% to 100%; a line marks the 50% a blind answerer would score. 40% 50% 60% 70% 80% 90% 100% 0% 25% 50% 75% 100% blind baseline Decisiveness: share of 70 assertions answered Correctness among decisive answers Bifrost CodeQL Semgrep CE

typescript is its own 70-assertion population. Analyzers without a typescript kernel entry are absent, never plotted as zero.

Exact figures
AnalyzerDecisivenessCorrectness when decisiveUnansweredEvidence
Bifrost 38/70 (54.3%) 38/38 (100.0%) 32
CodeQL 70/70 (100.0%) 57/70 (81.4%) 0
Semgrep CE 14/70 (20.0%) 12/14 (85.7%) 56

Confusion matrix · DataFlowBench v0.8.0

Decisive outcomes only. inconclusive, unsupported and runner-error are capability coverage: they are excluded from the matrix and from the rates, and are never converted into clean negatives.

What this result means

DataFlowBench measures whether analyzers correctly decide semantic data-flow questions — and whether they stay quiet when they should. Cases are balanced positive/negative pairs of language-neutral semantic templates (aliasing, kills, call context, branch joins, exception paths, …). The kernels and the modeling matrix run under a benchmark-controlled model profile, so the engines are compared under a common contract; the tool-native probes run under a tool-native profile, measuring the shipped product instead. The two profiles answer different questions and are never combined.

This snapshot's bounded claim covers the synthetic direct-flow breadth baseline and the propagation kernels of the 13 kernel languages above on the taint track, plus the modeling matrix and the tool-native probe set in java, javascript, and python. Cores are sized per language (28, 31, 34 or 35 templates), so each kernel is read on its own denominator; language-only constructs are reported in separate language-extension tiers on the snapshot pages. It does not estimate real-project accuracy or other languages' kernel behavior; performance is characterized separately in the latency tier, on its own terms, and never folded into these scores. The tool-native rows describe shipped coverage on six probe templates rather than product accuracy at large. inconclusive, unsupported, and runner-error are capability coverage and are never converted into clean negatives.

Accuracy and language coverage — current snapshot

Every analyzer is shown in the same field. Farther right means documented data-flow support for more languages; higher means more correct assertions within the kernels the analyzer covers. A specialist can therefore show its accuracy without hiding the cost of its narrower language support. The horizontal axis is an analyzer capability, not a count of the adapters DataFlowBench happens to implement. Benchmark kernel participation remains visible in the exact figures below. These are two independent dimensions, not a combined score.

Current analyzer accuracy by supported data-flow languages Scatter plot of all analyzers in DataFlowBench v0.8.0. The horizontal axis is the documented number of languages in which each analyzer performs data-flow analysis. The vertical axis is decisive-correct assertions divided by every assertion in the benchmark kernels that analyzer covers, including non-answers in the denominator. 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% 0 2 4 6 8 10 12 14 Languages with documented data-flow support Accuracy within covered kernels Bifrost CodeQL Joern Semgrep CE OpenTaint Infer FlowDroid Pysa
  • generalist
  • specialist
Exact figures
AnalyzerScopeSupported languagesBenchmark kernelsAccuracyCorrectWrongIncomplete
Bifrost generalist 13C, C++, C#, Go, Java, JavaScript, Kotlin, PHP, Python, Ruby, Rust, Scala, TypeScript 13/13 (100.0%) 494/886 (55.8%) 4940392
CodeQL generalist 12C, C++, C#, Go, Java, JavaScript, Kotlin, Python, Ruby, Rust, Swift, TypeScript 11/13 (84.6%) 499/746 (66.9%) 499120127
Joern generalist 11C, C++, C#, Go, Java, JavaScript, Kotlin, PHP, Python, Ruby, Swift 6/13 (46.2%) 323/412 (78.4%) 323890
Semgrep CE generalist 11C, C++, Go, Java, JavaScript, Kotlin, PHP, Python, Ruby, Rust, TypeScript 11/13 (84.6%) 132/746 (17.7%) 13222592
OpenTaint specialist 2Java, Kotlin 2/13 (15.4%) 121/140 (86.4%) 121190
Infer specialist 3C, C++, Java 3/13 (23.1%) 157/194 (80.9%) 157370
FlowDroid specialist 2Java, Kotlin 2/13 (15.4%) 122/140 (87.1%) 122180
Pysa specialist 1Python 1/13 (7.7%) 56/70 (80.0%) 56140

Modeling matrix — is the model surface load-bearing?

A separate population from the kernels above, on the same benchmark-controlled profile. Twelve preregistered templates in six balanced categories — declared sources and sinks, declared propagators, declared sanitizers, opaque summaries, framework entry points, persistence boundaries — ask whether each tool's own model-declaration surface can express a category and be made to carry the flow. A category a tool cannot express is unsupported, decided before the tool is invoked. The scored partition differs per adapter, so the denominators differ by construction and are never pooled or ranked. Read the scored column first: it is the load-bearing number, and the ratio inside it is only meaningful against that tool's own partition.

Bifrost 0.11.4

java

7/7 decided correctly · 1 inconclusive 8 scored · 16 declined of 24

javascript

4/4 decided correctly · 4 inconclusive 8 scored · 16 declined of 24

python

7/7 decided correctly · 1 inconclusive 8 scored · 16 declined of 24

CodeQL 2.27.0

java

24/24 decided correctly 24 scored · 0 declined of 24

javascript

24/24 decided correctly 24 scored · 0 declined of 24

python

24/24 decided correctly 24 scored · 0 declined of 24

Joern 4.0.628

java

14/16 decided correctly 16 scored · 8 declined of 24

javascript

14/16 decided correctly 16 scored · 8 declined of 24

python

14/16 decided correctly 16 scored · 8 declined of 24

Semgrep CE 1.177.0

java

10/10 decided correctly 10 scored · 14 declined of 24

javascript

10/10 decided correctly 10 scored · 14 declined of 24

python

10/10 decided correctly 10 scored · 14 declined of 24

OpenTaint v0.4.6

java

12/12 decided correctly 12 scored · 12 declined of 24

Infer v1.3.0

java

10/10 decided correctly 10 scored · 14 declined of 24

FlowDroid 2.15.1

java

14/14 decided correctly 14 scored · 10 declined of 24

Pysa 0.10.0

python

20/20 decided correctly 20 scored · 4 declined of 24

  • correct
  • wrong
  • inconclusive — coverage, not a wrong answer
  • declined — outside this tool's model surface, decided before it runs

Read the scored partition, not the ratio: across the 16 language tiers it runs from 8 to 24 of the 24 assertions in the tier, so no two of these bars are the same exam. The per-tier counts behind every bar are on the analyzers page, and the template-by-template model-* outcomes on the semantic templates page.

Tool-native probes — what ships and decides on its own

A third population, under the tool-native model profile: six templates run with nothing supplied by DataFlowBench, against whatever ruleset, semantics, or policy pack the product ships. This measures product coverage, not engine accuracy. Tool-native results are never pooled with the benchmark-controlled kernels or the modeling matrix, and never compared number-to-number with them. A row of unsupported is a declared decline — the tool ships no source or sink catalog for this tier — and is never counted as a wrong answer. Those runs still witness the identity of the binary and ruleset that produced them.

Bifrost 0.11.4

java

declines the tier — declared coverage 0 scored · 12 declined of 12

javascript

declines the tier — declared coverage 0 scored · 12 declined of 12

python

declines the tier — declared coverage 0 scored · 12 declined of 12

CodeQL 2.27.0

java

11/12 decided correctly · 1 FP 12 scored · 0 declined of 12

javascript

9/12 decided correctly · 2 FP · 1 FN 12 scored · 0 declined of 12

python

10/12 decided correctly · 2 FP 12 scored · 0 declined of 12

Joern 4.0.628

java

declines the tier — declared coverage 0 scored · 12 declined of 12

javascript

declines the tier — declared coverage 0 scored · 12 declined of 12

python

declines the tier — declared coverage 0 scored · 12 declined of 12

Semgrep CE 1.177.0

java

declines the tier — declared coverage 0 scored · 12 declined of 12

javascript

declines the tier — declared coverage 0 scored · 12 declined of 12

python

8/12 decided correctly · 4 FP 12 scored · 0 declined of 12

OpenTaint v0.4.6

java

declines the tier — declared coverage 0 scored · 12 declined of 12

Infer v1.3.0

java

declines the tier — declared coverage 0 scored · 12 declined of 12

FlowDroid 2.15.1

java

declines the tier — declared coverage 0 scored · 12 declined of 12

Pysa 0.10.0

python

6/12 decided correctly · 6 FN 12 scored · 0 declined of 12

  • correct
  • wrong
  • inconclusive — coverage, not a wrong answer
  • declined — the product ships no catalog for this tier

11 of the 16 tiers decline the tier outright — a declared decline drawn as declined coverage, never as 0/12 — and the 5 that do decide answer 60 assertions between them (CodeQL and Pysa and Semgrep CE). Per-tier coverage is on the analyzers page, the native-* template outcomes on the semantic templates page, and every case — including the runs that decide nothing — on the case evidence page.

How long an answer takes — beside the answers, never inside them

A descriptive characterization of per-case analyzer wall-clock, published under the contract preregistered in docs/latency-tier.md, which merged before a single timestamp was captured — because a latency page assembled after the numbers were known, by the vendor of one of the engines measured, would deserve exactly the skepticism it would get. It is not a score and not a ranking of engines. Nothing on this site pools it with correctness: there is no combined number, no efficiency-adjusted rate, and no leaderboard that blends the two, so "correct but slow" and "fast but wrong" stay independently visible. Read this section against the correctness sections above, not through them.

Cold per-case analyzer-invocation wall-clock over the bound latency population One row per analyzer, ordered fastest median first, on a logarithmic time axis. Each row draws the tenth to ninetieth percentile as a thin line, the interquartile range as a bar, and the median as a thick tick, with the median also printed as a number. Indented rows are the declared phases of adapters whose invocation exposes more than one subprocess, and belong to the adapter above them only. Materialization phases are explicitly labelled and excluded from the analyzer total. The latency page carries every value, including the minima and maxima this chart does not draw, in the data tables behind its "Show the data table" disclosures. A solid caret below a row is that adapter's separately measured warm marginal per case; a dashed span above a row is its estimated per-invocation overhead, an upper bound measured on a trivial no-flow fixture, drawn across the range its repeats spanned rather than at a point. Neither is a cold median, neither affects the ordering, and neither is subtracted from anything. 30 ms 100 ms 300 ms 1 s 3 s 10 s 30 s 100 s 300 s Wall-clock per invocation, logarithmic scale Bifrost — median 109 ms, IQR 94 ms to 329 ms, p10–p90 87 ms to 5.41 s, over 1031 timed analyzer invocations. Estimated per-invocation overhead, an upper bound from a trivial no-flow python fixture: 83 ms to 206 ms across repeats Bifrost 1031 timed 109 ms FlowDroid — median 683 ms, IQR 664 ms to 856 ms, p10–p90 625 ms to 879 ms, over 154 timed analyzer invocations. Estimated per-invocation overhead, an upper bound from a trivial no-flow java fixture: 584 ms to 615 ms across repeats FlowDroid 154 timed 683 ms FlowDroid phase total — median 767 ms, IQR 671 ms to 859 ms, p10–p90 644 ms to 883 ms. Included in this adapter's analyzer total; phase detail compares only within the adapter. total 767 ms FlowDroid phase compile — median 454 ms, IQR 446 ms to 462 ms, p10–p90 436 ms to 467 ms. Materialization phase excluded from the analyzer total by contract. compile · excluded 454 ms FlowDroid phase dex — median 297 ms, IQR 292 ms to 305 ms, p10–p90 291 ms to 313 ms. Materialization phase excluded from the analyzer total by contract. dex · excluded 297 ms FlowDroid phase analyze — median 517 ms, IQR 510 ms to 524 ms, p10–p90 498 ms to 525 ms. Included in this adapter's analyzer total; phase detail compares only within the adapter. analyze 517 ms Semgrep CE — median 803 ms, IQR 787 ms to 880 ms, p10–p90 772 ms to 1.03 s, over 196 timed analyzer invocations. Estimated per-invocation overhead, an upper bound from a trivial no-flow kotlin fixture: 772 ms to 968 ms across repeats Semgrep CE 196 timed 803 ms Infer — median 1.05 s, IQR 243 ms to 3.62 s, p10–p90 241 ms to 3.68 s, over 204 timed analyzer invocations. Estimated per-invocation overhead, an upper bound from a trivial no-flow c fixture: 261 ms to 1.02 s across repeats Infer 204 timed 1.05 s Infer phase capture — median 872 ms, IQR 61 ms to 3.43 s, p10–p90 61 ms to 3.50 s. Included in this adapter's analyzer total; phase detail compares only within the adapter. capture 872 ms Infer phase analyze — median 183 ms, IQR 181 ms to 185 ms, p10–p90 179 ms to 187 ms. Included in this adapter's analyzer total; phase detail compares only within the adapter. analyze 183 ms Pysa — median 2.60 s, IQR 2.56 s to 2.74 s, p10–p90 2.53 s to 4.47 s, over 102 timed analyzer invocations. Estimated per-invocation overhead, an upper bound from a trivial no-flow python fixture: 2.45 s to 2.49 s across repeats Pysa 102 timed 2.60 s OpenTaint — median 4.01 s, IQR 3.91 s to 4.11 s, p10–p90 3.78 s to 4.20 s, over 152 timed analyzer invocations. Estimated per-invocation overhead, an upper bound from a trivial no-flow kotlin fixture: 3.92 s to 5.26 s across repeats OpenTaint 152 timed 4.01 s Joern — median 5.33 s, IQR 3.88 s to 7.37 s, p10–p90 3.82 s to 7.91 s, over 460 timed analyzer invocations. Estimated per-invocation overhead, an upper bound from a trivial no-flow php fixture: 3.54 s to 4.95 s across repeats Joern 460 timed 5.33 s CodeQL — median 58.2 s, IQR 36.0 s to 65.5 s, p10–p90 27.5 s to 95.1 s, over 858 timed analyzer invocations CodeQL 858 timed 58.2 s CodeQL phase database-create — median 3.08 s, IQR 1.80 s to 6.22 s, p10–p90 1.27 s to 12.3 s. Included in this adapter's analyzer total; phase detail compares only within the adapter. database-create 3.08 s CodeQL phase database-analyze — median 51.7 s, IQR 30.7 s to 62.4 s, p10–p90 24.5 s to 76.3 s. Included in this adapter's analyzer total; phase detail compares only within the adapter. database-analyze 51.7 s

Every timed analyzer invocation the freeze binds — 3157 of them, across every score tier and both model profiles. This is the widest denominator on the site and the only one here that is not a single population: an adapter's median mixes the languages it runs on, whose fixtures differ in size and whose front ends differ in cost. The per-kernel views on the latency page hold the language fixed.

  • median, printed beside every row
  • interquartile range (Q1–Q3)
  • p10–p90; the minimum and maximum are in the tables, not the whiskers
  • indented rows: an adapter's own declared phases
  • estimated per-invocation overhead (trivial fixture, upper bound), drawn across the range its repeats spanned, where that range starts at or above 25% of the row's median
How to read this figure

Every bar above is cold per-invocation wall-clock, and the warm marginal is a different quantity measured separately. Cold is what this benchmark actually runs — one process per case, start-up inside the number, because start-up is not observable from inside a single invocation. Read across runtimes, though, those bars overstate the steady-state gap: a JVM engine's row carries a JVM start a long-lived deployment pays once. So the other quantity is measured directly rather than estimated and subtracted — k cases through one tool process, for increasing k, reporting the slope of batch wall-clock against k. No caret is drawn in this view: the whole-corpus view mixes languages, and a marginal measured on one kernel is not a claim about a mixed-language row. The per-kernel views carry the marks. Only adapters whose released CLI exposes a real multi-case batch have a figure at all; the rest are not observable with the released CLI, and the warm-marginal section records every verdict, measured and declined, with the evidence behind it.

The dashed spans are estimates, not measurements of the same kind as the bars. Each is the wall-clock of one complete adapter invocation — same pipeline, same policy, rule or query, same subprocess shape — over a trivial no-flow fixture: both benchmark endpoints declared, nothing connecting them, nothing to find. That is fixed per-invocation overhead plus the trivial fixture's own near-zero analysis, so it is an upper bound on what an adapter pays before it starts work, and it is a cold single-shot execution — the same posture the bars are measured in, and not a steady-state one. It is never subtracted from a median, never substituted for one, and never used to order the rows. The width of a span is the figure's precision, not a decoration: each measurement is repeated a fixed number of times and what is published is the range those repeats spanned — never a mean, never one chosen repeat, and never withheld for repeats that disagree, because a disagreement widens the range and that is the honest consequence of it. A span is drawn only where the range starts at or above 25% of that row's own cold median — a threshold fixed in the amendment before any estimate was measured, read at the low end so that no mark can appear on the strength of one slow repeat. The CodeQL estimate is below it and carries no mark — small, not unmeasured. Each estimate is measured in one named language, stated on the row it annotates, because boot cost is not language-free. The estimates table carries every value — every repeat behind every range, every unmarked row, and every adapter for which no estimate could be taken at all.

The axis is logarithmic. Each labelled tick is three times the one before it, so equal distances are equal ratios, not equal durations — the gap from 100 ms to 300 ms is drawn the same width as the gap from 10 s to 30 s. That is the only way this bound corpus's medians fit in one picture: they span 109 ms to 58.2 s, and on a linear axis every analyzer except the slowest would be a sliver against the origin. Because a log axis is easy to misread, every median is also printed as a number at the right of its own row.

The indented rows are phases, and they are not comparable across adapters. Only the 3 adapters whose preregistered row declares more than one subprocess have them. A phase mark sits on the same axis as the totals because it is the same kind of measurement — wall-clock of a subprocess — but reading one adapter's phase against another adapter's total is precisely the comparison the granularity rule forbids. Read a phase against the adapter it is indented under, and nothing else.

Ordering is not scoring, and this is never pooled with correctness. Rows are sorted by median because an unsorted ranking is unreadable, not because latency is a result. No correctness figure appears in this chart and no number here is blended with one: there is no combined score anywhere on this site, and a fast analyzer that answers wrongly is neither rewarded nor penalised by anything drawn above. Cases an analyzer declined before invocation are absent, not entered as zero — entering them as zero would make the analyzers that decline the most look the fastest, which is exactly backwards.

The conditions these numbers were produced under. A single developer machine under light concurrent load, running the benchmark's standing sequential-run discipline — one analyzer at a time, never two at once. Latency numbers are comparable within one environment and are not comparable across machines. The bars carry no repeated trials and no warm-up iterations: the case population is the sample and its spread is the statistics, so the whiskers are that spread and not error bars on a repeated measurement. Every number is a cold start: each case spawns a fresh process, so JVM tools (Joern, OpenTaint, FlowDroid) pay full JVM start-up in every invocation, and Pysa pays Pyre initialization. In a resident deployment those costs amortize to once per session, so cross-runtime comparisons here overstate the steady-state gap; the boot share of a single invocation is not adapter-observable and is included, stated rather than estimated. How much they overstate it is now measured, not left as a caveat — separately, as the warm marginal cost of one more case in a process that has already started, wherever the released CLI let it be measured. And what an invocation costs before it has anything to find is estimated for every adapter it could be estimated for — the per-invocation overhead of one complete invocation over a trivial no-flow fixture, an upper bound published across the range its repeats spanned, never subtracted from a bar above. Read this as characterization of what the benchmark actually costs to run, at the order-of-magnitude and shape level, and not as a precise figure for any engine. The full contract, the per-adapter granularity, the cache declarations and every quartile are on the latency page.

Progress over time

How the benchmark and its analyzers evolved

Generalists

One line per analyzer across every freeze. The vertical axis counts decisive-correct assertions; the grey step is the full benchmark core, which grew from 32 assertions in v0.1.0 to 886 in v0.8.0. Nothing is pooled with other result profiles or turned into a combined score.

Decisive-correct assertions per analyzer across DataFlowBench snapshots A stepped grey line shows each snapshot's total benchmark-controlled core population, growing from 32 assertions in v0.1.0 to 886 in v0.8.0. Beneath it, one line per analyzer plots that analyzer's decisive-correct count, with a faint stub reaching that analyzer's own covered population. Every marker also carries its exact figures in a tooltip, and the toggle above this chart has a Data table view that lists every value as text. 020040060080010003296316738738738738738738886v0.1.01kernelv0.2.03kernelsv0.3.010kernelsv0.4.013kernelsv0.5.013kernelsv0.6.013kernelsv0.6.113kernelsv0.7.013kernelsv0.7.113kernelsv0.8.013kernels Assertions
Decisive-correct share of each analyzer's own covered population across DataFlowBench snapshots One line per analyzer on a nought to one hundred percent axis. Each point is that analyzer's decisive-correct count divided by the assertions in the kernels it covers, with inconclusive and unsupported outcomes left in the denominator. Marker size grows with the share of the snapshot's kernels the analyzer covers, and every point is annotated with that kernel count. Every marker also carries its exact figures in a tooltip, and the toggle above this chart has a Data table view that lists every value as text, including the denominator each percentage is over. 0% 20% 40% 60% 80% 100% 1/13/310/1013/1313/1313/1313/1313/1313/1313/131/13/310/1011/1311/1311/1311/1311/1311/1311/136/136/136/136/136/136/136/1311/1311/1311/1311/1311/1311/1311/13v0.1.01kernelv0.2.03kernelsv0.3.010kernelsv0.4.013kernelsv0.5.013kernelsv0.6.013kernelsv0.6.113kernelsv0.7.013kernelsv0.7.113kernelsv0.8.013kernels Share of its covered population
  • Bifrost · decisive-correct
  • CodeQL · decisive-correct
  • Joern · decisive-correct
  • Semgrep CE · decisive-correct
  • benchmark core population (the denominator)
  • remainder of that analyzer's own covered population
  • Bifrost · decisive-correct ÷ its covered population
  • CodeQL · decisive-correct ÷ its covered population
  • Joern · decisive-correct ÷ its covered population
  • Semgrep CE · decisive-correct ÷ its covered population
  • marker size and the k/n label under it: kernels covered of the snapshot's kernels

Marker: decisive-correct. Cap or stub: that analyzer's covered population. Grey step: the full core population. The covered-kernel view keeps non-answers in its denominator and shows coverage as k/n; missing runs are absent, never zero. Exact figures are in the table.

SnapshotAnalyzerDecisive-correctDecidedCoverage outcomesIts covered populationCorrect ÷ its covered populationBenchmark core populationSecondary: correct ÷ decided
v0.1.0Bifrost17221032(1 of 1 kernel)53.1%(17 of 32 covered, 1 of 1 kernel)3277.3% (17 of 22 decided)
v0.1.0CodeQL2732032(1 of 1 kernel)84.4%(27 of 32 covered, 1 of 1 kernel)3284.4% (27 of 32 decided)
v0.2.0Bifrost52682896(3 of 3 kernels)54.2%(52 of 96 covered, 3 of 3 kernels)9676.5% (52 of 68 decided)
v0.2.0CodeQL8496096(3 of 3 kernels)87.5%(84 of 96 covered, 3 of 3 kernels)9687.5% (84 of 96 decided)
v0.3.0Bifrost163166150316(10 of 10 kernels)51.6%(163 of 316 covered, 10 of 10 kernels)31698.2% (163 of 166 decided)
v0.3.0CodeQL2763160316(10 of 10 kernels)87.3%(276 of 316 covered, 10 of 10 kernels)31687.3% (276 of 316 decided)
v0.4.0Bifrost222227511738(13 of 13 kernels)30.1%(222 of 738 covered, 13 of 13 kernels)73897.8% (222 of 227 decided)
v0.4.0CodeQL5066220622(11 of 13 kernels)81.4%(506 of 622 covered, 11 of 13 kernels)73881.4% (506 of 622 decided)
v0.4.0Joern2703440344(6 of 13 kernels)78.5%(270 of 344 covered, 6 of 13 kernels)73878.5% (270 of 344 decided)
v0.4.0Semgrep CE132154468622(11 of 13 kernels)21.2%(132 of 622 covered, 11 of 13 kernels)73885.7% (132 of 154 decided)
v0.5.0Bifrost435440298738(13 of 13 kernels)58.9%(435 of 738 covered, 13 of 13 kernels)73898.9% (435 of 440 decided)
v0.5.0CodeQL5066220622(11 of 13 kernels)81.4%(506 of 622 covered, 11 of 13 kernels)73881.4% (506 of 622 decided)
v0.5.0Joern2703440344(6 of 13 kernels)78.5%(270 of 344 covered, 6 of 13 kernels)73878.5% (270 of 344 decided)
v0.5.0Semgrep CE132154468622(11 of 13 kernels)21.2%(132 of 622 covered, 11 of 13 kernels)73885.7% (132 of 154 decided)
v0.6.0Bifrost435440298738(13 of 13 kernels)58.9%(435 of 738 covered, 13 of 13 kernels)73898.9% (435 of 440 decided)
v0.6.0CodeQL5066220622(11 of 13 kernels)81.4%(506 of 622 covered, 11 of 13 kernels)73881.4% (506 of 622 decided)
v0.6.0Joern2703440344(6 of 13 kernels)78.5%(270 of 344 covered, 6 of 13 kernels)73878.5% (270 of 344 decided)
v0.6.0Semgrep CE132154468622(11 of 13 kernels)21.2%(132 of 622 covered, 11 of 13 kernels)73885.7% (132 of 154 decided)
v0.6.1Bifrost446446292738(13 of 13 kernels)60.4%(446 of 738 covered, 13 of 13 kernels)738100.0% (446 of 446 decided)
v0.6.1CodeQL5066220622(11 of 13 kernels)81.4%(506 of 622 covered, 11 of 13 kernels)73881.4% (506 of 622 decided)
v0.6.1Joern2703440344(6 of 13 kernels)78.5%(270 of 344 covered, 6 of 13 kernels)73878.5% (270 of 344 decided)
v0.6.1Semgrep CE132154468622(11 of 13 kernels)21.2%(132 of 622 covered, 11 of 13 kernels)73885.7% (132 of 154 decided)
v0.7.0Bifrost450450288738(13 of 13 kernels)61.0%(450 of 738 covered, 13 of 13 kernels)738100.0% (450 of 450 decided)
v0.7.0CodeQL5016175622(11 of 13 kernels)80.5%(501 of 622 covered, 11 of 13 kernels)73881.2% (501 of 617 decided)
v0.7.0Joern2703440344(6 of 13 kernels)78.5%(270 of 344 covered, 6 of 13 kernels)73878.5% (270 of 344 decided)
v0.7.0Semgrep CE132154468622(11 of 13 kernels)21.2%(132 of 622 covered, 11 of 13 kernels)73885.7% (132 of 154 decided)
v0.7.1Bifrost452452286738(13 of 13 kernels)61.2%(452 of 738 covered, 13 of 13 kernels)738100.0% (452 of 452 decided)
v0.7.1CodeQL5016175622(11 of 13 kernels)80.5%(501 of 622 covered, 11 of 13 kernels)73881.2% (501 of 617 decided)
v0.7.1Joern2703440344(6 of 13 kernels)78.5%(270 of 344 covered, 6 of 13 kernels)73878.5% (270 of 344 decided)
v0.7.1Semgrep CE132154468622(11 of 13 kernels)21.2%(132 of 622 covered, 11 of 13 kernels)73885.7% (132 of 154 decided)
v0.8.0Bifrost494494392886(13 of 13 kernels)55.8%(494 of 886 covered, 13 of 13 kernels)886100.0% (494 of 494 decided)
v0.8.0CodeQL499619127746(11 of 13 kernels)66.9%(499 of 746 covered, 11 of 13 kernels)88680.6% (499 of 619 decided)
v0.8.0Joern3234120412(6 of 13 kernels)78.4%(323 of 412 covered, 6 of 13 kernels)88678.4% (323 of 412 decided)
v0.8.0Semgrep CE132154592746(11 of 13 kernels)17.7%(132 of 746 covered, 11 of 13 kernels)88685.7% (132 of 154 decided)

The exact values behind both chart views, one row per analyzer and snapshot.

Specialists

Analyzers deliberately focused on one ecosystem or a small related family, separated because breadth and specialization are not directly comparable. Uses the same percentage axis as the generalists' covered-kernel view.

Decisive-correct share of each analyzer's own covered population across DataFlowBench snapshots One line per analyzer on a nought to one hundred percent axis. Each point is that analyzer's decisive-correct count divided by the assertions in the kernels it covers, with inconclusive and unsupported outcomes left in the denominator. Marker size grows with the share of the snapshot's kernels the analyzer covers, and every point is annotated with that kernel count. Every marker also carries its exact figures in a tooltip, and the toggle above this chart has a Data table view that lists every value as text, including the denominator each percentage is over. 0% 20% 40% 60% 80% 100% 2/132/132/132/132/133/133/133/133/133/132/132/132/132/132/131/131/131/131/131/13v0.6.013kernelsv0.6.113kernelsv0.7.013kernelsv0.7.113kernelsv0.8.013kernels Share of its covered population
  • OpenTaint · decisive-correct ÷ its covered population
  • Infer · decisive-correct ÷ its covered population
  • FlowDroid · decisive-correct ÷ its covered population
  • Pysa · decisive-correct ÷ its covered population
  • marker size and the k/n label under it: kernels covered of the snapshot's kernels

Same scales and encoding as the generalists. The covered-kernel view keeps specialist scope visible as k/n; exact figures are in the table.

SnapshotAnalyzerDecisive-correctDecidedCoverage outcomesIts covered populationCorrect ÷ its covered populationBenchmark core populationSecondary: correct ÷ decided
v0.6.0OpenTaint991160116(2 of 13 kernels)85.3%(99 of 116 covered, 2 of 13 kernels)73885.3% (99 of 116 decided)
v0.6.0Infer1401620162(3 of 13 kernels)86.4%(140 of 162 covered, 3 of 13 kernels)73886.4% (140 of 162 decided)
v0.6.0FlowDroid981160116(2 of 13 kernels)84.5%(98 of 116 covered, 2 of 13 kernels)73884.5% (98 of 116 decided)
v0.6.0Pysa4758058(1 of 13 kernels)81.0%(47 of 58 covered, 1 of 13 kernels)73881.0% (47 of 58 decided)
v0.6.1OpenTaint991160116(2 of 13 kernels)85.3%(99 of 116 covered, 2 of 13 kernels)73885.3% (99 of 116 decided)
v0.6.1Infer1401620162(3 of 13 kernels)86.4%(140 of 162 covered, 3 of 13 kernels)73886.4% (140 of 162 decided)
v0.6.1FlowDroid981160116(2 of 13 kernels)84.5%(98 of 116 covered, 2 of 13 kernels)73884.5% (98 of 116 decided)
v0.6.1Pysa4758058(1 of 13 kernels)81.0%(47 of 58 covered, 1 of 13 kernels)73881.0% (47 of 58 decided)
v0.7.0OpenTaint991160116(2 of 13 kernels)85.3%(99 of 116 covered, 2 of 13 kernels)73885.3% (99 of 116 decided)
v0.7.0Infer1401620162(3 of 13 kernels)86.4%(140 of 162 covered, 3 of 13 kernels)73886.4% (140 of 162 decided)
v0.7.0FlowDroid981160116(2 of 13 kernels)84.5%(98 of 116 covered, 2 of 13 kernels)73884.5% (98 of 116 decided)
v0.7.0Pysa4758058(1 of 13 kernels)81.0%(47 of 58 covered, 1 of 13 kernels)73881.0% (47 of 58 decided)
v0.7.1OpenTaint1011160116(2 of 13 kernels)87.1%(101 of 116 covered, 2 of 13 kernels)73887.1% (101 of 116 decided)
v0.7.1Infer1401620162(3 of 13 kernels)86.4%(140 of 162 covered, 3 of 13 kernels)73886.4% (140 of 162 decided)
v0.7.1FlowDroid981160116(2 of 13 kernels)84.5%(98 of 116 covered, 2 of 13 kernels)73884.5% (98 of 116 decided)
v0.7.1Pysa4758058(1 of 13 kernels)81.0%(47 of 58 covered, 1 of 13 kernels)73881.0% (47 of 58 decided)
v0.8.0OpenTaint1211400140(2 of 13 kernels)86.4%(121 of 140 covered, 2 of 13 kernels)88686.4% (121 of 140 decided)
v0.8.0Infer1571940194(3 of 13 kernels)80.9%(157 of 194 covered, 3 of 13 kernels)88680.9% (157 of 194 decided)
v0.8.0FlowDroid1221400140(2 of 13 kernels)87.1%(122 of 140 covered, 2 of 13 kernels)88687.1% (122 of 140 decided)
v0.8.0Pysa5670070(1 of 13 kernels)80.0%(56 of 70 covered, 1 of 13 kernels)88680.0% (56 of 70 decided)

The exact values behind both chart views, one row per analyzer and snapshot.

Kernels at a glance

One bar per analyzer, one panel per kernel. Each bar is that analyzer's own core population, split into decisive-correct, decisive-wrong, incomplete, and declined — the last two are capability coverage, each drawn as its own segment so the remainder of a bar is never read as a wrong answer. Panels are separate populations with separate denominators: compare bars inside a panel, never across panels, and never against the modeling or tool-native sections above. Every kernel here is on the benchmark-controlled model profile.

c 56 assertions

Bifrost

42/56

CodeQL

0/56

Infer

47/56

Semgrep CE

12/56

cpp 68 assertions

Bifrost

30/68

CodeQL

0/68

Infer

53/68

Semgrep CE

12/68

csharp 70 assertions

Bifrost

36/70

CodeQL

55/70

go 70 assertions

Bifrost

40/70

CodeQL

54/70

Semgrep CE

12/70

java 70 assertions

Bifrost

46/70

CodeQL

57/70

FlowDroid

61/70

Infer

57/70

Joern

57/70

OpenTaint

61/70

Semgrep CE

12/70

javascript 70 assertions

Bifrost

40/70

CodeQL

57/70

Joern

55/70

Semgrep CE

12/70

kotlin 70 assertions

Bifrost

32/70

CodeQL

55/70

FlowDroid

61/70

OpenTaint

60/70

Semgrep CE

12/70

php 70 assertions

Bifrost

34/70

Joern

58/70

Semgrep CE

12/70

python 70 assertions

Bifrost

36/70

CodeQL

56/70

Joern

58/70

Pysa

56/70

Semgrep CE

12/70

ruby 70 assertions

Bifrost

36/70

CodeQL

57/70

Joern

46/70

Semgrep CE

12/70

rust 62 assertions

Bifrost

42/62

CodeQL

51/62

Joern

49/62

Semgrep CE

12/62

scala 70 assertions

Bifrost

42/70

typescript 70 assertions

Bifrost

38/70

CodeQL

57/70

Semgrep CE

12/70

  • decisive-correct
  • decisive-wrong
  • incomplete — inconclusive or runner-error, coverage rather than a wrong answer
  • declined — unsupported by declared capability, decided before the analyzer runs

None of the 49 analyzer rows published across these 13 kernels answers its whole core correctly, and 20 of them answer every assertion in their kernel definitively — nothing incomplete, nothing declined. The template-by-template outcome for every one of these rows — each balanced positive/negative pair, per analyzer — is on the semantic templates page, with per-case classifications and retained raw evidence on the case evidence page.

Direct-flow breadth across languages

The breadth baseline is one balanced direct-propagation pair per language, on Bifrost's all-language smoke run: a separate, much smaller population from the kernels above, and never pooled with them. Each chip is one language, with a mark for the positive and the negative assertion.

  • c
  • cpp
  • csharp
  • go
  • java
  • javascript
  • kotlin
  • php
  • python
  • ruby
  • rust
  • scala
  • typescript

Bifrost answers both assertions correctly in 13 of 13 languages. The exact reached / not-reached outcome for each of those assertions is on the semantic templates page, under the smoke scorecard's direct-propagation rows.

Why these numbers are trustworthy

Every count on this page is generated from the freeze manifest (582b25896487…), which digest-binds the benchmark revision, every case and fixture, every analyzer's identity, their normalized reports, and one retained raw artifact per result. CI regenerates the numbers from the manifest and fails on any drift; hand-authored prose cannot override a generated count. See reproduction to verify locally, or dive into the full snapshot — analyzers, languages, templates, and per-case evidence.