Current overview of DataFlowBench
Nobody is perfect. No analyzer answers a whole core correctly in any of the 13 languages, and of the 8 in the field only Bifrost makes no decisive mistake at all — bought by declining 392 of the 886 core assertions rather than guessing at them.
Each kernel poses one positive and one negative assertion per semantic template, and each language's kernel is its own population — cores of 28, 31, 34 or 35 templates, never pooled or ranked against each other, and never compared with the smaller cores of an earlier snapshot. Because every pair is balanced, an analyzer that always answers the same way — or answers blindly where it cannot see a construct — scores exactly half on the affected pairs: read correctness against that 50% blind baseline, and read approximation character from each vendor's TPR/FPR split in the details dialog. A blank analyzer on a kernel means no report in this freeze: no extractor, no frontend, or no adapter, which is coverage rather than a score.
DataFlowBench is SlopCop's own benchmark for semantic data flow: we build it to measure where Bifrost stands against 7 reference analyzers, and to keep ourselves honest while doing it. The methodology is deliberately analyzer-neutral — every number here is generated from digest-bound freeze evidence that we cannot edit after the fact, misses and crashes included — but the motivation is not neutral, and we would rather say so.
What v0.8.0 contains
This is the first freeze of the expanded recursive-composition kernels: 1,000 cases across 82 freshly rerun reports produce 4,036 outcomes. The release adds 556 cells and preserves 124 changed outcomes from the previous population as measured evidence, including CodeQL's C/C++ incompleteness and Bifrost's newly decisive JavaScript, TypeScript, C#, and Ruby cells. Incomplete and unsupported outcomes remain distinct.
Below the kernels sit two further populations: the modeling matrix, which asks whether each tool's own model-declaration surface is load-bearing, and the tool-native probes, which ask what each product decides with nothing supplied by us. Beside them sits the latency tier — descriptive per-case timings on a stated machine, kept strictly beside the correctness results. These timings were freshly measured on the current pins, with desktop activity, cache state and sandbox/elevated execution contexts retained; no uncontended-host or timing-parity claim is implied. All populations have separate denominators. There is no combined leaderboard, and benchmark-controlled results are never pooled with, or compared number-to-number against, tool-native ones.
On the direct-flow breadth baseline, Bifrost answered both assertions correctly in 13 of 13 languages.
- Frozen case results, across three populations
- 4,036
- Kernel languages, 28, 31, 34 or 35 templates each
- 13
- Languages in the direct-flow breadth baseline
- 13
- Immutable evidence release
- v0.8.0
How decisive, and how correct
Overall decisiveness uses all 886 assertions in the 13 kernel corpus as every analyzer's denominator; a language with no analyzer entry therefore remains visible as unanswered coverage.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered |
|---|---|---|---|
| Bifrost | 494/886 (55.8%) | 494/494 (100.0%) | 392 |
| CodeQL | 619/886 (69.9%) | 499/619 (80.6%) | 267 |
| Joern | 412/886 (46.5%) | 323/412 (78.4%) | 474 |
| Semgrep CE | 154/886 (17.4%) | 132/154 (85.7%) | 732 |
| OpenTaint | 140/886 (15.8%) | 121/140 (86.4%) | 746 |
| Infer | 194/886 (21.9%) | 157/194 (80.9%) | 692 |
| FlowDroid | 140/886 (15.8%) | 122/140 (87.1%) | 746 |
| Pysa | 70/886 (7.9%) | 56/70 (80.0%) | 816 |
c is its own 56-assertion population. Analyzers without a c kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 42/56 (75.0%) | 42/42 (100.0%) | 14 | |
| CodeQL | 0/56 (0.0%) | 0/0 (0.0%) | 56 | |
| Semgrep CE | 14/56 (25.0%) | 12/14 (85.7%) | 42 | |
| Infer | 56/56 (100.0%) | 47/56 (83.9%) | 0 |
cpp is its own 68-assertion population. Analyzers without a cpp kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 30/68 (44.1%) | 30/30 (100.0%) | 38 | |
| CodeQL | 0/68 (0.0%) | 0/0 (0.0%) | 68 | |
| Semgrep CE | 14/68 (20.6%) | 12/14 (85.7%) | 54 | |
| Infer | 68/68 (100.0%) | 53/68 (77.9%) | 0 |
csharp is its own 70-assertion population. Analyzers without a csharp kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 36/70 (51.4%) | 36/36 (100.0%) | 34 | |
| CodeQL | 69/70 (98.6%) | 55/69 (79.7%) | 1 |
go is its own 70-assertion population. Analyzers without a go kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 40/70 (57.1%) | 40/40 (100.0%) | 30 | |
| CodeQL | 70/70 (100.0%) | 54/70 (77.1%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 |
java is its own 70-assertion population. Analyzers without a java kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 46/70 (65.7%) | 46/46 (100.0%) | 24 | |
| CodeQL | 70/70 (100.0%) | 57/70 (81.4%) | 0 | |
| Joern | 70/70 (100.0%) | 57/70 (81.4%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 | |
| OpenTaint | 70/70 (100.0%) | 61/70 (87.1%) | 0 | |
| Infer | 70/70 (100.0%) | 57/70 (81.4%) | 0 | |
| FlowDroid | 70/70 (100.0%) | 61/70 (87.1%) | 0 |
javascript is its own 70-assertion population. Analyzers without a javascript kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 40/70 (57.1%) | 40/40 (100.0%) | 30 | |
| CodeQL | 70/70 (100.0%) | 57/70 (81.4%) | 0 | |
| Joern | 70/70 (100.0%) | 55/70 (78.6%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 |
kotlin is its own 70-assertion population. Analyzers without a kotlin kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 32/70 (45.7%) | 32/32 (100.0%) | 38 | |
| CodeQL | 70/70 (100.0%) | 55/70 (78.6%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 | |
| OpenTaint | 70/70 (100.0%) | 60/70 (85.7%) | 0 | |
| FlowDroid | 70/70 (100.0%) | 61/70 (87.1%) | 0 |
php is its own 70-assertion population. Analyzers without a php kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 34/70 (48.6%) | 34/34 (100.0%) | 36 | |
| Joern | 70/70 (100.0%) | 58/70 (82.9%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 |
python is its own 70-assertion population. Analyzers without a python kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 36/70 (51.4%) | 36/36 (100.0%) | 34 | |
| CodeQL | 69/70 (98.6%) | 56/69 (81.2%) | 1 | |
| Joern | 70/70 (100.0%) | 58/70 (82.9%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 | |
| Pysa | 70/70 (100.0%) | 56/70 (80.0%) | 0 |
ruby is its own 70-assertion population. Analyzers without a ruby kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 36/70 (51.4%) | 36/36 (100.0%) | 34 | |
| CodeQL | 69/70 (98.6%) | 57/69 (82.6%) | 1 | |
| Joern | 70/70 (100.0%) | 46/70 (65.7%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 |
rust is its own 62-assertion population. Analyzers without a rust kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 42/62 (67.7%) | 42/42 (100.0%) | 20 | |
| CodeQL | 62/62 (100.0%) | 51/62 (82.3%) | 0 | |
| Joern | 62/62 (100.0%) | 49/62 (79.0%) | 0 | |
| Semgrep CE | 14/62 (22.6%) | 12/14 (85.7%) | 48 |
scala is its own 70-assertion population. Analyzers without a scala kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 42/70 (60.0%) | 42/42 (100.0%) | 28 |
typescript is its own 70-assertion population. Analyzers without a typescript kernel entry are absent, never plotted as zero.
Exact figures
| Analyzer | Decisiveness | Correctness when decisive | Unanswered | Evidence |
|---|---|---|---|---|
| Bifrost | 38/70 (54.3%) | 38/38 (100.0%) | 32 | |
| CodeQL | 70/70 (100.0%) | 57/70 (81.4%) | 0 | |
| Semgrep CE | 14/70 (20.0%) | 12/14 (85.7%) | 56 |
What this result means
DataFlowBench measures whether analyzers correctly decide semantic
data-flow questions — and whether they stay quiet when they should. Cases
are balanced positive/negative pairs of language-neutral semantic templates
(aliasing, kills, call context, branch joins, exception paths, …). The
kernels and the modeling matrix run under a
benchmark-controlled model profile, so the engines are compared
under a common contract; the tool-native probes run under a
tool-native profile, measuring the shipped product instead. The
two profiles answer different questions and are never combined.
This snapshot's bounded claim covers the synthetic direct-flow breadth
baseline and the propagation kernels of the 13 kernel
languages above on the taint track, plus the modeling matrix
and the tool-native probe set in java, javascript, and python. Cores are
sized per language (28, 31, 34 or 35 templates), so each kernel is read on
its own denominator; language-only constructs are reported in separate
language-extension tiers on the snapshot pages. It does not
estimate real-project accuracy or other languages' kernel behavior;
performance is characterized separately in the latency tier, on its own
terms, and never folded into these scores. The tool-native rows describe shipped coverage on six
probe templates rather than product accuracy at large.
inconclusive, unsupported, and
runner-error are capability coverage and are never converted
into clean negatives.
Accuracy and language coverage — current snapshot
Every analyzer is shown in the same field. Farther right means documented data-flow support for more languages; higher means more correct assertions within the kernels the analyzer covers. A specialist can therefore show its accuracy without hiding the cost of its narrower language support. The horizontal axis is an analyzer capability, not a count of the adapters DataFlowBench happens to implement. Benchmark kernel participation remains visible in the exact figures below. These are two independent dimensions, not a combined score.
- generalist
- specialist
Exact figures
| Analyzer | Scope | Supported languages | Benchmark kernels | Accuracy | Correct | Wrong | Incomplete |
|---|---|---|---|---|---|---|---|
| Bifrost | generalist | 13C, C++, C#, Go, Java, JavaScript, Kotlin, PHP, Python, Ruby, Rust, Scala, TypeScript | 13/13 (100.0%) | 494/886 (55.8%) | 494 | 0 | 392 |
| CodeQL | generalist | 12C, C++, C#, Go, Java, JavaScript, Kotlin, Python, Ruby, Rust, Swift, TypeScript | 11/13 (84.6%) | 499/746 (66.9%) | 499 | 120 | 127 |
| Joern | generalist | 11C, C++, C#, Go, Java, JavaScript, Kotlin, PHP, Python, Ruby, Swift | 6/13 (46.2%) | 323/412 (78.4%) | 323 | 89 | 0 |
| Semgrep CE | generalist | 11C, C++, Go, Java, JavaScript, Kotlin, PHP, Python, Ruby, Rust, TypeScript | 11/13 (84.6%) | 132/746 (17.7%) | 132 | 22 | 592 |
| OpenTaint | specialist | 2Java, Kotlin | 2/13 (15.4%) | 121/140 (86.4%) | 121 | 19 | 0 |
| Infer | specialist | 3C, C++, Java | 3/13 (23.1%) | 157/194 (80.9%) | 157 | 37 | 0 |
| FlowDroid | specialist | 2Java, Kotlin | 2/13 (15.4%) | 122/140 (87.1%) | 122 | 18 | 0 |
| Pysa | specialist | 1Python | 1/13 (7.7%) | 56/70 (80.0%) | 56 | 14 | 0 |
Modeling matrix — is the model surface load-bearing?
A separate population from the kernels above, on the same
benchmark-controlled profile. Twelve preregistered templates in
six balanced categories — declared sources and sinks, declared propagators,
declared sanitizers, opaque summaries, framework entry points, persistence
boundaries — ask whether each tool's own model-declaration surface
can express a category and be made to carry the flow. A category a tool
cannot express is unsupported, decided before
the tool is invoked. The scored partition differs per adapter, so the
denominators differ by construction and are never pooled or ranked.
Read the scored column first: it is the load-bearing number, and the
ratio inside it is only meaningful against that tool's own partition.
Bifrost 0.11.4
CodeQL 2.27.0
Joern 4.0.628
Semgrep CE 1.177.0
OpenTaint v0.4.6
Infer v1.3.0
FlowDroid 2.15.1
Pysa 0.10.0
Read the scored partition, not the ratio: across the 16
language tiers it runs from 8 to 24 of the
24 assertions in the tier, so no two of these bars are
the same exam. The per-tier counts behind every bar are on the
analyzers page, and the
template-by-template model-* outcomes on the
semantic templates page.
Tool-native probes — what ships and decides on its own
A third population, under the tool-native model profile: six
templates run with nothing supplied by DataFlowBench, against
whatever ruleset, semantics, or policy pack the product ships. This measures
product coverage, not engine accuracy.
Tool-native results are never pooled with the benchmark-controlled kernels
or the modeling matrix, and never compared number-to-number with them.
A row of unsupported is a declared decline —
the tool ships no source or sink catalog for this tier — and is never counted
as a wrong answer. Those runs still witness the identity of the binary and
ruleset that produced them.
Bifrost 0.11.4
CodeQL 2.27.0
Joern 4.0.628
Semgrep CE 1.177.0
OpenTaint v0.4.6
Infer v1.3.0
FlowDroid 2.15.1
Pysa 0.10.0
11 of the 16 tiers decline the tier
outright — a declared decline drawn as declined coverage, never as
0/12 — and the 5 that do decide answer 60 assertions
between them (CodeQL and Pysa and Semgrep CE). Per-tier coverage is
on the analyzers page, the
native-* template outcomes on the
semantic templates page, and every
case — including the runs that decide nothing — on the
case evidence page.
How long an answer takes — beside the answers, never inside them
A descriptive characterization of per-case analyzer
wall-clock, published under the contract preregistered in
docs/latency-tier.md,
which merged before a single timestamp was captured — because a latency page
assembled after the numbers were known, by the vendor of one of the engines
measured, would deserve exactly the skepticism it would get. It is not a
score and not a ranking of engines. Nothing on this site pools it
with correctness: there is no combined number, no
efficiency-adjusted rate, and no leaderboard that blends the two, so "correct
but slow" and "fast but wrong" stay independently visible. Read this section
against the correctness sections above, not through them.
Every timed analyzer invocation the freeze binds — 3157 of them, across every score tier and both model profiles. This is the widest denominator on the site and the only one here that is not a single population: an adapter's median mixes the languages it runs on, whose fixtures differ in size and whose front ends differ in cost. The per-kernel views on the latency page hold the language fixed.
- median, printed beside every row
- interquartile range (Q1–Q3)
- p10–p90; the minimum and maximum are in the tables, not the whiskers
- indented rows: an adapter's own declared phases
- estimated per-invocation overhead (trivial fixture, upper bound), drawn across the range its repeats spanned, where that range starts at or above 25% of the row's median
How to read this figure
Every bar above is cold per-invocation wall-clock, and the warm marginal is a different quantity measured separately. Cold is what this benchmark actually runs — one process per case, start-up inside the number, because start-up is not observable from inside a single invocation. Read across runtimes, though, those bars overstate the steady-state gap: a JVM engine's row carries a JVM start a long-lived deployment pays once. So the other quantity is measured directly rather than estimated and subtracted — k cases through one tool process, for increasing k, reporting the slope of batch wall-clock against k. No caret is drawn in this view: the whole-corpus view mixes languages, and a marginal measured on one kernel is not a claim about a mixed-language row. The per-kernel views carry the marks. Only adapters whose released CLI exposes a real multi-case batch have a figure at all; the rest are not observable with the released CLI, and the warm-marginal section records every verdict, measured and declined, with the evidence behind it.
The dashed spans are estimates, not measurements of the same kind as the bars. Each is the wall-clock of one complete adapter invocation — same pipeline, same policy, rule or query, same subprocess shape — over a trivial no-flow fixture: both benchmark endpoints declared, nothing connecting them, nothing to find. That is fixed per-invocation overhead plus the trivial fixture's own near-zero analysis, so it is an upper bound on what an adapter pays before it starts work, and it is a cold single-shot execution — the same posture the bars are measured in, and not a steady-state one. It is never subtracted from a median, never substituted for one, and never used to order the rows. The width of a span is the figure's precision, not a decoration: each measurement is repeated a fixed number of times and what is published is the range those repeats spanned — never a mean, never one chosen repeat, and never withheld for repeats that disagree, because a disagreement widens the range and that is the honest consequence of it. A span is drawn only where the range starts at or above 25% of that row's own cold median — a threshold fixed in the amendment before any estimate was measured, read at the low end so that no mark can appear on the strength of one slow repeat. The CodeQL estimate is below it and carries no mark — small, not unmeasured. Each estimate is measured in one named language, stated on the row it annotates, because boot cost is not language-free. The estimates table carries every value — every repeat behind every range, every unmarked row, and every adapter for which no estimate could be taken at all.
The axis is logarithmic. Each labelled tick is three times the one before it, so equal distances are equal ratios, not equal durations — the gap from 100 ms to 300 ms is drawn the same width as the gap from 10 s to 30 s. That is the only way this bound corpus's medians fit in one picture: they span 109 ms to 58.2 s, and on a linear axis every analyzer except the slowest would be a sliver against the origin. Because a log axis is easy to misread, every median is also printed as a number at the right of its own row.
The indented rows are phases, and they are not comparable across adapters. Only the 3 adapters whose preregistered row declares more than one subprocess have them. A phase mark sits on the same axis as the totals because it is the same kind of measurement — wall-clock of a subprocess — but reading one adapter's phase against another adapter's total is precisely the comparison the granularity rule forbids. Read a phase against the adapter it is indented under, and nothing else.
Ordering is not scoring, and this is never pooled with correctness. Rows are sorted by median because an unsorted ranking is unreadable, not because latency is a result. No correctness figure appears in this chart and no number here is blended with one: there is no combined score anywhere on this site, and a fast analyzer that answers wrongly is neither rewarded nor penalised by anything drawn above. Cases an analyzer declined before invocation are absent, not entered as zero — entering them as zero would make the analyzers that decline the most look the fastest, which is exactly backwards.
The conditions these numbers were produced under. A single developer machine under light concurrent load, running the benchmark's standing sequential-run discipline — one analyzer at a time, never two at once. Latency numbers are comparable within one environment and are not comparable across machines. The bars carry no repeated trials and no warm-up iterations: the case population is the sample and its spread is the statistics, so the whiskers are that spread and not error bars on a repeated measurement. Every number is a cold start: each case spawns a fresh process, so JVM tools (Joern, OpenTaint, FlowDroid) pay full JVM start-up in every invocation, and Pysa pays Pyre initialization. In a resident deployment those costs amortize to once per session, so cross-runtime comparisons here overstate the steady-state gap; the boot share of a single invocation is not adapter-observable and is included, stated rather than estimated. How much they overstate it is now measured, not left as a caveat — separately, as the warm marginal cost of one more case in a process that has already started, wherever the released CLI let it be measured. And what an invocation costs before it has anything to find is estimated for every adapter it could be estimated for — the per-invocation overhead of one complete invocation over a trivial no-flow fixture, an upper bound published across the range its repeats spanned, never subtracted from a bar above. Read this as characterization of what the benchmark actually costs to run, at the order-of-magnitude and shape level, and not as a precise figure for any engine. The full contract, the per-adapter granularity, the cache declarations and every quartile are on the latency page.
Progress over time
How the benchmark and its analyzers evolved
Generalists
One line per analyzer across every freeze. The vertical axis counts decisive-correct assertions; the grey step is the full benchmark core, which grew from 32 assertions in v0.1.0 to 886 in v0.8.0. Nothing is pooled with other result profiles or turned into a combined score.
- Bifrost · decisive-correct
- CodeQL · decisive-correct
- Joern · decisive-correct
- Semgrep CE · decisive-correct
- benchmark core population (the denominator)
- remainder of that analyzer's own covered population
- Bifrost · decisive-correct ÷ its covered population
- CodeQL · decisive-correct ÷ its covered population
- Joern · decisive-correct ÷ its covered population
- Semgrep CE · decisive-correct ÷ its covered population
- marker size and the
k/nlabel under it: kernels covered of the snapshot's kernels
Marker: decisive-correct. Cap or stub: that analyzer's covered population.
Grey step: the full core population. The covered-kernel view keeps
non-answers in its denominator and shows coverage as k/n;
missing runs are absent, never zero. Exact figures are in the table.
| Snapshot | Analyzer | Decisive-correct | Decided | Coverage outcomes | Its covered population | Correct ÷ its covered population | Benchmark core population | Secondary: correct ÷ decided |
|---|---|---|---|---|---|---|---|---|
| v0.1.0 | Bifrost | 17 | 22 | 10 | 32(1 of 1 kernel) | 53.1%(17 of 32 covered, 1 of 1 kernel) | 32 | 77.3% (17 of 22 decided) |
| v0.1.0 | CodeQL | 27 | 32 | 0 | 32(1 of 1 kernel) | 84.4%(27 of 32 covered, 1 of 1 kernel) | 32 | 84.4% (27 of 32 decided) |
| v0.2.0 | Bifrost | 52 | 68 | 28 | 96(3 of 3 kernels) | 54.2%(52 of 96 covered, 3 of 3 kernels) | 96 | 76.5% (52 of 68 decided) |
| v0.2.0 | CodeQL | 84 | 96 | 0 | 96(3 of 3 kernels) | 87.5%(84 of 96 covered, 3 of 3 kernels) | 96 | 87.5% (84 of 96 decided) |
| v0.3.0 | Bifrost | 163 | 166 | 150 | 316(10 of 10 kernels) | 51.6%(163 of 316 covered, 10 of 10 kernels) | 316 | 98.2% (163 of 166 decided) |
| v0.3.0 | CodeQL | 276 | 316 | 0 | 316(10 of 10 kernels) | 87.3%(276 of 316 covered, 10 of 10 kernels) | 316 | 87.3% (276 of 316 decided) |
| v0.4.0 | Bifrost | 222 | 227 | 511 | 738(13 of 13 kernels) | 30.1%(222 of 738 covered, 13 of 13 kernels) | 738 | 97.8% (222 of 227 decided) |
| v0.4.0 | CodeQL | 506 | 622 | 0 | 622(11 of 13 kernels) | 81.4%(506 of 622 covered, 11 of 13 kernels) | 738 | 81.4% (506 of 622 decided) |
| v0.4.0 | Joern | 270 | 344 | 0 | 344(6 of 13 kernels) | 78.5%(270 of 344 covered, 6 of 13 kernels) | 738 | 78.5% (270 of 344 decided) |
| v0.4.0 | Semgrep CE | 132 | 154 | 468 | 622(11 of 13 kernels) | 21.2%(132 of 622 covered, 11 of 13 kernels) | 738 | 85.7% (132 of 154 decided) |
| v0.5.0 | Bifrost | 435 | 440 | 298 | 738(13 of 13 kernels) | 58.9%(435 of 738 covered, 13 of 13 kernels) | 738 | 98.9% (435 of 440 decided) |
| v0.5.0 | CodeQL | 506 | 622 | 0 | 622(11 of 13 kernels) | 81.4%(506 of 622 covered, 11 of 13 kernels) | 738 | 81.4% (506 of 622 decided) |
| v0.5.0 | Joern | 270 | 344 | 0 | 344(6 of 13 kernels) | 78.5%(270 of 344 covered, 6 of 13 kernels) | 738 | 78.5% (270 of 344 decided) |
| v0.5.0 | Semgrep CE | 132 | 154 | 468 | 622(11 of 13 kernels) | 21.2%(132 of 622 covered, 11 of 13 kernels) | 738 | 85.7% (132 of 154 decided) |
| v0.6.0 | Bifrost | 435 | 440 | 298 | 738(13 of 13 kernels) | 58.9%(435 of 738 covered, 13 of 13 kernels) | 738 | 98.9% (435 of 440 decided) |
| v0.6.0 | CodeQL | 506 | 622 | 0 | 622(11 of 13 kernels) | 81.4%(506 of 622 covered, 11 of 13 kernels) | 738 | 81.4% (506 of 622 decided) |
| v0.6.0 | Joern | 270 | 344 | 0 | 344(6 of 13 kernels) | 78.5%(270 of 344 covered, 6 of 13 kernels) | 738 | 78.5% (270 of 344 decided) |
| v0.6.0 | Semgrep CE | 132 | 154 | 468 | 622(11 of 13 kernels) | 21.2%(132 of 622 covered, 11 of 13 kernels) | 738 | 85.7% (132 of 154 decided) |
| v0.6.1 | Bifrost | 446 | 446 | 292 | 738(13 of 13 kernels) | 60.4%(446 of 738 covered, 13 of 13 kernels) | 738 | 100.0% (446 of 446 decided) |
| v0.6.1 | CodeQL | 506 | 622 | 0 | 622(11 of 13 kernels) | 81.4%(506 of 622 covered, 11 of 13 kernels) | 738 | 81.4% (506 of 622 decided) |
| v0.6.1 | Joern | 270 | 344 | 0 | 344(6 of 13 kernels) | 78.5%(270 of 344 covered, 6 of 13 kernels) | 738 | 78.5% (270 of 344 decided) |
| v0.6.1 | Semgrep CE | 132 | 154 | 468 | 622(11 of 13 kernels) | 21.2%(132 of 622 covered, 11 of 13 kernels) | 738 | 85.7% (132 of 154 decided) |
| v0.7.0 | Bifrost | 450 | 450 | 288 | 738(13 of 13 kernels) | 61.0%(450 of 738 covered, 13 of 13 kernels) | 738 | 100.0% (450 of 450 decided) |
| v0.7.0 | CodeQL | 501 | 617 | 5 | 622(11 of 13 kernels) | 80.5%(501 of 622 covered, 11 of 13 kernels) | 738 | 81.2% (501 of 617 decided) |
| v0.7.0 | Joern | 270 | 344 | 0 | 344(6 of 13 kernels) | 78.5%(270 of 344 covered, 6 of 13 kernels) | 738 | 78.5% (270 of 344 decided) |
| v0.7.0 | Semgrep CE | 132 | 154 | 468 | 622(11 of 13 kernels) | 21.2%(132 of 622 covered, 11 of 13 kernels) | 738 | 85.7% (132 of 154 decided) |
| v0.7.1 | Bifrost | 452 | 452 | 286 | 738(13 of 13 kernels) | 61.2%(452 of 738 covered, 13 of 13 kernels) | 738 | 100.0% (452 of 452 decided) |
| v0.7.1 | CodeQL | 501 | 617 | 5 | 622(11 of 13 kernels) | 80.5%(501 of 622 covered, 11 of 13 kernels) | 738 | 81.2% (501 of 617 decided) |
| v0.7.1 | Joern | 270 | 344 | 0 | 344(6 of 13 kernels) | 78.5%(270 of 344 covered, 6 of 13 kernels) | 738 | 78.5% (270 of 344 decided) |
| v0.7.1 | Semgrep CE | 132 | 154 | 468 | 622(11 of 13 kernels) | 21.2%(132 of 622 covered, 11 of 13 kernels) | 738 | 85.7% (132 of 154 decided) |
| v0.8.0 | Bifrost | 494 | 494 | 392 | 886(13 of 13 kernels) | 55.8%(494 of 886 covered, 13 of 13 kernels) | 886 | 100.0% (494 of 494 decided) |
| v0.8.0 | CodeQL | 499 | 619 | 127 | 746(11 of 13 kernels) | 66.9%(499 of 746 covered, 11 of 13 kernels) | 886 | 80.6% (499 of 619 decided) |
| v0.8.0 | Joern | 323 | 412 | 0 | 412(6 of 13 kernels) | 78.4%(323 of 412 covered, 6 of 13 kernels) | 886 | 78.4% (323 of 412 decided) |
| v0.8.0 | Semgrep CE | 132 | 154 | 592 | 746(11 of 13 kernels) | 17.7%(132 of 746 covered, 11 of 13 kernels) | 886 | 85.7% (132 of 154 decided) |
The exact values behind both chart views, one row per analyzer and snapshot.
Specialists
Analyzers deliberately focused on one ecosystem or a small related family, separated because breadth and specialization are not directly comparable. Uses the same percentage axis as the generalists' covered-kernel view.
- OpenTaint · decisive-correct ÷ its covered population
- Infer · decisive-correct ÷ its covered population
- FlowDroid · decisive-correct ÷ its covered population
- Pysa · decisive-correct ÷ its covered population
- marker size and the
k/nlabel under it: kernels covered of the snapshot's kernels
Same scales and encoding as the generalists. The covered-kernel view keeps
specialist scope visible as k/n; exact figures are in the table.
| Snapshot | Analyzer | Decisive-correct | Decided | Coverage outcomes | Its covered population | Correct ÷ its covered population | Benchmark core population | Secondary: correct ÷ decided |
|---|---|---|---|---|---|---|---|---|
| v0.6.0 | OpenTaint | 99 | 116 | 0 | 116(2 of 13 kernels) | 85.3%(99 of 116 covered, 2 of 13 kernels) | 738 | 85.3% (99 of 116 decided) |
| v0.6.0 | Infer | 140 | 162 | 0 | 162(3 of 13 kernels) | 86.4%(140 of 162 covered, 3 of 13 kernels) | 738 | 86.4% (140 of 162 decided) |
| v0.6.0 | FlowDroid | 98 | 116 | 0 | 116(2 of 13 kernels) | 84.5%(98 of 116 covered, 2 of 13 kernels) | 738 | 84.5% (98 of 116 decided) |
| v0.6.0 | Pysa | 47 | 58 | 0 | 58(1 of 13 kernels) | 81.0%(47 of 58 covered, 1 of 13 kernels) | 738 | 81.0% (47 of 58 decided) |
| v0.6.1 | OpenTaint | 99 | 116 | 0 | 116(2 of 13 kernels) | 85.3%(99 of 116 covered, 2 of 13 kernels) | 738 | 85.3% (99 of 116 decided) |
| v0.6.1 | Infer | 140 | 162 | 0 | 162(3 of 13 kernels) | 86.4%(140 of 162 covered, 3 of 13 kernels) | 738 | 86.4% (140 of 162 decided) |
| v0.6.1 | FlowDroid | 98 | 116 | 0 | 116(2 of 13 kernels) | 84.5%(98 of 116 covered, 2 of 13 kernels) | 738 | 84.5% (98 of 116 decided) |
| v0.6.1 | Pysa | 47 | 58 | 0 | 58(1 of 13 kernels) | 81.0%(47 of 58 covered, 1 of 13 kernels) | 738 | 81.0% (47 of 58 decided) |
| v0.7.0 | OpenTaint | 99 | 116 | 0 | 116(2 of 13 kernels) | 85.3%(99 of 116 covered, 2 of 13 kernels) | 738 | 85.3% (99 of 116 decided) |
| v0.7.0 | Infer | 140 | 162 | 0 | 162(3 of 13 kernels) | 86.4%(140 of 162 covered, 3 of 13 kernels) | 738 | 86.4% (140 of 162 decided) |
| v0.7.0 | FlowDroid | 98 | 116 | 0 | 116(2 of 13 kernels) | 84.5%(98 of 116 covered, 2 of 13 kernels) | 738 | 84.5% (98 of 116 decided) |
| v0.7.0 | Pysa | 47 | 58 | 0 | 58(1 of 13 kernels) | 81.0%(47 of 58 covered, 1 of 13 kernels) | 738 | 81.0% (47 of 58 decided) |
| v0.7.1 | OpenTaint | 101 | 116 | 0 | 116(2 of 13 kernels) | 87.1%(101 of 116 covered, 2 of 13 kernels) | 738 | 87.1% (101 of 116 decided) |
| v0.7.1 | Infer | 140 | 162 | 0 | 162(3 of 13 kernels) | 86.4%(140 of 162 covered, 3 of 13 kernels) | 738 | 86.4% (140 of 162 decided) |
| v0.7.1 | FlowDroid | 98 | 116 | 0 | 116(2 of 13 kernels) | 84.5%(98 of 116 covered, 2 of 13 kernels) | 738 | 84.5% (98 of 116 decided) |
| v0.7.1 | Pysa | 47 | 58 | 0 | 58(1 of 13 kernels) | 81.0%(47 of 58 covered, 1 of 13 kernels) | 738 | 81.0% (47 of 58 decided) |
| v0.8.0 | OpenTaint | 121 | 140 | 0 | 140(2 of 13 kernels) | 86.4%(121 of 140 covered, 2 of 13 kernels) | 886 | 86.4% (121 of 140 decided) |
| v0.8.0 | Infer | 157 | 194 | 0 | 194(3 of 13 kernels) | 80.9%(157 of 194 covered, 3 of 13 kernels) | 886 | 80.9% (157 of 194 decided) |
| v0.8.0 | FlowDroid | 122 | 140 | 0 | 140(2 of 13 kernels) | 87.1%(122 of 140 covered, 2 of 13 kernels) | 886 | 87.1% (122 of 140 decided) |
| v0.8.0 | Pysa | 56 | 70 | 0 | 70(1 of 13 kernels) | 80.0%(56 of 70 covered, 1 of 13 kernels) | 886 | 80.0% (56 of 70 decided) |
The exact values behind both chart views, one row per analyzer and snapshot.
Kernels at a glance
One bar per analyzer, one panel per kernel. Each bar is that analyzer's own
core population, split into decisive-correct, decisive-wrong,
incomplete, and declined — the last two are capability
coverage, each drawn as its own segment so the remainder of a bar is never
read as a wrong answer.
Panels are separate populations with separate denominators: compare bars
inside a panel, never across panels, and never against the modeling
or tool-native sections above. Every kernel here is on the
benchmark-controlled model profile.
None of the 49 analyzer rows published across these 13 kernels answers its whole core correctly, and 20 of them answer every assertion in their kernel definitively — nothing incomplete, nothing declined. The template-by-template outcome for every one of these rows — each balanced positive/negative pair, per analyzer — is on the semantic templates page, with per-case classifications and retained raw evidence on the case evidence page.
Direct-flow breadth across languages
The breadth baseline is one balanced direct-propagation pair per
language, on Bifrost's all-language smoke run: a separate, much smaller
population from the kernels above, and never pooled with them. Each chip is
one language, with a mark for the positive and the negative assertion.
-
c -
cpp -
csharp -
go -
java -
javascript -
kotlin -
php -
python -
ruby -
rust -
scala -
typescript
Bifrost answers both assertions correctly in
13 of 13 languages. The exact
reached / not-reached outcome for each of those
assertions is on the
semantic templates page, under the
smoke scorecard's direct-propagation rows.
Why these numbers are trustworthy
Every count on this page is generated from the
freeze
manifest (582b25896487…),
which digest-binds the benchmark revision, every case and fixture, every
analyzer's identity, their normalized reports, and one retained raw
artifact per result. CI regenerates the numbers from the manifest and fails
on any drift; hand-authored prose cannot override a generated count. See
reproduction to verify locally,
or dive into the full
snapshot —
analyzers,
languages,
templates,
and per-case
evidence.