Skip to content

Analyzer profiles — v0.8.0

This exploratory view draws one small radar per analyzer over the benchmark-controlled core populations. It combines no axes into a score and makes no dominance claim.

The correctness axes use v0.8.0’s freshly rerun evidence. The latency axis uses this release’s separately qualified characterization data. Their denominators remain different, and no calculation on this page combines them.

Bifrost
Four-axis profile for Bifrost: speed 98.5% of the log scale, recall 55.8%, precision 100.0%, kernel coverage 100.0%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
111 ms median 886 timed invocations
recall
55.8% 247 of 443 positive assertions it covers
precision
100.0% 247 of 247 decided positives
coverage
100.0% 13 of 13 kernels, 886 assertions
CodeQL
Four-axis profile for CodeQL: speed 7.3% of the log scale, recall 63.5%, precision 83.7%, kernel coverage 84.6%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
60.4 s median 746 timed invocations
recall
63.5% 237 of 373 positive assertions it covers
precision
83.7% 237 of 283 decided positives
coverage
84.6% 11 of 13 kernels, 746 assertions
Semgrep CE
Four-axis profile for Semgrep CE: speed 69.9% of the log scale, recall 20.6%, precision 77.8%, kernel coverage 84.6%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
800 ms median 154 timed invocations
recall
20.6% 77 of 373 positive assertions it covers
precision
77.8% 77 of 99 decided positives
coverage
84.6% 11 of 13 kernels, 746 assertions
Joern
Four-axis profile for Joern: speed 42.1% of the log scale, recall 72.8%, precision 82.0%, kernel coverage 46.2%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
5.44 s median 412 timed invocations
recall
72.8% 150 of 206 positive assertions it covers
precision
82.0% 150 of 183 decided positives
coverage
46.2% 6 of 13 kernels, 412 assertions
Infer
Four-axis profile for Infer: speed 84.0% of the log scale, recall 62.9%, precision 98.4%, kernel coverage 23.1%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
302 ms median 194 timed invocations
recall
62.9% 61 of 97 positive assertions it covers
precision
98.4% 61 of 62 decided positives
coverage
23.1% 3 of 13 kernels, 194 assertions
FlowDroid
Four-axis profile for FlowDroid: speed 70.5% of the log scale, recall 87.1%, precision 87.1%, kernel coverage 15.4%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
767 ms median 140 timed invocations
recall
87.1% 61 of 70 positive assertions it covers
precision
87.1% 61 of 70 decided positives
coverage
15.4% 2 of 13 kernels, 140 assertions
OpenTaint
Four-axis profile for OpenTaint: speed 46.5% of the log scale, recall 88.6%, precision 84.9%, kernel coverage 15.4%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
4.01 s median 140 timed invocations
recall
88.6% 62 of 70 positive assertions it covers
precision
84.9% 62 of 73 decided positives
coverage
15.4% 2 of 13 kernels, 140 assertions
Pysa
Four-axis profile for Pysa: speed 52.9% of the log scale, recall 62.9%, precision 95.7%, kernel coverage 7.7%. The list beneath the shape carries every value and its denominator. The shape is a profile, not a score, and its area means nothing. speed recall precision coverage
speed
2.59 s median 70 timed invocations
recall
62.9% 22 of 35 positive assertions it covers
precision
95.7% 22 of 23 decided positives
coverage
7.7% 1 of 13 kernels, 70 assertions

These are profiles, not scores, and their areas are not comparable. A radar's enclosed area depends on the order the axes happen to be drawn in — swap two of them and the area changes without a single number changing — so nothing on this page computes one, and a shape that looks "bigger" is not a better analyzer. Nor are two shapes comparable as wholes: read one axis at a time, against the denominator printed under it. There is no combined score here or anywhere else on this site.

What each axis is over, since they are four different denominators.

  • speed — the median analyzer-invocation wall-clock over these same kernel populations, log-normalized and inverted so that faster reads as larger. The scale is absolute, not relative to the analyzers present: 100 ms maps to the outer ring and 100.0 s to the centre, the same three decades the latency chart's axis spans. Adding or removing an analyzer cannot move anyone else's mark. It is the only axis whose spacing is not linear, so equal distances along it are equal ratios of time. Every latency caveat on the latency page applies to it unchanged: one machine, one environment stamp, no repeated trials, per-invocation start-up costs inside the number.
  • recall — true positives over every positive-polarity assertion in the kernels the analyzer covers. Non-answers count against it: a case it answered inconclusive or declined as unsupported stays in the denominator, because it is a case the analyzer took on and did not resolve.
  • precision — true positives over decided positives, true plus false. Non-answers are absent from this denominator entirely. That asymmetry with recall is deliberate and it is the thing most easily misread on this page: the very same case can lower an analyzer's recall while being invisible to its precision, so an analyzer that decides little can hold a high precision and a low recall at once.
  • coverage — kernels covered of the snapshot's 13. It is on the picture precisely because the two rates beside it are over each tool's own covered slice: recall and precision from a one-kernel slice and from a thirteen-kernel slice are not the same exam, and this axis is how far apart those exams are. It is the same honesty note the cross-snapshot evolution chart's normalized view carries, and it binds harder here, because there the coverage was an annotation and here it is an axis of the same shape.

One population, so that the four numbers describe one exam. Every axis is computed over the benchmark-controlled core kernel populations — the same no-pooling filter the landing page's cards read — and not over the whole freeze. The latency figure here is therefore narrower than the whole-corpus median on the latency page, which also includes the modeling and tool-native runs; where the two differ, they differ because they are over different populations and both say which.

Latency and correctness are still never pooled. Placing them on one picture is not combining them: no arithmetic here crosses an axis boundary, no correctness outcome in this freeze was derived from or tie-broken by a timing value, and "correct but slow" and "fast but wrong" are as independently visible on these shapes as they are in the tables. If a reader wants a single number that ranks these analyzers, this page does not have one, and that is the point of it.

Correctness comes from the frozen results/results.json core population. Latency comes from the v0.8.0 timing evidence under docs/latency-tier.md. No correctness outcome is conditioned on or tie-broken by timing.