Analyzer profiles — v0.8.0
This exploratory view draws one small radar per analyzer over the benchmark-controlled core populations. It combines no axes into a score and makes no dominance claim.
The correctness axes use v0.8.0’s freshly rerun evidence. The latency axis uses this release’s separately qualified characterization data. Their denominators remain different, and no calculation on this page combines them.
- speed
- 111 ms median 886 timed invocations
- recall
- 55.8% 247 of 443 positive assertions it covers
- precision
- 100.0% 247 of 247 decided positives
- coverage
- 100.0% 13 of 13 kernels, 886 assertions
- speed
- 60.4 s median 746 timed invocations
- recall
- 63.5% 237 of 373 positive assertions it covers
- precision
- 83.7% 237 of 283 decided positives
- coverage
- 84.6% 11 of 13 kernels, 746 assertions
- speed
- 800 ms median 154 timed invocations
- recall
- 20.6% 77 of 373 positive assertions it covers
- precision
- 77.8% 77 of 99 decided positives
- coverage
- 84.6% 11 of 13 kernels, 746 assertions
- speed
- 5.44 s median 412 timed invocations
- recall
- 72.8% 150 of 206 positive assertions it covers
- precision
- 82.0% 150 of 183 decided positives
- coverage
- 46.2% 6 of 13 kernels, 412 assertions
- speed
- 302 ms median 194 timed invocations
- recall
- 62.9% 61 of 97 positive assertions it covers
- precision
- 98.4% 61 of 62 decided positives
- coverage
- 23.1% 3 of 13 kernels, 194 assertions
- speed
- 767 ms median 140 timed invocations
- recall
- 87.1% 61 of 70 positive assertions it covers
- precision
- 87.1% 61 of 70 decided positives
- coverage
- 15.4% 2 of 13 kernels, 140 assertions
- speed
- 4.01 s median 140 timed invocations
- recall
- 88.6% 62 of 70 positive assertions it covers
- precision
- 84.9% 62 of 73 decided positives
- coverage
- 15.4% 2 of 13 kernels, 140 assertions
- speed
- 2.59 s median 70 timed invocations
- recall
- 62.9% 22 of 35 positive assertions it covers
- precision
- 95.7% 22 of 23 decided positives
- coverage
- 7.7% 1 of 13 kernels, 70 assertions
These are profiles, not scores, and their areas are not comparable. A radar's enclosed area depends on the order the axes happen to be drawn in — swap two of them and the area changes without a single number changing — so nothing on this page computes one, and a shape that looks "bigger" is not a better analyzer. Nor are two shapes comparable as wholes: read one axis at a time, against the denominator printed under it. There is no combined score here or anywhere else on this site.
What each axis is over, since they are four different denominators.
- speed — the median analyzer-invocation wall-clock over these same kernel populations, log-normalized and inverted so that faster reads as larger. The scale is absolute, not relative to the analyzers present: 100 ms maps to the outer ring and 100.0 s to the centre, the same three decades the latency chart's axis spans. Adding or removing an analyzer cannot move anyone else's mark. It is the only axis whose spacing is not linear, so equal distances along it are equal ratios of time. Every latency caveat on the latency page applies to it unchanged: one machine, one environment stamp, no repeated trials, per-invocation start-up costs inside the number.
- recall — true positives over every
positive-polarity assertion in the kernels the analyzer covers.
Non-answers count against it: a case it answered
inconclusiveor declined asunsupportedstays in the denominator, because it is a case the analyzer took on and did not resolve. - precision — true positives over decided positives, true plus false. Non-answers are absent from this denominator entirely. That asymmetry with recall is deliberate and it is the thing most easily misread on this page: the very same case can lower an analyzer's recall while being invisible to its precision, so an analyzer that decides little can hold a high precision and a low recall at once.
- coverage — kernels covered of the snapshot's 13. It is on the picture precisely because the two rates beside it are over each tool's own covered slice: recall and precision from a one-kernel slice and from a thirteen-kernel slice are not the same exam, and this axis is how far apart those exams are. It is the same honesty note the cross-snapshot evolution chart's normalized view carries, and it binds harder here, because there the coverage was an annotation and here it is an axis of the same shape.
One population, so that the four numbers describe one exam.
Every axis is computed over the benchmark-controlled core kernel
populations — the same no-pooling filter the landing page's cards read — and
not over the whole freeze. The latency figure here is therefore narrower than
the whole-corpus median on the latency page, which also includes the modeling
and tool-native runs; where the two differ, they differ because they are over
different populations and both say which.
Latency and correctness are still never pooled. Placing them on one picture is not combining them: no arithmetic here crosses an axis boundary, no correctness outcome in this freeze was derived from or tie-broken by a timing value, and "correct but slow" and "fast but wrong" are as independently visible on these shapes as they are in the tables. If a reader wants a single number that ranks these analyzers, this page does not have one, and that is the point of it.
Where these numbers come from
Section titled “Where these numbers come from”Correctness comes from the frozen results/results.json core population.
Latency comes from the v0.8.0 timing evidence under
docs/latency-tier.md.
No correctness outcome is conditioned on or tie-broken by timing.