Four tools were asked the same questions over the same repositories and scored against the same gold-derived expected files. The report publishes the means. This page publishes the run: every question it scored, what each side answered, and where in that answer the first expected file actually sat. Filter it, sort it, and open a row to check the arithmetic yourself.
Generated from retrieval result retrieval-1787589820180 (schema v3) at 2026-08-24T16:43:40.181355Z. Every figure below is rendered from that document; none is typed.
A side with no measurement for a question reads n/a and is excluded from every mean on this page — never folded in as 0.000. A side that ran and returned no expected file anywhere is a real zero and says so in different words. The two are different claims and are printed differently, here and in the report.
Rows arrive in the run's own order — repository, then question id — and never by score. Every filter and every sort applies to the four sides identically, and no ordering is offered that would surface one side's good rows before another's. A question this project loses looks exactly like one it wins.
The two graph tools' names differ by two letters, so neither is ever written bare in a column header, a filter or a verdict: each column is labelled in full, with the third-party tool marked as such. A result favouring any side is an outcome of the measurement, not an error in it.
This measures which files a tool puts in front of you — not what an agent then does with them. The questions and the gold facts they are scored against were written by this project, which is the largest thing to discount for, and each figure is one run rather than a distribution. The full report states the methodology, the caveats and every skip.
One row per question. The four side columns show whichever metric is selected below; the verdict names every side that reached an expected file first, ties included. Click any row — or focus it and press Enter — to open the gold file set and each side's own ranked answer beneath the table.
A figure marked like this is a side that reached one of the expected files somewhere in its own answer. A figure marked like this is a side that ran and reached none of them — a measured zero, and a real result. n/a is a side with no measurement for that question at all: excluded from every mean, never counted as zero.
The same aggregates the report publishes, so this page stands on its own. The headline pool excludes the negative-control questions, which are reported separately and never averaged into it; the per-category and per-repository tables slice the whole scored set, negative controls included. Where a column shows n/a, that side is not in this run — again, not a zero.
Everything above is rendered in your browser from a JSON block embedded in this page's own source, which is generated from the committed result document named at the top. Nothing is fetched, from anywhere: open this file from a local clone with the network off and it behaves identically. Nothing on this page is typed by hand — a figure that moves in the next measurement moves here by regeneration or not at all.