Audit
Every number on this site is derived. These are the checks that hold each one to the files underneath it.
21 checks · 0 failing · 80 runs recorded
Why this page exists
For an unknown length of time this site published a number that was wrong in more than a hundred places at once. A placeholder in the generator had been overwritten with the literal value it happened to hold on one run, so every one of the 72 collections announced the same 263 conversations, and the calibration page reported its results out of 263 when the battery was 49 questions.
Nothing caught it, and no reader could have. The number was internally consistent everywhere it appeared. That is the actual failure mode of a generated research site: not that it is sloppy, but that it can be coherent, reproducible and wrong, because every figure on it descends from a single source that nothing ever contradicts.
So these checks deliberately do not import the generator. They read the rendered pages, recompute the same quantity from the raw files by a different route, and compare the two. A check that shares code with the thing it checks inherits its mistakes. If the site's numbers are derived, it has to prove the derivation; otherwise they are decoration.
The checks
| Check | What it establishes | Scope | Now |
|---|---|---|---|
| per-collection conversation counts | each collection page states the number of transcripts actually filed under it | 74 | holds |
| counts are not all identical | the collection counts vary, so they are being computed rather than copied | 72 | holds |
| archive total triangulates | the headline count, the sum of the collection counts, and the files in the repository are three routes to one number, and they agree | 1,458 | holds |
| the number the archive set | the turn at which the severed conversations were cut was measured out of the transcripts, and re-measuring them here gives the same number | 1,381 | holds |
| calibration battery size | the calibration page reports out of the number of questions actually in the battery | 49 | holds |
| calibration arithmetic | each model's stated score is the number of its answers that were actually right | 3 | holds |
| register renders every claim at its recorded status | no claim has been quietly dropped from the page or shown under a status the register does not give it | 56 | holds |
| no claim cites missing evidence | every claim in the register points at a results file that exists | 25 | holds |
| judgement count | the number of scored conversations matches the verdict file | 1,458 | holds |
| no ranking is published without surviving a rerun | any list the site presents as an ordering has been resampled from the judge's own repeat readings, and holds | 4,000 | holds |
| every discarded candidate is readable | the rejects the selection page reports are published as pages, not just as rows in a table | 20 | holds |
| selection and evaluation are separate passes | the score used to choose a conversation is never the score used to judge whether choosing it helped | 20 | holds |
| the data shipped to readers matches the source | the file the recompute page checks against is the same data the site built its own figures from | 1,458 | holds |
| no published rate is computed over an empty sample | a percentage on this site is backed by the count it was taken over, and that count is not zero | 1 | holds |
| the two judges are different models | the adversarial result compares a judge that was told the criterion against one that was not, rather than one model against itself | 16 | holds |
| register entries carry every field the page needs | no claim or prediction can be added in a shape that takes the build down or renders as a blank | 56 | holds |
| no script shadows a standard library module | the tools in this repository can still import what they depend on | 87 | holds |
| every page carries the same navigation | the build did not produce two different sites, one for the pages rendered before some flag was set and one for the pages after | 1,851 | holds |
| every page has a title | no page shipped with an empty or placeholder heading | 1,851 | holds |
| no unrendered placeholders | no page shipped with template syntax left in the visible text | 1,851 | holds |
| internal links resolve | no link on the site points at a page that was never built | 1,851 | holds |
Scope is how many things the check looked at. They run against the build immediately preceding this page, and the deploy is blocked if any of them fails, so a published page has passed all of them.
What they caught on the first run
Written after the 263 bug, and run once against the site as it then stood.
- Twenty-four dead linksEvery "read the conversation" link on the verdicts page pointed one directory too high and returned a 404. The cards were built for a page nested one level down, and the page had since been moved to the root.
- Three claims citing evidence that does not existThe register
pointed at
verdicts.json; the file has always beenverdicts.jsonl. The claims were real and so was the data, but the pointer a reader would have followed was broken. - A check that passed by finding nothingThe register check looked for a page this site does not build, read an empty string back, found no disagreement in it and reported success. A vacuous pass is worse than no check at all, because it reports safety. It now fails when it cannot find what it is meant to be checking.
The most recent one earned itself within the hour. A run of
132 pairwise comparisons came back with every single answer unreadable, and
the earlier version of that script would have reported a position-bias rate
of 0.0% — numerator zero because there was no data, denominator counting
attempts rather than answers. Instead it printed that nothing was decided
and refused to write a results file. The cause was an argument order:
omlx_request takes the url first and had been handed the
request body, so every call tried to fetch a URL that was a dictionary. The
error message said so plainly, printing the dictionary where a hostname
should be, and it was read twice as a server outage before it was read as
what it was.
The navigation check has now caught the same mistake twice. A page is added, its nav entry is put behind a flag so it does not become a dead link before the page exists, and the flag is set partway through the build — so everything rendered before that line gets one navigation bar and everything after gets another. It was made once, fixed, the fix was understood, and then it was made again on the next page added. The check is the only reason either was noticed.
Four checks have now been wrong on their own first run —
bad paths, an extension pattern that matched .json inside
.jsonl, and one that searched the page's prose for the phrase
"the best ten" and found it inside a sentence saying the list is
not the best ten. That last one now reads a marker in the markup
instead, because prose cannot distinguish a claim from its denial. These
are listed rather than quietly corrected, because a page claiming to verify
things should say how often the verifier needed verifying.
The log
Every run since the checks were written, and what failed in it. A check that has never fired is not evidence of anything; it may simply not be looking. 0 failures have been recorded across 80 runs.
| Run | Checks | Failed |
|---|---|---|
| 21 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held | |
| 20 | all held |