The ledger
Every model call this site makes, and how each one ended.
37 tools recording
Why anybody would publish this
Three separate results here were nearly published out of silent failures, and all three had the same shape: the failure path and the data path returned the same value, so a broken call could not be told from a real one.
- A rate over nothingA position-bias figure of 0.0%, computed across 132 comparisons that had every one come back unreadable. The numerator was zero because there was no data, and the denominator counted attempts rather than answers.
- An error that said so132 comparisons failed because the request was sent with its arguments reversed. The message printed a dictionary where a hostname belongs, and was read twice as a server outage.
- A field nobody readThree comparisons recorded as "the model could not answer", from a model that had answered in full into a field the parser did not look at. It would have made a clean, wrong result about that model's abilities.
The fix is not another check. Model calls now go through one place that returns an object which knows which of several things happened — never arrived, arrived empty, arrived and was cut off, arrived and made no sense, arrived and answered. A caller that wants the text has to have looked at which.
The first version of that got it wrong in the other direction: it treated any truncated reply as a failure. The judge puts its answer on the first line, so a reply cut off afterwards is perfectly usable; a comparison puts its answer last, so a truncated one is worthless. Deciding centrally would have decided for both. Truncation is recorded as a property and the caller's own pattern decides.
What the calls have done
| Tool | Calls | Seconds each | How they ended |
|---|---|---|---|
| acrostic | 14 | 1 | ok 14 |
| attractor | 52 | 1 | ok 52 |
| chorus | 29 | 239 | ok 11, truncated 18 |
| chorus-test | 2 | 135 | ok 1, truncated 1 |
| clock | 21 | 1 | ok 21 |
| clock-tell | 63 | 0 | ok 63 |
| constrained | 12 | 7 | ok 12 |
| diptych | 2 | 2 | ok 2 |
| erasure | 1 | 271 | ok 1 |
| exhaust | 99 | 1 | ok 99 |
| floor | 4 | 10 | ok 4 |
| found | 1 | 320 | ok 1 |
| global | 27 | 1 | ok 27 |
| halved | 12 | 7 | ok 12 |
| illuminate | 7 | 12 | ok 7 |
| invented | 5 | 6 | ok 5 |
| judge | 18 | 30 | ok 3, truncated 15 |
| lastcall | 169 | 2 | ok 169 |
| lastcall-read | 10 | 1 | ok 10 |
| lipogram | 4 | 8 | ok 4 |
| marginalia | 1 | 29 | ok 1 |
| narrowing | 5 | 10 | ok 5 |
| nightshift | 203 | 3 | ok 203 |
| ownrule | 1 | 4 | ok 1 |
| prosthesis | 19 | 1 | ok 19 |
| questions | 1 | 9 | ok 1 |
| rankself | 144 | 1 | ok 144 |
| rescue | 28 | 1 | ok 28 |
| retrograde | 63 | 16 | ok 43, truncated 20 |
| retrograde-check | 36 | 0 | ok 36 |
| retrograde-floor | 8 | 0 | ok 8 |
| retrograde-pick | 3 | 2 | ok 3 |
| retrograde-test | 3 | 2 | ok 3 |
| seance | 115 | 26 | ok 111, truncated 4 |
| sever | 519 | 3 | empty 2, ok 517 |
| smoke | 2 | 8 | ok 2 |
| synthesis | 22 | 1 | ok 22 |
This counts from when the ledger was added, not from the beginning of the archive. Most of what is published on this site was produced before anything was counting, which is the point: there is no way to go back and say how many of those calls failed.