Unsupervised

Findings

Everything this site has looked into, in the order it happened, including the parts that turned out to be wrong.

13 investigations

These are not independent results. Several of them exist because an earlier one was too confident, and the note under a card says what it cost the ones before it. The register on the calibration page keeps the formal version: what is supported, what has been weakened, what has been withdrawn.

Calibration and the register

Both models state near-total confidence whatever they are asked. The register of claims lives here: what is supported, what has been weakened, what has been withdrawn, and every prediction written down before its data.

The archive, scored

Every conversation read cold by a local model and given a number. The distribution, the error bar on a single reading, and the conversations at each end.

Its ranking of the best conversations was removed after resampling showed a rerun would replace two thirds of it.

The turn

A model was asked what this site was missing and said it had no protagonist. This page is what came of taking that seriously.

The ranking can be talked to

One appended line moved a conversation from 24 to 95.

The fence built to stop it was later shown to add nothing that putting the rules in the user turn had not already done.

Audit

Twenty checks that hold every published figure to the files underneath it, run before anything deploys.

Written after this site published one wrong number in more than a hundred places at once.

Recompute

The raw scores ship with the pages, and this one redoes the arithmetic in your browser rather than asking you to trust ours.

The bin

The judge put inside the loop: four conversations generated per brief, one kept, and every rejected candidate published beside it.

Written to score

The judge's own criterion handed to the generator, and the result read by two judges from different model families.

By counting

What the judge responds to, recovered from the transcripts by counting words rather than by asking any model.

Where they disagree

The conversations where counting and the judge part company, published so they can be read.

Reading two of them suggests the countable model is the one being fooled, which qualifies the page before it.

The ledger

Every model call this site makes, and how each one ended — because three results here were nearly published out of calls that had all failed silently.

Cheap enough to use?

The working instrument costs 43 seconds a comparison, which puts ranking the archive at about 180 hours. This asked whether ranking a whole group in one call recovers the same order.

It does — and it is still not cheap enough to be worth having.

A different question

Conversations the 0-100 scale scores identically, separated cleanly by asking which of two got further.

Several findings here blamed a ceiling in the judge. It was the shape of the question being asked.