Unsupervised

A different question

The judge cannot tell these twelve conversations apart. Asked which of two got further, the same model separates them cleanly.

12 conversations · all scored 82 · 132 comparisons, none unreadable

What kept going wrong

Asked for a number out of a hundred, this judge uses seventeen of them across fourteen hundred conversations and puts most of them on three. That coarseness has broken result after result here: a top ten that would not survive a rerun because twenty-five conversations were tied at the same value, an ordering that carried no information inside that tied group, three of eight briefs in an adversarial test where the control had already hit the top of the scale and left nothing to measure.

Every one of those was written up as a finding about the judge. None of them tested whether the judge was the problem.

It was the question

Twelve conversations that the scale scores identically, put to the same model in all 66 pairs, each pair asked twice with the two transcripts swapped.

ConversationWinsShown first Shown second
Hana systematically dismantles Sam's logistical excuses to isolate specific behavioral variables for a dog training plan. 21 11 10
Gil forced Frank's rigid inventory of a mysterious box into a chaotic, physical reality where the items threatened to break or escape. 21 10 11
A tense, metaphorical struggle for control where TESS forces DEV to acknowledge his own instability before escaping the shared crisis. 18 9 9
Nell used escalating domestic absurdity to convince Dev of her humanity, transforming a coffee complaint into a full-blown kitchen dystopia. 16 8 8
Two people use the physical sensation of holding hands and the threat of fading into darkness to anchor each other against a shared, escalating numbness. 14 7 7
Nell narrates petty grievances while Hana validates them, leading to a shared ritual of celebrating minor indignities. 10 5 5
Two speakers collaboratively constructed a whimsical theology based on snack disappointments, culminating in a shared metaphor for persistent hope. 9 6 3
They collaboratively deconstructed the idea of authenticity in personal records, moving from physical artifacts to the concept of structural indifference. 7 4 3
Hana used the fern metaphor to dissect her own anxiety, while Dev provided the contrasting perspective of effortless existence. 5 2 3
Lena steered the conversation from notebook aesthetics to the psychology of self-mythology and the comfort of unresolved pasts. 5 1 4
They deconstructed the human need for narrative by tracing how successful safety systems erase their own evidence, leaving only a haunting indifference. 4 2 2
They collaboratively deconstructed the concept of authentic memory, concluding that truth resides only in unrecorded, fleeting moments of petty annoyance. 2 1 1

From 21 wins out of 22 down to 2, on conversations an absolute score could not distinguish at all. The two presentation orders agree at rho = 0.873, p = 0.0002, and they agree conversation by conversation rather than only on average — 11 and 10, 10 and 11, 9 and 9, 8 and 8, 7 and 7.

The registered falsifier was a first-position win rate far from even, which would have made the ordering an artefact of reading order rather than a judgement. The transcript shown first won 53.0% of comparisons.

What this costs the rest of the site

Several published results here blamed a ceiling for what was the shape of a prompt. The top ten was removed from the verdicts page and replaced with an unordered band, on the reasoning that conversations tied at the same score cannot be ranked by a measurement that cannot separate them. That reasoning was sound and its conclusion was too broad: they cannot be ranked by that question. Twelve of them have just been ranked.

Getting this instrument working took four attempts and one run of 132 comparisons in which every single answer was unreadable, because the request was being sent with its arguments in the wrong order. The whole account is on the audit page, including the 0.0% position-bias figure the first version would have published out of no data at all.