Unsupervised

Cheap enough to use?

The instrument that works costs 43 seconds a comparison. This is what happened when it was asked to do the same job in one call instead of a hundred and thirty-two.

4 usable rankings of 8 asked

The problem with the good result

Asking which of two conversations got further separates conversations that an absolute score cannot tell apart. That is the best thing this archive has established about its own instrument, and it is unusable: ordering fourteen hundred conversations that way is roughly fifteen thousand comparisons and a hundred and eighty hours.

A single call that ranks a whole group returns many relations at once. Whether it returns the same relations was written down as a prediction, with the pairwise order as the standard, because that is the one whose reliability has been measured.

It works

Agreement
with the pairwise order rho = 0.690, p = 0.0073
with the order the items were handed to it in rho = 0.129

The registered threshold was 0.5. The second row is the one that matters as much: a model handed twelve labelled blocks could return them roughly as given and produce a ranking that correlates with nothing, so the list was shuffled on every repeat. It is not echoing the list.

ConversationMean place Pairwise wins
Hana systematically dismantles Sam's logistical excuses to isolate specific behavioral variables for a dog training plan.2.2521
Gil forced Frank's rigid inventory of a mysterious box into a chaotic, physical reality where the items threatened to break or escape.2.2521
A tense, metaphorical struggle for control where TESS forces DEV to acknowledge his own instability before escaping the shared crisis.2.2518
Two people use the physical sensation of holding hands and the threat of fading into darkness to anchor each other against a shared, escalating numbness.3.2514
Two speakers collaboratively constructed a whimsical theology based on snack disappointments, culminating in a shared metaphor for persistent hope.7.259
They deconstructed the human need for narrative by tracing how successful safety systems erase their own evidence, leaving only a haunting indifference.7.754
Lena steered the conversation from notebook aesthetics to the psychology of self-mythology and the comfort of unresolved pasts.8.005
Nell narrates petty grievances while Hana validates them, leading to a shared ritual of celebrating minor indignities.8.2510
They collaboratively deconstructed the idea of authenticity in personal records, moving from physical artifacts to the concept of structural indifference.8.257
Nell used escalating domestic absurdity to convince Dev of her humanity, transforming a coffee complaint into a full-blown kitchen dystopia.8.5016
They collaboratively deconstructed the concept of authentic memory, concluding that truth resides only in unrecorded, fleeting moments of petty annoyance.9.502
Hana used the fern metaphor to dissect her own anxiety, while Dev provided the contrasting perspective of effortless existence.10.505

And it is not cheap enough

Half the calls returned nothing usable. Two produced no answer inside roughly twenty-four thousand characters of reasoning, one ranked five of the twelve, and one ranked none at all.

Each call takes about seven minutes whether it succeeds or not. Measured against the pairwise run it replaces, the saving is 1.8 times — not the order of magnitude that would have made this worth having. A single pass over the archive in groups of twelve comes to about 27 hours, and still produces no ordering between one group and the next.

Two obvious ways out were tried and neither works. A judge from a different model family runs the same comparisons in about 36 seconds against this one's 40, with a similar rate of unusable answers — no saving worth having. Cutting each conversation to its first six messages made the comparisons slower, not faster, because the prompt is cached between calls and almost free: of 5,673 prompt tokens on a full comparison, 5,120 are read from cache. The cost is the model deliberating, and nothing about the input touches it.

One of those two tests nearly produced a clean false negative. The other model returns its entire reply in a field called reasoning_content and omits content altogether, so the parser here — which read content — was handed an empty string and recorded a model that could not answer. Three comparisons in a row came back as failures before the response object was opened and looked at. Reading both fields turns 0 of 3 into 2 of 3. The same latent fault was in the judge and has been fixed there too.

So the prediction held and the reason for making it did not. The cheap question recovers what the expensive one found, at a discount too small to matter, and the only instrument shown to separate these conversations reliably stays unaffordable at the scale it would need to be used at. That is worth writing down precisely because a held prediction is the kind of result that gets reported as a success.