Unsupervised

By counting

What the judge is responding to, worked out from the transcripts without asking it anything.

613 conversations at 82 · 375 at 12 · no model calls

Seven rounds on the instrument, none on the archive

Everything published here so far has been about the judge — how stable it is, what it can rank, whether it can be gamed, whether its scale has room. The archive underneath is fourteen hundred conversations and none of that work has looked at them.

So: take every conversation the judge scored 82 and every one it scored 12, its two commonest verdicts, and try to tell them apart by counting. Words per message, how much each speaker reuses the other's vocabulary, how fast new words arrive, question marks. Nothing that needs a model, nothing a reader could not check by hand.

It works, and it only just clears the line

Fitted on half the conversations and scored on the half it never saw, plain counting tells the two apart 70.4% of the time. The registered threshold was 70%, so the prediction holds — by four tenths of a point.

That margin deserves a harder look than it usually gets. Over 200 splits the mean sits 2.5 standard errors above the line, so it is reliably above 70 for this archive. But only 62% of individual splits clear 70 on their own, and the worst lands at 62.9. A single run of this experiment would have missed its own threshold more than a third of the time. The number to carry away is that counting recovers most of the distinction, not that it cleared a bar.

And thirty per cent is not recovered. Whatever else the judge is doing, these eight counts do not capture it.

That remainder turns out to matter more than the seventy. Scoring the whole archive both ways and reading the conversations where the two part company hardest, the countable model fails in a consistent direction: it reads two speakers using each other's words as circling when it can be the sound of an actual argument, and reads a constant supply of new vocabulary as movement when it can be one speaker changing the subject to avoid the other. Those disagreements are published, conversations and all.

What is actually different

CountedScored 82Scored 12 At 82
length
words per message
85.152 89.450 lower
novelty
how much of each message is vocabulary the conversation has not used before
0.332 0.304 higher
echo
how much of each message is vocabulary the other speaker just used
0.406 0.455 lower
growth
whether messages get longer or shorter as it goes on
-0.148 -0.060 lower
turns
messages in the conversation
24.261 24.147 higher
spread
how uneven the message lengths are
0.276 0.267 higher
variety
distinct words as a share of all words
0.316 0.288 higher
questions
question marks per message
0.608 0.887 lower

The conversations this judge scores highly introduce more vocabulary the conversation has not used, reuse the other speaker's words less, and ask fewer questions. That is a coherent reading of its own criterion — circling politely looks like echoing, and getting somewhere looks like new words arriving — arrived at without asking it anything.

The table is raw group averages, which is the only form worth showing. The fitted coefficients disagree with it about message length: length carries a positive weight in the model while the high-scoring conversations are marginally the shorter ones. That is what correlated features do, and it is the reason no coefficient is printed here. A weight in a model with eight related inputs is not the effect of its feature.