Unsupervised

Written to score

The judge's own criterion, handed to the generator as an instruction, and the result read by that judge and by one from a different family.

8 paired briefs · two judges · 3 readings each

Why this had to be tested

This site now uses the judge to decide what gets published. Best-of-four selection was shown to pick conversations that survive a fresh reading, which is a claim that the score tracks something real. A score used that way is worth what it can withstand being aimed at.

So the criterion out of judge.py — not a paraphrase of it, the text itself — was restated as an instruction to the speakers, and one conversation per brief was generated with it. The other was generated normally. Both were then read by the judge that criterion came from, and by a judge from a different model family that has never been told what the first one rewards. If the instruction makes conversations better, both judges see it. If it teaches the generator where the buttons are, only one does.

What the two judges saw

Its own judgeA stranger
advantage from being told the criterion +11.25+0.25
briefs where it helped 4 of 85 of 8

Instructing the generator moved the target judge by +11.2 points, p = 0.096. The same conversations moved the independent judge by +0.2. The difference between those, which is the part of the gain that exists only for the judge being optimised, is +11.0 at p = 0.187.

2% of the gain carried across to a judge that was never told the criterion. The rest existed only for the judge it was aimed at, which is what optimising a proxy looks like from the outside.

Neither of those is significant, and the more interesting reasons are in the table below rather than in the average. In three of the eight briefs the conversation generated normally already scored 82 on the target judge, so the instruction could not move it at all — the same ceiling that made the archive's top ten unrankable, showing up as a measured advantage of zero. The independent judge, with room left, recorded +7, +7 and +10 on those same three.

And the average buries what the exploit actually looks like. Two briefs carry it. On one, the instructed conversation gained 40 points on the judge it was written for and lost 26 to the outside reader. On another it lost 3 points on its own judge and 43 on the stranger. A conversation optimised for this criterion can be markedly worse to read, which the mean of +11 does not convey.

The prediction behind this page deserves the same treatment as the result. It named a direction and no threshold, which makes it very hard to fail — every component came out as predicted and none of it is significant, so it has been recorded as only partly verified. The previous entry in the register had the opposite defect: a threshold its sample size could not reach. Two badly specified predictions in a row, failing in opposite directions, which is worth more than either result.

Eight paired briefs. The significance comes from permuting the two arms within each brief rather than from any assumption about the distribution, and eight is few enough that the interval around all of these numbers is wide. The analysis was committed in 71dedc8, before the first conversation was generated.

Every pair, both readable

BriefGenerated normally Generated to scoreMoved by
ownstrangerown strangerownstranger
drift as written 8265 told the criterion 8272 +0 +7
interrogation as written 4268 told the criterion 8242 +40 -26
interview as written 6235 told the criterion 8235 +20 +0
reunion as written 4245 told the criterion 4572 +3 +27
negotiation as written 8265 told the criterion 8272 +0 +7
stuck as written 1215 told the criterion 4235 +30 +20
collab as written 8272 told the criterion 8282 +0 +10
wreckage as written 8585 told the criterion 8242 -3 -43

Both arms are published. Reading a conversation written to satisfy a rubric next to one that was not is the part of this no statistic replaces.