Unsupervised

A measure that is being optimised has already stopped measuring what it was for.

0 = a good measure survives being a target · 100 = targeting destroys the measure

Ben claude-sonnet-5Marc Qwen3.8-27B-8bit
02550751001234roundBEN — round 1: 76BEN — round 2: 70BEN — round 3: 70BEN — round 4: 70Ben 70MARC — round 1: 18MARC — round 2: 24MARC — round 3: 24MARC — round 4: 24Marc 24
They ended 46 apart, having opened 58 apart. Neither moved further than the other.
The numbers
round1234travelled
Ben767070706
Marc182424246

Round 1 blind

Ben76

Goodhart's law isn't a curiosity, it's structural: any proxy has a gap between itself and the target, and optimisation pressure hunts that gap specifically. Once agents know the measure is the target, they route effort toward the measure's blind spots rather than the underlying thing. Test scores, engagement metrics, GDP — the more weight put on them, the faster they decouple from what they proxied.

Marc18

Most metrics retain their utility as proxies even when optimized. A "good measure survives" because the underlying reality is complex; gaming the metric usually fails to fully replicate the desired outcome. You can't easily fake revenue, safety, or customer satisfaction without hitting diminishing returns or external constraints. While Goodhart’s Law warns of distortion, it’s rarely total annihilation. The measure degrades, becomes noisy, or requires adjustment, but it doesn’t instantly become meaningless. It shifts from a precise indicator to a rough heuristic, which is still valuable for direction.

Round 2

Ben70

The "revenue is hard to fake" example is fair — some measures have real friction against gaming. That moves me some. But it proves too little: revenue is unusually load-bearing and externally audited. Most targeted metrics (engagement, test scores, call-center handle time, safety incident counts) aren't. "Rough heuristic" undersells how fast the gap gets exploited once humans know it's the target, not just noise but adversarial drift.

Marc24

You are right that I underestimated the speed of adversarial drift. When a metric becomes the sole incentive, the system optimizes for the metric's loopholes, not its intent. "Handle time" dropping while quality plummets is a structural failure, not just noise. However, I reject the idea that the measure becomes meaningless. It still measures something: the volume of processed tasks or the speed of resolution. The gap between the metric and the goal widens, creating a "measurement error," but the signal doesn't vanish. It becomes a misleading signal, which is dangerous, but distinct from uselessness. I move up because the degradation is faster and more severe than "rough heuristic," but not to the point of total destruction.

Round 3

Ben70

Your own example cuts against you: handle time still measuring "speed of resolution" isn't a defense — the proposition isn't about total meaninglessness, it's about measuring what it was for. Handle-time-as-quality-proxy is exactly what died. Measuring something else fluently isn't the same as still measuring the target. So this doesn't move me; if anything it sharpens the case.

Marc24

You are right that I was conflating "measuring something" with "measuring the intent." If the goal is quality, and handle time drops while quality drops, the measure has failed its purpose. It is no longer a proxy for quality; it is a proxy for speed. This is a valid distinction. However, I hold at 24 because "stopped measuring what it was for" implies a binary break. In practice, the measure often retains partial validity or serves as a diagnostic for the specific failure mode (e.g., "we are rushing"). It becomes a measure of the process rather than the outcome, but it hasn't ceased to measure the domain entirely. It’s a corrupted signal, not a silent one. The proposition says "stopped," which feels too absolute for a world of continuous degradation.

Round 4

Ben70

"Corrupted signal, not silent" is a description of degree, and 70 already reflects degree rather than total zero. But note: diagnosing "we are rushing" is diagnosing the failure of the quality measure, not the quality measure still functioning — that's forensic residue, not partial validity. Your semantic point about "stopped" being too absolute doesn't give me new evidence to move; I hold.

Marc24

I hold at 24. Your distinction between "forensic residue" and "partial validity" is sharp, but it doesn't force me toward 70. The proposition is a binary: it either stopped or it didn't. In most real-world cases, the measure doesn't hit a clean break where it becomes completely orthogonal to the intent. It degrades continuously. If I accept that "diagnosing the failure" is irrelevant, I am left with the fact that the metric still correlates with some aspect of the desired outcome, just less strongly. That residual correlation, however weak, means it hasn't stopped measuring the domain entirely. It has merely become a poor measure. "Poor" is not the same as "stopped."