5 min read
Goodhart’s Law: The Metrics We Told You to Watch Can Be Gamed Too
Charles Goodhart was a British economist, and in 1975 he wrote a sentence about monetary policy that turned out to describe almost everything else too: when a measure becomes a target, it stops being a good measure. He meant it about interest rates and money supply. It applies just as cleanly to a community dashboard.
I’ve written twenty-five pieces for this site telling you what to watch. North Star Metric. Activation Rate. the retention curve. The viral coefficient, K. The 90-9-1 split. Every one of those is real, and every one of them is useful, and I stand behind all of it. But I never wrote the piece that says what happens the moment one of those numbers stops being something you watch and becomes something someone is told to hit. This is that piece. It’s overdue.
The gap between watching and hitting
A metric you watch is a diagnostic. It tells you something true about your community, and you go look at the actual humans behind the number to understand why it moved. A metric you hit is a target, and the moment it’s a target, the fastest way to move it is rarely the honest way.
A leaderboard like this one is the clearest visible example of a metric everyone can see and no one can resist optimizing for its own sake. Rank by reputation and post count is a genuinely useful signal, right up until someone starts posting for the rank instead of for the reason the rank was tracking in the first place.
Take Activation Rate, a number I told you to watch two pieces ago. The definition matters: the percentage of new members who complete one specific, meaningful action inside a defined window. Watched as a diagnostic, a dropping Activation Rate tells you something real is broken in the first 48 hours. Handed to a growth team as a quarterly target, the fastest fix is to make the “meaningful action” trivial. Redefine it from “posted in a Space” to “clicked a Space once.” The number climbs. The dashboard turns green. And the thing Activation Rate was built to predict, whether a new member actually sticks, has quietly stopped correlating with the number at all, because the action being measured no longer costs anything to complete.
Retention does this too, and more subtly. A team under pressure to flatten a decaying retention curve has an easy lever: email every inactive member a discount, a badge, a “we miss you” nudge timed to land right before the reporting window closes. Some of them click. Some of them technically return, in the sense that a login event fires. The curve on the dashboard flattens. Nothing about the actual health of the community changed. You’ve optimized the measurement instrument, not the thing it was measuring.
The same failure shows up outside anything I’ve written about directly. A moderation team measured on “tickets resolved per day” will resolve more tickets per day. Some of that is genuinely faster service. Some of it is closing threads the moment a reply lands, whether or not the actual problem got fixed, because closed-fast counts the same as closed-well on the only number anyone’s checking.
None of this is a lie, exactly
What makes Goodhart’s Law uncomfortable instead of just obvious is that nobody in any of these examples did anything you could call cheating. The activation threshold really was lowered by a well-meaning product decision. The re-engagement email really was sent with good intentions. The support agent really did close the ticket the customer said was fine. Each individual choice is defensible on its own. The law isn’t about bad actors gaming a system. It’s about what happens to any number, made by anyone, the instant it stops being an observation and starts being someone’s job to move.
That’s also why “just don’t measure anything” isn’t the answer, and I’m not writing that piece either. A community with no metrics is flying blind, and everything else on this site assumes you’re watching something. The fix isn’t fewer numbers. It’s being honest about which side of the line each number is standing on.
Where the line actually is
A metric stays a diagnostic as long as it’s something a team looks at together to ask “why,” not something an individual is handed as a personal number to hit. North Star Metric, the very first piece I wrote for this project, argued for exactly one number worth obsessing over. I stand by that, but I’d add the caveat now that the piece was missing: obsessing over a number as a team, in the sense of building your whole roadmap around understanding what moves it, is different from handing that same number to one person as a quota. The first produces insight. The second produces gaming, reliably, regardless of who the person is or how well-intentioned they are.
A cheap test: watch what happens the moment the number stops moving in the direction you want. A diagnostic metric gets investigated. Someone asks what changed, looks at individual member behavior, tries to find the real cause. A target metric gets defended, or explained away, or quietly redefined until it moves again. If your team’s instinct when a number goes the wrong way is to figure out how to make it go the right way rather than to figure out what it’s actually telling you, the number has already crossed over into target territory, whether anyone announced that or not.
A structural fix, not a willpower one
Telling people to try harder not to game a number doesn’t work, because nobody experiences what they’re doing as gaming. Everyone in the earlier examples believed their own decision was reasonable in the moment. The fix that actually holds is structural, not motivational: keep the number a team looks at, and the number a person’s compensation or standing depends on, physically separate. Different meeting, different document, ideally a different cadence entirely.
Put concretely: a community team can watch Activation Rate weekly, as a group, arguing about what it means. That’s healthy. The moment Activation Rate also appears on someone’s individual performance review as a target with a number attached, the same metric is now doing two incompatible jobs at once, and the review-driven behavior will win, because it’s the one with consequences attached to it. Split them. Let the team’s diagnostic dashboard be genuinely, uncomfortably honest, precisely because nobody’s paycheck depends on what it says this week.
Sit with that for a minute
Every dashboard eventually gets checked by someone whose job depends on what it says. That’s not a design flaw you can engineer around. It’s the actual mechanism Goodhart described, showing up again in a system he never heard of, because it isn’t really about economics or communities at all. It’s about what happens to any number the instant a person’s outcome is tied to it.
The honest response isn’t to stop measuring your community. It’s to keep the measuring and the accountability in separate rooms, on purpose, and to notice the exact moment someone tries to move them into the same one.
Related reading