Quarter one: the KPI describes the work. Quarter two: the work is redesigned to produce the KPI. Goodhart arrives somewhere in between.
I tweeted these lines earlier in the week after having spent most of the day reviewing some strategy slides for an upcoming meeting. That thought stuck and here we are with the long version.
Goodhart's law is usually stated like this: when a measure becomes a target, it ceases to be a good measure. In most companies it has the status of folklore. Something wise to nod at in an operating review, filed next to “culture eats strategy for breakfast.”
No villain required
The folklore version sounds like a warning about cynical employees gaming badly designed incentives. That version does happen. The most infamous example perhaps is of Wells Fargo where it paid bonuses on cross-sold accounts, and its employees opened millions of accounts that customers may never have authorized.
But the usual version needs no conspiracy and very little cynicism. A KPI begins as an honest description. The company notices that its best support teams resolve tickets quickly. Resolution time becomes the measure. And once a number decides what gets reviewed, funded and rewarded, the work starts arranging itself around the number. The sales team measured on number of meetings booked fills the calendar with conversations that go nowhere. People respond, quite reasonably, to whatever the company has made consequential.
Up to a point, that's the idea. It's what a KPI is for: to direct attention and change the work. Goodhart arrives at the moment improving the number becomes easier than improving the thing it was meant to describe. That support team can cut resolution time by solving problems faster. It can also cut it by transferring the hard cases, or by closing tickets early and letting the customer open new ones tomorrow. The dashboard records all three as faster resolutions. In quarter one the metric observed the work. By quarter two it instructs the work.
Stakes and speed
Charlie Munger compressed a lifetime of this into one line: "Show me the incentive and I will show you the outcome." A number nobody is paid on decays slowly. Attach compensation to it and the redesign of the work begins immediately.
How fast a metric corrupts also depends on who is doing the optimizing. In the past, gaming a metric took effort, coordination, and a certain willingness to be cynical, so the corruption took quarters to arrive. Nobody designed that lag, but it protected companies all the same.
AI and agent driven workflows remove all three. An agent looks like a diligent employee that does the task without complaining. Underneath, it's an optimizer. It pursues the goal you wrote down without any of the friction that used to protect you. It doesn't get tired. It doesn't get embarrassed. It feels no loyalty to the spirit of the goal, only to its letter.
Google DeepMind once trained an agent to stack a red Lego block on a blue one, rewarding it on the height of the red block's bottom face. The agent flipped the block over and collected the reward. It didn't misunderstand the goal. The goal was just incomplete, the way every written goal is. Reduce handling time, but don't make the customer call back next week. Process more claims, but don't shunt the ambiguous ones into someone else's queue. A person fills those gaps without being told. An agent gets only what you wrote down. Every measure wears out under pressure, and that pressure just got cheap.
There is no holdout.
The machine learning field has this same disease and calls it overfitting. It never cured it. What it built instead is a defense: the held-out test set. A slice of data locked away where the optimizer never sees it, used only to check the model's work. Never train on the test set is the closest thing the field has to a commandment.
The point of the holdout is independence. Companies, however, run things the other way. Every number on the dashboard is announced in advance, to the people being graded on it. Everyone gets to study the test before taking it, and the same measure that directs the work then certifies the work is good.
That is not to say companies don’t produce independent views. They do, though usually by accident. A leader listens to a customer call that was never selected for review. A manager walks the shop floor without an agenda. Someone orders a surprise FP&A audit. These moments matter because nobody prepared for them.
The tempting fix is a secret KPI. It doesn't work. For the support team, a hidden copy of resolution time is exactly as incomplete as the public one. You need a different kind of evidence: whether the customer calls back about the same problem, the rework that shows up downstream three weeks later. Slower, sometimes it's just a conversation, and that's fine. Its job is to tell you whether the dashboard still means anything.
With people you got a warning, and it was usually free. Somebody grumbles that the target is stupid. A manager notices the hard tickets all landing in one queue. An AI agent files no complaint. It gets better at the number, night after night, while the rework and the callbacks pile up somewhere nobody is looking. No threshold is crossed. No alarm goes off.
So build the independent view before you switch the agent on, not after the quarter looks good. Write down what the metric misses. Have the humans in the loop watch that instead of the number. Then expect it to wear out too. The moment your independent evidence becomes visible and consequential, it's one more number to manage.
Measures stay honest only while nobody is optimizing them.