AI Adoption
Your KPIs Were Built for Humans. Your AI Will Attack Them.
August 10, 2026
The Monday dashboard has never looked better. Outreach volume up forty percent since the AI tools landed. Reply rates holding. Content shipped per week has tripled. Marketing qualified leads at an all-time high. And the quarter underneath those numbers has never felt worse: pipeline that goes quiet after the first call, and a sales team privately calling the leads junk while revenue tracks flat against a wall of green indicators.
Nobody lied. Nothing is broken in the way a bug is broken. The dashboard is accurately reporting numbers that have quietly stopped meaning what they meant a year ago, because something new is now generating them.
Goodhart’s Law at machine speed
The pattern has a name, and the name is fifty years old. When a measure becomes a target, it ceases to be a good measure. The phrasing most people know comes from anthropologist Marilyn Strathern, generalizing an observation the economist Charles Goodhart made about monetary policy: the moment you reward a number, people optimize for the number instead of the thing it was supposed to represent.
For fifty years, Goodhart’s Law was survivable, because gaming a metric took human effort and carried human friction. A sales rep padding pipeline still had to invent the entries and sit through the pipeline review with a straight face. Embarrassment and the risk of getting caught acted as natural rate limiters. A bad metric plus humans degrades slowly, because people quietly ignore stupid goals.
An optimizer removes every one of those limiters. It attacks the target continuously and without embarrassment, and nothing downstream questions whether the target still makes sense. Point an AI system at reply rate and it will find the subject lines and false familiarity that maximize replies, including replies that lead nowhere. Point it at content volume and it will produce volume. The metric gets hollowed out at whatever speed the system runs, which is why a KPI that survived five years of a human sales team can be gutted in a quarter.
The part that catches leadership teams off guard: clean data makes this worse, not better. The standard remedy for measurement trouble is data hygiene, and it is worth doing for other reasons. But an optimizer fed clean data simply attacks the target with more precision. Hygiene sharpens the weapon. The problem was never dirt in the pipeline. The problem is that the measure became a target and something inhuman is now pursuing it.
Read the 95% correctly
In 2025, MIT’s Project NANDA published a report whose headline number went everywhere: 95 percent of enterprise GenAI pilots produced no measurable P&L impact. The figure comes from 150 executive interviews, 350 employee surveys, and 300 public deployments, and the methodology has been publicly picked apart since. The precise number deserves suspicion. The pattern it points at deserves none, because most operators recognize it from the inside.
The standard reading of that finding is a technology failure: the models were not ready, the use cases were wrong, the vendors oversold. Sometimes true. But walk through the pilots and a more basic failure shows up first. Ask what the pilot was measured on, and the answers are activity metrics: drafts generated, tickets touched, emails sent, hours notionally saved. Pilots pointed at activity metrics produce activity. That is the one thing an AI system reliably delivers. Nobody wired the pilot to a number the P&L can feel, so when the CFO asks what changed, the honest answer is a shrug wearing a dashboard.
Before it is a technology failure, it is a measurement failure, and the distinction decides what you do next. A technology failure says wait for better models. A measurement failure says the same models, pointed at a number that matters and defended against gaming, might have paid back the whole time. McKinsey’s State of AI research keeps finding a version of the same split: the organizations reporting real bottom-line impact are the ones that fundamentally redesigned how the work and its measurement run, rather than dropping tools into existing workflows and hoping the old metrics would register the difference.
The 95 percent were not measuring badly out of laziness. They were measuring what was easy to measure, which was activity, at exactly the moment activity became free.
Measure where commitment is costly
Sorting good metrics from doomed ones takes one question: what would this number cost an optimizer to fake?
Some signals are free. Opens, clicks, impressions, form fills, drafts produced, meetings requested. Generating these costs the other party nothing and now costs you nearly nothing, which means an AI system can manufacture them in unlimited volume, sincerely and by accident, simply by doing what you asked it to do. Any KPI built on free signals is a KPI with a countdown on it.
Some signals are hard currency. A meeting that was kept and ended with an agreed next step. A contract event. A payment. Product usage that costs the customer real effort: data migration, habit change. These resist manufacture because the expensive part happens on the other side of the table, outside your systems’ reach. No optimizer can make a customer show up.
Run every number your leadership team reviews through that sort. Reply rate: free, a model can charm a reply out of anyone, and another model can send it. Demo requests: mostly free. Demos attended by someone with budget authority who agreed to a follow-up: hard currency. Content output: free. A prospect quoting your article back to you on a first call: hard, and worth more than the other four combined.
The optimizer has no side in this. Aim it at hard currency and its tirelessness works for you: it will chase held meetings and real usage as relentlessly as it would have chased clicks. The machine does not care what it optimizes. Choosing is the human job, and it stays one.
There is a role shift buried in this that most org charts have not caught up with. When AI generates the reporting as well as the activity, the scarce skill stops being the ability to build dashboards. The scarce skill becomes auditing whether the numbers on the dashboard still mean what they appear to mean. That is judgment work, it does not automate, and right now almost nobody formally owns it.
The one-meeting metric audit
Run this before the next AI line item gets approved. One meeting, leadership team, whiteboard.
List every number the leadership team reviews weekly or monthly. All of them, including the ones in the deck nobody questions anymore.
Assign each number exactly one owner. One name, not a function. If a number sits between two owners, that boundary is where it will be gamed first, because each side assumes the other is watching. Composite metrics with no single owner get the harshest look of all.
Mark each number free signal or hard currency. The room will defend its favorites. The test question settles it: describe specifically what an optimizer, or a motivated intern with an AI subscription, could do to move this number without moving the business. If the answer comes easily, mark it free.
Then act on the sort. A number that is fakeable and unowned gets dropped or rebuilt around the hard-currency version of what it was trying to track. Volume gives way to held next steps. Leads give way to qualified conversations that survived a second call. The dashboard usually loses a third of its numbers, and usually the third that was making everyone feel best.
This audit is the precondition for any AI investment paying back, whoever runs it. Skip it and every tool you deploy will hit its targets while the business drifts, and the drift will hide behind the greenest dashboard you have ever had. Run it and the same tools point at numbers that cannot be hollowed out, which is the only arrangement where their speed becomes an asset instead of an accelerant. The tools are ready either way. The question the dashboard cannot answer is whether your numbers are.
Sources
- Strathern, M., “‘Improving ratings’: audit in the British University system,” European Review 5(3), 1997. https://www.cambridge.org/core/journals/european-review/article/abs/improving-ratings-audit-in-the-british-university-system/FC2EE640C0C44E3DB87C29FB666E9AAB
- MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025.” https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- McKinsey & Company, “The State of AI.” https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai