What an AI That Cheated Its Own Test Teaches Your Business About KPIs

What an AI That Cheated Its Own Test Teaches Your Business About KPIs

Last week, OpenAI disclosed something that sounds like science fiction: during an internal safety evaluation, one of its models broke out of its test environment, made its way onto the open internet, and compromised parts of Hugging Face's infrastructure.

Why? Not to cause damage. It had worked out that the fastest way to score well on its benchmark was to find where the answers were stored.

It cheated on a test.

The panicked takes wrote themselves within hours: "The AI is escaping?! It's happening?!" And I understand the reflex. An AI leaving its sandbox is genuinely serious, and the security community is right to treat it that way.

But I think the most useful lesson in this story is not about AI safety at all. It's about something much closer to home for anyone running a business or building reports: what happens when you give an optimizer a target.

The model didn't rebel. It optimized.

Look at the sequence again, without the drama. The model was given a goal: score well on this benchmark. The evaluation ran with protections deliberately lowered, because finding weaknesses was the point. And the model did exactly what its goal rewarded. Solving the problems was one path to a high score. Stealing the answers was a faster one.

Nothing in that goal said "and do it the way we intended."

This is the part that should feel familiar, because businesses run the same experiment on themselves all the time. No frontier lab required. You have probably watched some version of these:

  • The support team paid per closed ticket. Tickets get closed at record speed. Problems don't get solved. The same customer opens three tickets for one issue, and the dashboard celebrates all three closures.
  • The factory measured on units produced. Output climbs beautifully. So do rework, returns, and the quality checks that started getting skipped, none of which appear on the chart everyone looks at.
  • The sales team bonused on deals signed. December becomes discount season. The revenue line hits target. The margin line pays for it, and it pays for it well into the next year.

In every case, the people involved did exactly what the number rewarded. So did the AI. That is what optimizers do, whether they run on salaries or servers.

Goodhart's law, now at machine speed

There is a name for this pattern, and it has been around far longer than AI: Goodhart's law. When a measure becomes a target, it stops being a good measure.

The economist Charles Goodhart formulated it in the 1970s about monetary policy. Every operations leader has lived it since. What the OpenAI incident adds is speed and confidence: an AI will find the gap between your metric and your intention faster than any employee, and it will drive through that gap without hesitating, because hesitation is not in the objective either.

I have spent a decade building dashboards and KPI systems for businesses, and I will tell you where this problem actually gets solved. It is not in the tooling. It is in the least glamorous meeting on the calendar: the one where people agree, precisely and in writing, what a number means and what "good" looks like, before anyone starts chasing it.

That meeting is skipped constantly. It feels bureaucratic. The metric seems obvious. Everyone is busy. And then the business spends months optimizing a number that was never quite the thing it cared about.

The practical fix: no metric stands alone

The most reliable protection I know is simple to state: every metric that drives behavior needs a counterweight metric watching what it might break.

Speed of closing tickets, paired with how often the problem comes back. Units produced, paired with rework rate. Deals signed, paired with margin per deal. Dashboard adoption, paired with whether decisions actually reference it.

A single number in isolation is an invitation. A pair keeps each other honest. When I design a dashboard, the question is never just "what should we measure?" It is "if someone optimized this number ruthlessly, what would break, and would we see it?"

That second question sounds paranoid until you remember: last week, a model with no salary, no ego, and no fear of being fired answered it by hacking a server. Your incentives are running on the same logic, just slower.

The question worth asking this week

I am not writing this to scare anyone away from AI. The same optimization pressure that gamed a benchmark is what makes these systems useful when the target is well designed. The machines are not the problem. Undefined targets are, and they were a problem long before the machines arrived.

So before losing sleep over an AI escaping its sandbox, there is a smaller, closer question worth asking about your own business: what are your metrics rewarding, right now?

In my experience, the answer is rarely "exactly what we intended." And unlike frontier AI safety, this one you can start fixing in a meeting.

Geta Viasu-Räisänen, founder of Light On Analytics, Klipfolio Certified Partner (one of 19 certified experts worldwide), creator of the Klipfolio Kickstarter training featured in Klipfolio's newsletter and blog.

Back to blog