← In the News
The Decoder / FT

Amazon kills its internal AI leaderboard after staff gamed it

Uprovd Take Amazon measured activity, not value - and got more activity

  • Outcome Metrics
  • Measurement
  • Renew Watch Kill
Read the original at The Decoder / FT

Reporting describes how Amazon retired an internal leaderboard that ranked teams by AI usage after employees learned to run up activity with low-value tasks - behavior nicknamed “tokenmaxxing.” The metric rewarded volume, so volume is what people optimized for.

It’s a clean illustration of a measurement trap: when you score AI by how much it is used, you get more usage, not more value. The number goes up while the business case stays exactly where it was. A leaderboard cannot tell you which tool to keep, because “most used” and “worth paying for” are different questions.

The question a leaderboard can’t answer is the one a CFO actually asks: of everything we’re paying for, what do we renew, what do we watch, and what do we kill? That needs value attributed to spend per tool, ordered by the money at stake - and it needs to be auditable, because a number nobody can trace is just a leaderboard with better branding.

Uprovd’s position is that the only durable metric is outcome, traced to its source record and carried with an honest confidence score. Activity can be gamed; a measured change in cost, time, or quality, traceable down to the commit or ticket that produced it, cannot. That’s the AI portfolio verdict - and you can get a provisional read on your own tools in two minutes with the free Keep-or-Kill Assessment.

This is Uprovd's analysis of third-party reporting. Original article linked above.

5 SPOTS REMAINING · FOUNDING CUSTOMER PROGRAM

The headlines are arriving at our thesis.
Prove your AI value before your board asks.

Apply to be one of our 5 Founding Customers - or try the demo to see the platform first.