
Uber didn’t plan to burn through its entire 2026 AI budget in four months. But with engineering adoption of agentic coding tools jumping from a third of its workforce to over 80% in a matter of weeks, that’s exactly what happened – some engineers were spending $500 to $2,000 a month each, and leaderboards that rewarded heavy usage only accelerated the burn!
In another incident, Microsoft pulled Claude Code licenses from its Experiences & Devices division after costs hit roughly $2,000 per engineer per month. At a sports technology company, one engineer was driving $600,000 a year in spend across 40 different AI models, nobody in finance or engineering knew until a third-party audit surfaced it.
The pattern in every one of these cases is similar: a genuinely valuable tool rolled out without the financial guardrails that traditional per-seat software licensing used to provide automatically. This now shows up in nearly every customer conversation: the first question isn’t whether to use AI, it’s what the tokens will cost!
So how is industry solving this? Google kept a human reviewer in the loop for invoice reconciliation process rather than chasing 100% automation on day one, and generated $30 million in savings.
Here are the five strategies that we are advising our customers to implement immediately:
1. Govern before you scale, not after the invoice: The core governance model is simple: set a spend cap per employee, per month. When someone crosses it, flag it, and have a real conversation. Pair that with cost attribution so finance can see where every dollar is going, not just a lump sum from “OpenAI” or “Anthropic” on the P&L.
2. Standardize on a small number of approved models: The $600,000-a-year spend was spread across 40 different models, a spread no one could see. Pick two or three approved models spanning a cheap tier for routine tasks and a premium tier for genuinely hard problems, and route work to the cheapest model that can do the job. This alone closes off most of the “shadow AI” risk.
3. Centralize how prompts and tools are managed to benefit from prompt caching: Instead of every team writing and maintaining its own system prompts, route requests through a central gateway. This is what makes prompt caching actually work. Every major AI provider now discounts repeated context by up to 90% on a cache hit, but only if that context is genuinely identical across calls. Decentralized prompts break the cache; a centralized setup captures the discount by design.
4. Make usage visible to the people spending it. Awareness is a strategy in its own right. Put cost-per-task or cost-per-team dashboards in front of the teams generating the spend, not just finance. Retire leaderboards that reward raw token volume, and replace them with ones that reward efficiency — output per dollar, not tokens burned.
5. Review value, not just spend. Use batch processing for anything that isn’t real-time (routine discounts of 25–50%), and hold a regular review of what the spend is actually producing.
Finally, none of this requires slowing down adoption. It requires treating AI spend as a metered, governed line item from day one, the same discipline that turned cloud computing from a budget wildcard into a predictable cost center. The companies doing this well right now aren’t spending less on AI. They’re the ones who are extracting the most value from AI.
