Managing AI Token Costs Now That Your Bill Has a Water Meter

Managing AI token costs used to be almost embarrassingly simple. For the last couple of years, you multiplied $20 by your headcount. You handed the number to your CFO, and moved on. It was predictable. It was clean. We all knew it wasn’t going to last. It was a lovely fiction.

That fiction is unwinding fast. GitHub Copilot moved its plans to usage-based billing on June 1, 2026. That replaced flat monthly pricing with metering tied to actual token consumption, now other vendors are following the same path.

Dave and I spent a recent podcast episode picking this apart. What does the shift actually mean for how organizations think about AI? Honestly, it's a bigger conversation than "the price changed."

The Flat Fee Made Everyone a Little Lazy

This is really why managing AI token costs is suddenly urgent. Flat pricing was convenient. But it hid what was really going on. Nobody had to ask what all that AI usage was actually for. Picture two companies on the same flat-fee plan. In one, ten people use the tool constantly. In the other, two hundred people barely touch it. Same bill either way. So there's no pressure to look closely. Usage itself became the goal. Companies pushed adoption hard. "Our people are using AI" felt like progress. Nobody could say what that usage was actually producing.

Dave brought in an interesting metaphor. His apartment building is fighting over whether to install a water meter. Right now, water is included as a flat fee. Nobody thinks twice about a long shower. The moment a meter goes in, every shower becomes a line item. Long showers and quick showers used to cost the same. Now they don't. Suddenly the building has to figure out who's actually using more than they need.

That's exactly the position a lot of companies find themselves in right now. The meter just got installed on AI spend. The flat-fee assumptions of the last couple of years don't hold up anymore.

More Usage Was Never the Same Thing as More Value

Token consumption and value created were never the same measurement. They just happened to be invisible together under flat pricing. Nobody had to reckon with the gap. Xodiac's piece on AI spend governance puts it well. Most technology leaders can't say what their organization spent on AI last month. They can't break it down by team, tool, or use case. That's not a technology problem. It's a governance problem, and it's exactly why managing AI token costs takes more than a spreadsheet.

In my own work with clients, I keep running into a pattern. A small slice of users account for roughly half the total token spend. Something like a tenth of the people with access. It's tempting to look at that number and assume you've found your problem. Cut the outliers. Rein in the heavy users. Done.

But that instinct skips an important question. Where is the genuinely valuable work being done? Some of those heavy users might be your most effective people. Or they might not be. Nobody bothered to check. From a spend report alone, those two groups look identical. You can't tell the difference by staring at a chart. You have to go ask people what they're actually building.

Usage as a KPI Stops Working the Moment It Starts Working

Early in AI adoption, tracking usage as a KPI makes some sense. It's a rough signal. People are exploring a new capability. They're getting comfortable, finding where it fits into their work.

But that signal has a shelf life. At some point, your organization genuinely adopts AI. A meaningful chunk of your people start using it every day. Once that happens, usage volume stops telling you anything useful. Worse, it starts rewarding the wrong behavior. Pick the biggest model. Run it on everything. Keep the number climbing. Whether or not any of it moves the needle.

If your dashboard still treats "more tokens used" as a win, ask yourself something honest. Has that KPI quietly outlived its purpose? Dave has made a version of this same argument about delivery in general. In Most delivery problems aren't execution problems, he makes the case. Teams get into trouble measuring activity instead of real value. AI usage is just the newest place that pattern shows up.

The Real Work Is Connecting Spend to Value, Not Cutting It

None of this means you shouldn't have controls. Spending caps, approval thresholds, visibility into where the money is going. These are basic administration. Most organizations need more of them than they currently have.

But controls alone answer the wrong question. Managing AI token costs well means figuring out what that spend is actually buying you. Is it increasing revenue somewhere? Reducing cost somewhere else? Managing a risk you'd otherwise be carrying? Those are the questions a CFO is really asking when the bill is "going through the roof." Not "make it smaller." "Tell me what it's for."

Getting there means talking to people honestly. It shouldn't feel like an interrogation. The people who leaned hardest into AI adoption are very likely your highest token users right now. They took "go use these tools" seriously. That was the outcome the organization asked for. If your first response to a rising bill is to flag those people as a cost problem, think again. They'll either hide their usage, or quietly stop experimenting. Neither one helps you.

Where This Leaves You

If you're trying to manage AI token costs as pricing shifts, this is what matters. The flat-fee era gave organizations a false sense of control. It made AI spend predictable without ever making it meaningful. Now that the meter is on, there's an obvious temptation. Solve it the way most of us solve unexpected bills. Cut usage. Set the strictest cap you can get away with. Move on.

That's the safe move, but it's the wrong one. Managing AI token costs well doesn't mean chasing the lowest bill. The organizations that come out ahead will be the ones who figured out where their AI spend was creating value. And where it wasn't. Those are the harder conversations, but worth having.

A Few Questions We Keep Getting

Why did AI tools switch from flat-fee to usage-based pricing? A small share of users started consuming far more AI resources than a flat rate could cover. That's why flat-fee pricing didn't hold up. Vendors including GitHub Copilot moved to usage-based billing, tied to actual token consumption, starting June 1, 2026.

How do I start managing AI token costs across my team? Start with visibility, not cuts. A spend report alone can't tell you why usage is high. It could be your most effective people doing valuable work. Or it could be someone running an expensive model on tasks a cheaper one could handle. Ask people directly what they're using the tokens for before you set caps.

Should I cut off my highest AI token users? Not automatically. Your highest token users are often the people who took AI adoption seriously in the first place. Cutting them off risks punishing your most engaged employees. It also teaches everyone else to hide their usage. Managing AI token costs well means understanding that difference before you act on it.

This post draws from a recent Definitely Maybe Agile episode with Peter Maddison and Dave Sharrock. Listen to the full episode at definitelymaybeagile.com.