AI vendors are shifting from flat-fee pricing to consumption-based billing, and most organizations have no idea what their token spend is actually producing.
Dave Sharrock and Peter Maddison break down what's driving the shift to AI token economics, why the old "$20 per user" budget model is breaking down, and why usage alone is the wrong thing to optimize for. They dig into the pattern showing up across organizations, where a small share of users account for half the token spend, and why chasing that number down misses the real question: what value did that spend create? The conversation covers KPI traps, model selection tradeoffs, and how to build the kind of honest, open culture that lets you actually govern AI spend without punishing your best people.
This week's takeaways:
- Token usage by itself is a bad KPI once your organization has moved past early AI adoption, because it stops measuring exploration and starts driving the wrong behavior.
- The 10% of users driving 50% of the token spend aren't automatically the problem. Some are generating outsized value, and the only way to know is to ask them directly.
- Managing AI cost well means pairing spend visibility and caps with an honest conversation about the value that spend is producing, not just sorting a table by usage.
Listen to the full episode at definitelymaybeagile.com
Subscribe so you never miss an episode.
Have a question or topic you'd like us to cover? Reach out at feedback@definitelymaybeagile.com
New episodes released every Thursday to challenge your thinking and inspire action.
Listen and subscribe:
Welcome And Token Economics Setup
Peter [0:04]: Welcome to Definitely Maybe Agile, the podcast where Peter Maddison and Dave Sharrock discuss the complexities of adopting new ways of working at scale.
Dave [0:13]: Hello, Peter. How are you today? Great to catch up.
Peter [0:17]: So how are things? What are we going to talk about today? Today we're going to talk about token economics, because well, you know, that sounds so fun.
Dave [0:25]: The phrase itself just brings that little bit of like, "really, we've got to talk about money and accounting and numbers and things like this." So what is it? And there are loads of conversations right now, discussions around token maximization and the rumors going around of various companies spending millions and millions of dollars unexpectedly as they're shifting, or being shifted, to a consumption-based model rather than... I'd frame it like this.
The Flat Fee Token Paradise
Peter [0:56]: I'd say that, and I'm sure we can find a lot of this online, but earlier in the year, or over the last few years, we've been living in this paradise where you could pay $20 per user and get, well, not unlimited, but enough tokens to enable you to do AI-driven work. A token, in simple terms, is basically the unit of currency for engaging with an AI. Tokens translate to the amount of AI you're using or consuming. And we had this wonderful idea that we could budget for our company as $20 times the number of engineers or people we wanted to give access to these tools, and that would basically be our budget. So the CFO goes, "That sounds great, this is how much money I need." Then, of course, the AI companies realized, "Well, no, we're not making any money anymore." In fact, they weren't really ever making much money in a lot of cases, and they needed to generate more revenue. They couldn't keep charging a fixed price when some users were consuming astronomically more than others. So they had to switch to what we'd call a consumption-based model. And during that $20 fixed-cost era, you had this concept of "token maxing" you brought up before, which was basically companies driving people to use as many tokens as they could, because that showed they were "doing things" with AI, and doing things with AI supposedly meant generating value for the company. It's a bit like the idea that if you put a hundred monkeys in a room, you'll eventually get Shakespeare. Same concept, just a lot of noise.
Dave [2:49]: The more people using more cycles of AI, the more likely you make that transition to some sort of AI-native organization. And then we hit a KPI.
Why Vendors Shift To Consumption
Peter [3:01]: Yes. And then we hit the first of June. Microsoft, well, they'd announced it before that, but as of June 1st, 2026, they announced they'd shift to this consumption-based model, which immediately threw all of this up in the air. And we've kind of gone from token maxing to token mining, if you like. Reduce the number of tokens, quick, quick, because all of a sudden we've got this hockey-stick curve as people start consuming more and more tokens. And through all of this, we're still not necessarily seeing the value at the other end. There you go.
The Water Meter Moment
Dave [3:37]: I'll jump in here. As you're describing that, I'm sitting in my apartment right now, and it reminds me of a conversation going on in our strata about the building I'm in having to put a water meter in. Until now, there's been no water meter, and everyone can use the water they want for a flat fee. Obviously, the building is fighting getting a water meter installed, because the moment you do, you're paying for usage, not just access. And what's interesting, and this is what I keep thinking about, is there are people who take long showers and use lots and lots of water, and then there are people who nip in, have a quick shower, and they're out again. Now an organization has to figure out, metaphorically, who's taking really long showers and using all those tokens excessively, versus who's using water because they actually need it.
More Tokens Do Not Mean Value
Peter [4:44]: I think there was always a fallacy built into that. The fallacy was this idea that if you use more AI, you produce more value. There was this equation that didn't really make sense. In fact, it's been in various articles, and I've seen the same thing across multiple organizations: 10% of users accounting for 50% of the cost. In some ways, you'd think it would be surprising to see this pattern repeat so consistently across different organizations and industries. Although there's a certain logic to it, since the early adopters are the ones who jump on this and use it most.
KPI Traps And The 10% Effect
Dave [5:29]: A couple of things come out of this. Number one is the poor use of KPIs. In an organization with no AI usage, measuring AI usage as some kind of directional signal kind of makes sense. It's a proxy for "we're beginning to explore and play around with this new technology." But the moment you move to an organization that's already using AI in a number of ways, that KPI becomes useless. In fact, it drives poor behavior right out of the gate. So that's one point, that the KPI needs to be relevant to the behavior you're trying to change. But the other side is that 10%. If we just look at averages and go after the 10% with 50% of the usage, that's not necessarily solving the problem, because some of that 10% might be having a tremendous impact. They might be leveraging far more tokens than average, but the return they're generating could be outsized. And then there'll be another part of that population that isn't making good use of the tokens they're burning through.
Finding Impact In Complex Systems
Peter [6:50]: Yes. And that's kind of the piece, right? You have to go ask. You have to identify who that 10% is and ask them. Go interview them, find out what they've actually been doing with all those tokens. What have they been building? What have they seen?
Dave [7:09]: I'm not even sure it's as simple as talking to the people using those tokens, because we've talked about value recognition many times before. Value and impact are very difficult things to measure in most cases. There are scenarios where it's relatively straightforward, but most organizations you and I work with struggle to find a simple metric that pins down where value is being created. In reality, they're operating in a complex, multi-dimensional system for creating value, and it's difficult to tell.
Peter [7:46]: I think the piece there is: talk to the people first, because you might find out what they've actually been working on. What have they been using the tokens for? Have they been building business features, or have they been grabbing the most expensive model available and applying it to every mundane task, then never bothering to turn it off?
Dave [8:03]: Writing emails using Mythos or Fable.
Peter [8:08]: In which case they're going to burn through a lot of tokens when they could have used a smaller, simpler model. Which gets you into this idea, and I agree with you, that the ultimate goal is being able to say: what value did this token produce? How do I connect the spend I'm generating to the value I'm creating for the organization? Is there actually a connection between the two? That's where the art comes in, because what the CFO is looking for is, "Hey, I went from a fixed cost to a variable cost and it's going through the roof. Now I really need to know what this spend is producing. What value am I creating?"
Controls, Caps And Spend Visibility
Dave [8:57]: I think there are a couple of things here. First, you need visibility into that spend. As a CFO, the first question is: who's monitoring this? How are we capping usage? How are we preventing the kind of stories we're hearing about, companies unexpectedly spending millions, or tens of millions, depending on which article you read, because the consumption isn't being tracked or capped?
Peter [9:25]: There's some basic administration, some basic controls that should be in place: how much can you spend before you need to go ask? And then there are cultural pieces around that too. What needs to be true for someone to get their limits raised? Who gets the extra budget, and what does that look like overall? Because I need to be able to see where that money is going.
Dave [9:51]: And that becomes the second piece: how is value being created, what's the impact of that spend, and where can you track it? We've seen this pattern before, with the shift to the cloud and various other transitions. It's a journey a lot of executives have already been through.
Peter [10:12]: Yeah, it comes down to: am I increasing my top line? Am I reducing my bottom line? Am I managing risk? That's the hundred-thousand-foot view, the simplest way to think about it. But those then break down into more specific questions. How are the tokens being consumed actually benefiting the organization?
Dave [10:31]: One of the things that jumps to mind with these problems is the tendency to flatten the view, to just look at averages, say tokens spent per individual, and make decisions based on usage per person, rather than understanding the other side of it: what's the impact of that work, and how do you tie the two together?
Build An Honest Culture Around Usage
Peter [10:57]: And that's not straightforward. It's not a simple, single metric to look at. Which brings us to the concept of the honest organization, because if you're going to go and ask people about their usage, it can't feel like you're shooting the messenger. People won't respond honestly if that's the case.
Dave [11:15]: The people who responded to the call to use more AI, to really dig into it and use it to maximize what they're doing, they're probably many of your highest token users. And yet they're doing exactly what the organization asked them to do. The risk is, if you're not having good conversations around this, they end up being penalized.
Peter [11:42]: Yes, which is not what you want. Now, as we were saying, there can be people who need some help understanding how to use the tools better, how to use the right models for the right purposes. But again, that has to be a conversation handled well and openly, because it helps. It's a learning exercise. These are new tools, new capabilities. They're very powerful, there's a lot they can do, but there's a lot to learn about how and when to use them, and in what different ways.
Model Choice Guidelines And Tradeoffs
Peter [12:14]: So I think that's a good thing.
Dave [12:14]: Guidelines need to be put in place, things like this. It's interesting, I don't see this very often. I see a lot of encouragement to use tools and models, but a lot less guidance on how to optimize usage across different models.
Peter [12:32]: I know of one organization whose policy, and the culture they've built around it, essentially says: use the smallest model first, and only go up from there if you really need to. Start small. Even that on its own can help, though it's a tricky balance.
Dave [12:58]: My mind immediately goes to: if the smallest model isn't solving the problem, you're more likely to give up trying. It almost feels like there's a benefit to starting with the strongest model and then paring it down once you've proven you can solve the problem, until you land on something you're happy with.
Peter [13:18]: The problem there is you get hooked on the top-tier reasoning model, and there are so many directions this can go. There's a lot to AI we haven't even touched on, caching, temperature, prompting. There are a lot of levers you can pull to change how these engines behave. But if you've encouraged people to use as many tokens as possible, the more expensive models obviously help them do exactly that.
Dave [13:52]: So that's another thing that can spiral quickly.
Peter [13:55]: Yeah. I think there are some genuinely interesting pieces to think through around how token economics get managed.
Budgeting, Forecasting And Maturity Steps
Peter [14:06]: How do you help an organization think about budgeting and forecasting? We've got tools and frameworks that help organizations think through this. There's a maturity curve to it, starting with visibility, then moving up into optimizing spend or anticipating what you're going to spend, building on each layer to get better at using these tools within the system. I think of it as another part of the delivery system, and how that delivery system is set up to succeed.
Dave [14:46]: It's interesting that a year ago, and in earlier conversations we had on the podcast, we focused on things like governance, on just the basic setup. And it's interesting that in the last three or four months we've seen a real shift toward operational guidelines and policies, how to use and optimize these different models, because the cost model changed.
Peter [15:13]: Yes. And what's interesting is companies now have a real need, which is also what's driving this. You need to get a handle on the spend. We were talking to people about this before, it just wasn't as urgent as it's suddenly become.
Dave [15:34]: How do we wrap it up?
Key Takeaways For CFOs And Teams
Peter [15:35]: I'd wrap it up with a couple of key points. One, get a handle on where the tokens are going, who's using them, and start to understand what they're using them for, so you can put the right guidelines in place, properly govern the system, give the CFO what they need to run the business, and make sure costs don't spiral out of control as you adopt these tools. That's one of the main pieces. The other key piece is understanding the value add. It's not just about counting tokens, it's about looking at what you're doing with them. What value is being created, and which part of the organization is it impacting? Without that, you probably won't get the most out of bringing in the technology.
Dave [16:39]: If I'd add anything, it's this: if you're working with delivery teams, development teams, or the operational side, it's important to think about how you communicate to the CFO what's being done across the organization by different teams and individuals. I'm conscious of how easy it is to sort a table by token usage and suddenly penalize the most interesting, innovative, valuable contributors for actually trying to make a difference. There's a nuance to making sure budgets are managed the right way.
Closing And Subscribe Reminder
Peter [17:30]: Yes. Awesome. Well, pleasure to have the conversation as always, Dave. Looking forward to next time. And if you're listening, don't forget to tell your friends and hit subscribe. Okay, until next time.
Dave [17:43]: Thanks again.
Peter [17:44]: You've been listening to Definitely Maybe Agile, the podcast where your hosts Peter Maddison and Dave Sharrock focus on the art and science of digital, agile, and DevOps at scale.



