Why AI Agents Need Room to Fail Before They Learn
Definitely, Maybe AgileJuly 30, 2026x
228
00:17:1611.89 MB

Why AI Agents Need Room to Fail Before They Learn

Giving an AI agent real autonomy means accepting it will fail early and often before it gets good, the same curve organizations hit during any real change. Peter Maddison brings a stuck OAuth problem to the table: an AI agent that kept going in circles and couldn't find its way through. That leads into a conversation from Dave Sharrock's local AI meetup about an AlphaGo-style approach to AI agent autonomy: instead of specifying every step, you define hard constraints and let the model work o...

Giving an AI agent real autonomy means accepting it will fail early and often before it gets good, the same curve organizations hit during any real change.

Peter Maddison brings a stuck OAuth problem to the table: an AI agent that kept going in circles and couldn't find its way through. That leads into a conversation from Dave Sharrock's local AI meetup about an AlphaGo-style approach to AI agent autonomy: instead of specifying every step, you define hard constraints and let the model work out its own strategy inside them. Peter and Dave connect this to the Virginia Satir change curve, the same dip in performance that shows up when an organization tries a new way of working, and to the difference between using AI to optimize what you already do versus using it to rethink the business itself. They also get into how experiments like Andon Labs' AI-run cafes and vending machines use small dollar constraints to let a model learn from failure without real financial risk.

This week's takeaways:
- A well-articulated objective with clear guardrails lets an AI agent find its own path to a solution, even one you didn't expect or fully understand.
- Real learning, whether it's an AI agent or an organization adopting a new way of working, comes with an unavoidable dip in performance that can't be planned away.
- The bigger opportunity with AI isn't squeezing more efficiency out of an existing process, it's using AI to test entirely different ways a business could operate.

Listen to the full episode at definitelymaybeagile.com
Subscribe so you never miss an episode.
Have a question or topic you'd like us to cover? Reach out at feedback@definitelymaybeagile.com

New episodes released every Thursday to challenge your thinking and inspire action.

Listen and subscribe:

Welcome To Definitely Maybe Agile 

Peter 0:04 Welcome to Definitely Maybe Agile, the podcast where Peter Maddison and Dave Sharrock discuss the complexities of adopting new ways of working at scale.

Dave 0:13 Hello, Dave. How are you today? Excellent, Peter. Good to see you again. So what's new in your world?

AI Experiments And OAuth Frustration 

Peter 0:20 Oh, a ton of things. I've been playing around with AI as a game in all sorts of different ways, and having a bunch of fun, and some not so fun, since it keeps running into problems that are fairly well known, problems it should have been able to solve and then isn't able to. But there are things in there that are interesting. And I know your AI meetup, where you live, had a very interesting presentation last week about what I'd call AI autonomy. I'm curious, because I've described to you the problem I was having with OAuth, with the agents just going in circles and not being able to figure it out. Do you think that if I'd given the AI more autonomy, it would have been able to solve my

What Autonomy Really Changes 

Peter 1:06 problem?

Dave 1:06 It would have solved it in some way or other, yes. Wouldn't it? It depends how you articulate the problem, I think. The AI discussion you're talking about: Callum, my son, came in and spoke a bit about the work he's doing, and on that note, we've had Callum on as a guest before, and he's doing work around evals and safety. We were learning a bit about how they structure problems. We've talked a lot in these conversations about context and agents, and about how to go about implementing AI in a business process or environment. I think one of the key takeaways a lot of the audience walked away with from last week's conversation was that, let's say in Silicon Valley, or at least in certain places, they're not trying to constrain the problem. They're taking this sort of AlphaGo approach: go learn how to do the problem as best you can, and let the LLMs steer. So to your point about the OAuth problem, if you hand that problem over, articulated well, and then get out of the way, the LLM will come back with a solution. Not necessarily the one you're looking for, or one you can explain, but a solution.

Articulate The Objective Clearly 

Peter 2:37 So in this case, I think the key part is the "articulated well." And the piece that Andon Labs is playing with is an idea we've talked about for years: if you've got a good understanding of what the objective is, and you can make it measurable, you can say, "go after that." You're setting a direction, giving it a place to go, and saying, okay, make your decisions. Now you can start making decisions toward that target. And as long as your decisions are moving you toward that target, and it's the right measure and the right direction, you should be successful at whatever that objective is.

Dave 3:19 Well, if it's

Measuring Success In Messy Work 

Dave 3:20 relative, like say a complicated problem. If I'm taking a set of accounts and trying to reconcile what's happening there, the rules are relatively straightforward, and can be defined pretty cleanly to figure out mathematically what's going on. It's much harder, I think, in a real-world context. Take customer support, for example. What's the definition of a satisfied support call? There's the obvious stuff, the call was this long, the customer left with their problem resolved, and so on. But there are a lot of nuances around exactly how satisfied someone was with how the problem was solved, and whether it satisfied the problem long term or was just a short-term fix.

Peter 4:08 Well, that brings to mind Big Hero 6, you know: "Are you satisfied with your care?" The conceptual idea there, well, you can ask afterward, sort of like, did you get what you were looking for from this conversation? With AI, we can of course start looking for things like sentiment. What was the tone in the conversation? Did the tone change during the conversation? There are all sorts of things you can look for in those conversations. But to your point, it's not cut and dried. You can't just have a number and say, "go solve for this number."

Dave 4:53 Well, and I think that's where, again, I'm not speaking for Andon Labs here, but from what I understood, if you're taking, say, a cafe, and the key metric is you should make money over the long run, or whatever a cafe experiment might be, they're controlling and simplifying the environment so they can get data on how the system performs and changes. That completely makes sense. Once you're dealing with as many variables as you get in a real business context, those variables grow and grow, and you end up with these really complex, nuanced relationships with customers, partners, whatever it might be. Those problems become a lot harder to articulate in a way that has a simple reward function.

Peter 5:37 Right, exactly. And one of the pieces you mentioned that people found surprising was the idea that they'd essentially characterize it as giving the agent free rein.

Dave 5:51 Yeah,

AlphaGo Style Learning With Guardrails 

Dave 5:52 with guardrails and governance constraints and things like that. A lot of the conversation we've had is around creating that context, locking it down, and spending time clearly defining what's important in the context of the problem they're trying to solve. Now, the AlphaGo model, the foundational idea behind a lot of the work being discussed, is about not sitting down and explaining through context what the rules are and what the best plays are. It's more about defining clear boundaries: you've got to stay on the board, there's only so many moves you can make. Those are the hard constraints. But then you let the LLM actually figure out what strategy looks like within that. By doing that, they lose a lot of games in the early stages, because they're trying things and failing, trying things and failing. But through that, they start learning what works and doesn't, and develop their own interpretation of how to approach it. That's exactly what Callum was talking about with the cafe and vending machine examples: initially there's a massive drop, they lose a lot of money at the beginning because those early tries fail more often than they succeed. But in the long run, they start turning the corner and getting better and better at whatever it is they're doing.

Peter 7:25 So there's a learning loop from the prior runs, essentially. They feed in what happened before: here's what you tried last time, this succeeded, this failed, these are the things that looked promising. Then they feed that in as guidance for the next iteration: here's another thousand bucks, see if you do any better this time.

Dave 7:46 Well,

The Satir Curve Of Change Costs 

Dave 7:47 what's interesting, as you describe that, is I'm thinking of the Virginia Satir curve of organizational change. We've bumped into this a hundred times in a different context: we want change in the organization, but we don't want it to cost anything. We don't want that dip in productivity while people experiment and learn a new way of working. That's basically what was being described in this conversation. And yet what we tend to do, and we've seen this in organizational change, is go in and say we're going to make a process change, and the desire is to make it cost nothing at all. So you end up defining every part of the process, because somehow we're going to make this lift and shift and, ta-da, it's going to be better. But we're not going to incur the learning cost of having the opportunity to use this platform, this technology, whatever it is that's new, in a way that's even better than how we used to do it.

Peter 8:48 Which is far more disruptive to the organization, and far more likely to fail. That's why the last thing you want to do is build out the entire thing and then move everybody there overnight. "Our entire process changes on Monday." It'll be chaos. It always is.

Dave 9:04 And people get caught up in it, because there are individuals somewhere making decisions that affect lots of other people, and they don't have the knowledge, the experience, or the specific context that each of those other individuals has. So how do you make these changes without a learning curve? Well, you don't. You need a learning curve.

Peter 9:24 Yeah, and you do it through small, incremental changes. You don't make massive changes across the whole organization at once, because that's incredibly disruptive. You need to introduce new things in small increments, learn, experiment, and see where the problems are going to show up. That's the right way to go about these kinds of changes in organizations. It's interesting, what you're describing with the agent learning from its prior iterations, it's basically becoming an improving loop, learning and improving through that reinforcement-learning process. That's not new. We've used that in machine learning for a very long time. I think some of the newer pieces, allowing agents to have those self-reinforcing loops, is good, and seeing where they end up with that is interesting in itself. I do think there's something to this, especially in software development, where there are so many variables and so many ways you can go. A lot of the spec-driven world says, define everything up front. That's not always appropriate. And as the models get, I'll put "smarter" in quotes, and make fewer mistakes, they're starting to respond well to a less fully specified description and being given more freedom to decide how something gets implemented. There are still guardrails you need around that. You don't want your agents running off and doing things in unintended, or potentially damaging, ways. That's why even in the Andon Labs example, their $1,000 constraint is actually a pretty large one. They're not saying, "here's my credit card and bank account, spend whatever you want, just make me money."

Dave 11:34 Yeah,

Beyond Optimization Toward Business Rethink 

Dave 11:35 that's for sure. As you're describing that, I keep coming back to something from a lot of our conversations: we keep talking about this optimization mindset. We look at it through an optimization lens, where organizations trying to improve their AI ROI are thinking, "I have a process, it could be more efficient, how can I use AI to make it more efficient?" And that becomes a conversation about constraints, about defining context so you're effectively ring-fencing this AI technology to make things quicker, faster, better in some way. What we keep coming back to is that, instead of optimization, there's a lot more to look at, like how you can really rethink how your business operates. That feels like, okay, let's find some system or ecosystem, some simulation world, and let the LLM run it and see what it comes up with strategically, something different from what our executive team, with all their experience and knowledge, would come up with as a potential solution. And start exploring that simulation world. I don't think you want to put an LLM at the top of Ford and say, go figure out how to design, sell, and support cars going forward. But on the simulation side, you might well want to explore that, because there are a lot of industries ripe for a rethink.

Peter 13:13 So, what if? Here's a concept for you. Take something like the quarterly prospectus of an organization, which describes its financial operations, and use that to build a digital twin of the structures, decision-making, and ways that organization makes money. Then deploy an agent into it and say, okay, if you were in charge, what decisions would you make?

Dave 13:43 Yeah, I don't know that you'd start just with that, but yes, there's something to explore there, say with a digital marketing agency or some other self-contained business model, and see how that would be done differently. I think it's a pretty interesting one.

Peter 13:58 I think that'd be a fun little experiment to run. Maybe I'll try it if I ever have a free weekend. But yes.

Dave 14:04 Yeah, for sure. Now, what do we take away from this conversation?

Context Windows Takeaways And Wrap 

Peter 14:10 For me, there are two big pieces. One is making sure you're taking advantage of the capabilities an LLM has for thinking outside the box, understanding those capabilities more, and using them in the right places. I'm not sure everybody does. And the other, one we've talked about many times before, is that while you can take a non-deterministic engine like an LLM and put it into an existing process to improve it, if used the right way, some of the real value comes from stepping back and asking: what if we did this in an entirely different way? What if we used something new to disrupt how we do business today, and potentially even disrupt our industry?

Dave 15:07 I'll lean on the Satir curve idea I introduced earlier. One of the things striking me about this conversation is the graphs that were being shared last week: the performance of these LLMs is laughably poor at the beginning. There are lots of stories about LLMs making a whole bunch of daft decisions when you tell them to go run something. But the interesting thing is there's a bottom to that curve, and then they start learning what their context is and how to make decisions. That, to me, mirrors the Satir curve of organizational change. Organizations don't like incurring cost, they don't want to see a dip in performance while new skills are being developed, used, and learned, until those new skills lead to something effective.

Peter 16:13 Yeah, I think one of the interesting things they're probably managing is how they handle the context windows of what gets fed back in. Because you don't want to feed in the entire transcript, you need to summarize it. So how do you pull out the parts that are valuable to feed into the next iteration, and stop it from becoming unbearable because there's too much information and only so much context window to work with? But that's a whole other set of conversations. So, with that, let's wrap up for this week. As always, tell your friends, hit subscribe. And if you'd like to reach out to us, you can do that at feedback@definitelymaybeagile.com.

Dave 17:00 Excellent. Until next time, Peter. Thanks again.

Peter 17:02 Thanks.

Peter 17:03 You've been listening to Definitely Maybe Agile, the podcast where your hosts, Peter Maddison and Dave Sharrock, focus on the art and science of digital, agile, and DevOps at scale.



Definitely Maybe Agile,agile podcast,devops podcast,peter maddison,David Sharrock,Agile Transformation,Business Agility, AI agent autonomy, AlphaGo strategy, Satir change curve, Reinforcement Learning,organizational change, AI in business, digital,