How to Validate AI Output Before You Trust It
Definitely, Maybe AgileOctober 01, 2026x
231
00:14:4610.18 MB

How to Validate AI Output Before You Trust It

AI agents will tell you everything works. Here is how to validate AI output when the mistakes are ones no human would ever make. Peter and Dave compare notes from building software, training material, and presentations with AI. The problems they keep running into are strange ones: the wrong database ID dropped into a system, 236 passing tests that test nothing, and five layers of abstraction added to a solution that needed none of them. Their answer borrows from good agile practice. Define w...

AI agents will tell you everything works. Here is how to validate AI output when the mistakes are ones no human would ever make.

Peter and Dave compare notes from building software, training material, and presentations with AI. The problems they keep running into are strange ones: the wrong database ID dropped into a system, 236 passing tests that test nothing, and five layers of abstraction added to a solution that needed none of them. Their answer borrows from good agile practice. Define what validation means before the agent starts, the same way a team agrees on its definition of done. Keep checking that the context you gave the AI is still relevant as the work moves. And keep someone in the loop who knows the problem well enough to look at the result and say, "Actually, you got this wrong."

This week's takeaways:
- Decide what "done" and "valid" mean before you hand work to an AI agent, and use a separate agent to grade what comes back.
- A test suite that pings components or mocks everything out can report hundreds of passes and prove nothing, so ask for synthetic transactions that run through the whole system.
- AI will happily over-engineer a small problem, so someone needs enough knowledge to ask whether the solution fits the size and scale of what you are actually trying to solve.

Listen to the full episode at definitelymaybeagile.com
Subscribe so you never miss an episode.
Have a question or topic you'd like us to cover? Reach out at feedback@definitelymaybeagile.com

New episodes released every Thursday to challenge your thinking and inspire action.

Listen and subscribe:

Why AI Output Needs Proof

Introduction [0:04]

Peter [0:04]: Welcome to Definitely Maybe Agile, the podcast where Peter Maddison and Dave Sharrock discuss the complexities of adopting new ways of working at scale. Hello, Dave. How are you?

Dave [0:13]: Excellent. Good to see you again. So we've got an interesting topic for today.

Peter [0:19]: Yes. We're going to have a chat about this: AI can go and build you a piece of software. Actually, it doesn't have to be software. AI can go and create something for you. But how do you know that it's right, that it works, that it's the right thing? Especially if you're going to do this autonomously and have one agent pass it to another, how do you go about verifying it?

Strange Mistakes Humans Rarely Make [0:43]

Dave [0:43]: I think this is coming up because you and I have both had experiences where we're building something, whether it's software products, presentations, training, whatever it might be. And we're reassured that everything has been taken care of. Being somewhat experienced at this point, we don't just nod our heads and say, "Great, everything's working." We start prodding, and we find out some pretty simple things have been missed.

Peter [1:11]: Yeah, and you're going, "A human could never make this mistake."

Dave [1:16]: Yes, let's say that.

Peter [1:17]: Honestly, some of the things I've seen it do, I would say a human would not be able to make some of these mistakes. They're not accidental things. They're just not the kind of mistakes a human would make. No human would go and use completely the wrong database ID in the system. If they were the one building it, they'd know exactly which one to use.

Dave [1:46]: I think that's the interesting thing. These are often situations where, at least from my side, it feels like they have all the information needed to make an informed, good-quality decision. And then you find that the prompt, or the way we're building this, has blanked out on a particular part of the system or missed something completely. That's the bit we want to check on. How do we know what we're building is a coherent whole, is a step forward, and is doing what we think it should be doing?

Context Windows and Garbage Out [2:23]

Peter [2:23]: There are a lot of pieces to this, but we can talk about some of the parts we've seen. One part is obviously about context. Ensuring context is well defined and accessible, and that what the agent pulls into its context window is actually valuable and relevant to the task at hand, rather than pulling in the entire universe every single time. One of the other things we know is that as context windows start to fill up, you get more and more garbage out the other end. So it's still garbage in, garbage out, just like with anything, really. But there are other pieces we've seen over and over again. Even with the right context and the right information on a small task, it decides, "I'm just going to do something else." Or, "I'm going to totally over-engineer this." Even when I've given it the general context, so it knows the size of the organization and the solution I'm building for, it goes and builds something with 70 different layers and far, far too much engineering for the size of the problem you're trying to tackle.

Dave [3:45]: I think you've introduced two distinct pieces there. One is understanding context, and before we move off that, I wanted to add two things. First, I'm finding you have to continually go back, reread it, and validate that it's still relevant, because you add more and more context as you work through the problem. It isn't one and done. It's something you're continuously monitoring. Second, and this goes slightly against the first point, there are studies now showing you can get rid of a lot of context because much of it is, let's call it common sense for the moment. You can reduce the amount of context and still get strong, coherent results. I don't think there's a clear understanding yet of where that comes from or how it's happening, but I think there's going to be a shift in how context is used. The study I'm referring to is from Claude, where they reduced context by up to 80% without impacting the quality of the outcome.

Peter [4:46]: Some of that is the tailoring of the model, its ability to respond to the right type of prompt, especially as it learns about you and the environment it's operating in. A lot of that behavior is already built in, so it can create a much more relevant outcome from much less context. We definitely see that. But even with the latest, greatest frontier models, I've still had it create absolute craziness for the intended purpose. So that brings us back to where we started. What else can we do to validate, as we're building this, that it's actually doing what we want or going in the right direction?

Defining Validation With a Grader [5:40]

Dave [5:40]: What I've noticed, and I want to be careful here, is that it's almost an assumption. We think it's obvious there will be some sort of validation piece in there, because we talk about it a lot and we capture it. But it's not always validated in the way we would expect. So there's a real benefit in clearly articulating, as part of the planning, exactly what we mean by validation. In the agent loops that are being talked about a lot more now, this is the grader function: how we actually grade what comes out as high-quality output, or as the outcome we're looking for. We've seen this in a number of the things you and I have worked on together. There's a need for some sort of assessment or verification prompt or agent that comes back in, looks at what's being created, and says, "Yes, we like where this is going," or, "No, we don't."

Definition of Done Meets Test Automation [6:40]

Peter [6:40]: That comes down to the skills, knowledge, and tools we've always used to identify the metrics and targets. What are we trying to achieve here? How can I clearly define what success looks like? What's my definition of done? If I can actually define my definition of done, then I can set the agents off to go achieve it.

Dave [7:06]: I love that, and I'm smiling because it comes straight back to agile and the things we were putting in place with agile teams. Do you have a definition of done? Are you validating that definition of done? I'd bring in test automation as a significant part of this discussion, because there are lots of layers in test automation. What I find is that I can't just assume that's happening. I have to direct it a little to make sure it's testing in the right way, so I can validate that functionality persists from one version to the next across the whole system we're working on.

Peter [7:43]: And then you have a separate agent looking at the tests, making sure your first agent didn't just say, "Hey, I passed the tests," by mocking them all out. Or it tells you, "I've got 236 passing tests." Well, sure you do. Except most of those tests don't actually do anything, and none of them really test the system. I've had more luck building automation frameworks with something like Playwright if there's a front end, or building an automation framework for the back end that actually tests the system. Even then, I've had to argue with it a little. "No, I don't want you to build something that just pings each individual component. I want you to create a synthetic transaction, run it through the system, and look at what happens."

Dave [8:31]: Use cases and behavior-driven testing, not just functional.

Peter [8:37]: Which was always good practice anyway. It's not new. These tools just make it easier to do, once you talk them into doing it.

Overengineering and Extra Abstraction Layers [8:46]

Dave [8:46]: For sure. Can you talk a little about the level of complexity you see? This is something I see a lot too, but I do less of the coding and more of the artifact generation, training, and the various other things we pull together. On the software side, you mentioned enterprise, and you mentioned simplicity. I think of agile and the art of simplicity, keeping things simple.

Peter [9:13]: Part of this is, on occasion, me not keeping the context as clean as I should. But it's clear that it'll look to solve a problem and, even knowing the size or scale of the problem, decide the right solution is another abstraction layer, because that's what best practice says when it looks in its model. So it introduces yet another abstraction layer, and that just creates another level of complexity. If what you're building has a very small use case, or not a large number of users, that sometimes just isn't the right architecture. You look at it and think, "Wait a minute, this is way more complicated than it needs to be, with far too many layers of abstraction and contracts between different components." It's like, "I'm just trying to run a simple website with a single form on it. Why on earth are you building all this?" It wasn't quite that simple, and I won't go into the details of what it was, but it built many layers of abstraction into a system. I went back, refactored it, and I think I removed five different layers. You need an understanding of what you're doing and whether the solution it's building is appropriate for the use case. Which maybe brings you back to the same thing: understanding the problem.

Dave [10:40]: I'm struck by how expensive that can be if every user is going through that same learning curve of over-delivering on the system and then having to pare it back. That's a very expensive learning curve if it isn't being shared and addressed across the organization.

Peter [11:02]: If everyone's doing that, yes. But I think in a lot of organizations it is shared and addressed, and there are common pieces. Some of what I'm describing is my own experimentation as I build these things out. I'm using it as a learning exercise so I can help clients figure out how to solve this and where the limits are. And yes, figuring that out can be very expensive. I wonder, though, once everybody gets this figured out, does that mean they suddenly start using it less? Or do they finally reach the utopian world we're being sold, where we can type in a single sentence and produce an entire software suite?

Dave [11:48]: I'm not sure that's happening just yet, but we'll see.

Peter [11:51]: Not any time soon, from what I've seen, despite what they would love to tell us. My experience is that's not the case. Not unless you've got unlimited compute, essentially. And money.

Three Takeaways and Holding the Reins [12:06]

Dave [12:06]: Three takeaways?

Peter [12:08]: I think the first is that the importance of managing context is very, very clear. Second, there's still a deep need to understand the problem you're solving, and to have somebody with enough expertise, or at least the ability to ask the right questions, to challenge what's being produced. Is this fit for purpose? Is what's being built, produced, or managed actually targeting the size and scale of the problem I'm trying to address? Or have we done something far more complicated than needed? Even just asking that question can be valuable, and it's worth seeing what you get back. Of course, it wants to keep you happy, so it's always going to find something.

Dave [12:59]: I was going to add to that. One of the things that struck me during this conversation is how much we still have to define what the problem is and how it's going to be solved. To your last comment, we're nowhere close to writing a single line that says, "Go build a product like this," and having that happen. I think there's a huge amount of holding on to the reins that we sometimes forget we have to do. Whether it's understanding context and making sure what's captured is still relevant and correct, right through to what we should be validating, what we should be building, what it should look like, and what knowledge we bring to the table.

Peter [13:49]: There's a judgment piece there. You can have it present things for you to judge: "These are the decisions I think need to be made." But you still need the judgment to look at it and say, "Actually, you got this wrong." And that requires knowing enough about the problem yourself. That's the tricky piece.

Closing [14:17]

Peter [14:17]: Well, thank you as always, Dave. It's always fun having these chats, and I look forward to next time. For all our listeners, don't forget to hit subscribe, and you can reach us at feedback@definitelymaybeagile.com. Until next time, thanks again. You've been listening to Definitely Maybe Agile, the podcast where your hosts Peter Maddison and Dave Sharrock focus on the art and science of digital, agile, and DevOps at scale.



validate AI output,AI agents,AI software delivery,Definition of Done, test automation, context management, behavior-driven testing, human judgment,Definitely Maybe Agile,agile podcast,devops podcast,peter maddison,dave sharrock, agile transformati,