Is Spec-Driven Development Just Big Upfront Design?
Definitely, Maybe AgileOctober 08, 2026x
232
00:19:3113.44 MB

Is Spec-Driven Development Just Big Upfront Design?

Spec-driven development with AI is getting a lot of attention. Peter and Dave ask when that structure actually helps you, and when it slows you down. They start with the AI SDLC white paper Anthropic recently published. It lays out a structured, spec-driven approach built for consistency. Peter sees it working best on large legacy systems, the kind with millions of lines of code written over 15 years by thousands of people. For greenfield work, a looser approach of prototyping and exploring o...

Spec-driven development with AI is getting a lot of attention. Peter and Dave ask when that structure actually helps you, and when it slows you down.

They start with the AI SDLC white paper Anthropic recently published. It lays out a structured, spec-driven approach built for consistency. Peter sees it working best on large legacy systems, the kind with millions of lines of code written over 15 years by thousands of people. For greenfield work, a looser approach of prototyping and exploring often fits better. Dave pushes on the piece most teams drop: the feedback loop. Everyone agrees to review and validate, and then the next shiny feature shows up. They also cover what happens when an AI agent follows an over-detailed spec too literally, why systems thinking beats chasing every edge case, and how explore versus exploit plays out differently for a startup and a company with 10 million customers.

This week’s takeaways:

  • A structured, spec-driven SDLC pays off on large, complex systems where stability matters, but it can feel like far too much when you are just prototyping something new.
  • Define enough up front to guide the work, then rely on regular check-ins and feedback loops to course-correct, because no spec can predict every case and an AI agent will happily take the shortcut you never meant to allow.
  • Explore and exploit are not either/or. Organizations can now run both under one roof, as long as they know where they sit on the path from startup to established player and use judgment about how much structure to add.

Listen to the full episode at definitelymaybeagile.com
 Subscribe so you never miss an episode.

 Have a question or topic you’d like us to cover? Reach out at feedback@definitelymaybeagile.com

New episodes released every Thursday to challenge your thinking and inspire action.

Listen and subscribe:

Welcome and Today’s Question [0:04]

Peter Maddison [0:04]: Welcome to Definitely Maybe Agile, the podcast where Peter Maddison and Dave Sharrock discuss the complexities of adopting new ways of working at scale. Hello, Dave. How are you?

Dave Sharrock [0:13]: Peter, great to catch up again. I’m looking forward to a bit of a meander today on SDLCs and explore versus exploit. I think we’ve been trying to find a nugget there, and I think we’ve got something. How would you frame it?

Anthropic’s SDLC and Structured Delivery [0:29]

Peter Maddison [0:29]: So this partly came out of a white paper Anthropic published on the SDLC. It’s had quite a bit of attention, and it has come up in a lot of the conversations I’ve been having with customers and others. When I read through it, my first reaction was, yeah, this is pretty much what I’ve been saying for a long time. They did frame it well, though. As we were talking about before we started, it’s a structured approach to the SDLC, and it’s very much intended to drive consistency out of the delivery process. And as we’ve discussed on the podcast, there are different ways to go about developing software, with AI or without it. With AI especially, there’s the spec-driven approach: define all of my intent, my spec and my context, get everything curated and managed, and only then deliver anything. That’s very different from saying, here’s a fuzzy idea of where I think I want this system to go. I want to explore it, shape the concepts, have AI build some prototypes, and then feed those prototypes back in to help direct it further as it builds out and evolves the solution.

Spec-Driven Work vs. Small Increments [1:57]

Dave Sharrock [1:57]: Can I jump in? I haven’t gone through the whole paper yet. I’ve seen it come through, but I haven’t sat down with it. A couple of things come to mind from the way you’re describing it. We’ve spent a lot of our careers encouraging small changes, the cumulative effect of small changes over time, instead of big upfront design. When you talk about a spec-driven approach and an SDLC framework being put in place, it just screams big upfront design to me. So how are you looking at that? Is the upfront design flexible enough, and does it accept that things will change?

Peter Maddison [2:42]: I think this is where the fine line comes in, as it always has. We would do the exploration and document what things should look like. Then we’d take the plan and turn it into a spec. We’d take that spec, break it into smaller pieces, and execute against those in sequence. That’s not so different from saying, here’s a bunch of requirements, here are the Jira tickets to execute against them, and here are all the tasks, subtasks and messy component pieces underneath. That’s how a lot of organizations at scale have been operating for quite some time.

Dave Sharrock [3:29]: And we should say successfully operating. Whenever I think about the big organizations I’ve worked with, you need some level of upfront structure and thinking going in. Even with the features and Jira tickets you’re describing, there’s a big chunk of work that happens before that. It gives context and a broader sense of what we’re trying to achieve strategically, and what constraints we’re working within in terms of the systems we use. That thinking has to be done beforehand.

Peter Maddison [4:01]: Yes.

The Feedback Loop Teams Keep Dropping [4:02]

Peter Maddison [4:02]: But there’s a piece we’re not digging into yet, and it’s the part a lot of organizations then fail on. That’s the actual feedback loop. You build a piece of this, a slice that actually has value, and then you talk about it and decide whether you need to build the next bit. The only times you don’t want to work that way are when experimenting is very expensive because it consumes a lot of materials, or when you’ve already done it 10,000 times and you know exactly what the change will be. That’s never the case in software development, because there’s always something new. We know this. It’s old hat to an extent. The conceptual question, as you were describing, is how we build these parts out and where we draw the line on defining too much up front.

Dave Sharrock [5:21]: I want to pick up on the feedback loop piece, because in many of the organizations we work with, this is agreed. Yes, we need to review. We need some sort of retrospective. We need to validate our architecture, our approach, our solution, whatever it is. The reality is that it gets dropped in the excitement of the next shiny object. We go grab the next feature, and very quickly we don’t have time for that validation or health check to make sure the ideas we had at the beginning still hold water. But now it’s a lot cheaper to build that in. You can have a health check, or a set of agents that go off and validate that the assumptions we made when building the SDLC framework are still being met.

Peter Maddison [6:05]: And the Anthropic paper talks to that. I think there are six or so steps, and at the end they go back and validate. They’ve also got operations in there, so they’re looking at not just the immediate validation but the ongoing validation that the service is stable and working the way it was intended. Those are very key attributes.

Legacy Stability vs. Greenfield Discovery [6:30]

Peter Maddison [6:30]: And again, there are two schools of thought. What I’ve seen out there is that spec-driven development works very well when you’ve got a large, complex, older system. Lots of components, millions of lines of code written over 15 years by thousands of different people. In that kind of system, taking the time and care to understand exactly what’s going to change, how it’s going to change, and what the parts are works very well. When you’ve got something greenfield, new or less well understood, the other approach of exploring and prototyping can work very well. That’s not to say the two worlds never cross. In general, each one just fits a little better in its own space.

Dave Sharrock [7:27]: I also wanted to ask about the explore versus exploit angle. When we look at how a system is performing and whether something has to change, it’s very easy to focus only on technical performance. Is it behaving the way we expected when we defined it? But as soon as you add customers, things get a lot more open. Customers might suddenly decide to stop paying subscription fees, or their needs shift in how they use the product. Those changes are much harder to pick out. So whether you’re at the start of product development looking for product-market fit, or you’re established in the market and trying to protect a product and keep it valuable, what do you see happening there?

Peter Maddison [8:29]: I think the ultimate feedback loop is getting feedback from customers. No plan survives contact with the enemy, as they say. Customers will use the system in ways you don’t expect and challenge it in ways you don’t expect. That’s why the sooner you can get something in front of them, the better. But I still know companies that say, we’re not releasing it to customers until it’s perfect, or until this other feature is in. And then, because we’re already delayed, let’s add one more feature too. No. You’ve got to put it in front of the customer. Otherwise you’ll spend three to six months fixing problems you didn’t know you had.

Dave Sharrock [9:18]: Right, problems you don’t know about. One point I’d pick up on: we talked about big upfront planning, and there’s a lot of good in pausing before rushing off to build something. Thinking about how you might build it, and what governance and regulatory requirements you have to work within. The risk is when that turns into big upfront design leading to huge, feature-heavy releases. Then we’re not validating that our customers care about, can use, or need the solution we’re putting in front of them.

AI Requirements and Unintended Consequences [10:00]

Peter Maddison [10:00]: Another interesting way to think about this: if you define the specification in too much detail, you can accidentally steer an AI agent, which operates at much greater speed, in a direction you didn’t want. Say you write that you need to access the database remotely from certain servers. The agent might decide the easiest way to do that is to remove the firewalls and let anything access the database server. That isn’t what you wanted. So you get unintended consequences, and with AI systems that’s a real problem. It’s the same problem you have with requirements. You can’t define all of them up front, because there’s no way to know every possible combination of things that needs to happen. We can define general rules and the things we want to see, but we’ll never define every last detail. So we need to define enough to have confidence it will guide the work, then learn as it builds and course-correct as we go.

Systems Thinking Over Edge Cases [11:17]

Dave Sharrock [11:17]: As you’re describing that, I’m thinking about systems thinking. Systems thinking isn’t about defining how every part of the system behaves, which is what we easily get caught up in. I need to define exactly what happens in this edge case, and this one, and this one. Take the classic example of a thermostat managing the temperature in a room. The point isn’t to list every reason the room might cool down or warm up. The point is to create a mechanism that responds in a predictable, stable way to fluctuations within a certain range. I don’t think we’re good at that. The project management mindset, and the business analysis mindset too, pushes us to define every requirement and every edge case. We end up managing as many use cases as possible instead of designing a mechanism or feedback loop that keeps behavior within a range. That’s what we’re really looking for.

Peter Maddison [12:29]: Yes, exactly. That’s the critical piece, and it’s about finding the balance. What set of systemic guidelines can we give it that keep it on track, the rails it should stay within? Then let it go and figure out how best to do the work. But make sure you’ve got enough check-ins, validations and feedback loops to ask, are we still going in the right direction?

Explore vs. Exploit Across Company Stages [13:02]

Dave Sharrock [13:02]: Can we go a little further into explore versus exploit? I feel like we’re skirting around something really interesting in what you were describing. We talk about experiments and exploring, but what we really mean is validating ideas as we go. That sits at the beginning and the end of the SDLC rather than in the middle. We come up with an idea, pass it through the SDLC, then validate and close the loop. What do you see happening there, or what came out of the paper?

Peter Maddison [13:31]: The Anthropic paper feels like it was designed for a large enterprise, where a lot of these pieces are known quantities. To define them, you need some idea of what they are, and that’s not always the case when you’re starting something new and just heading in a particular direction. For a smaller organization, those definitions will be very light. In a larger, more complex system, the customer needs are also very different. If I’ve got 10 million customers, they want consistency and stability in their experience. They don’t want it changing every day. You do get benefits for experimentation, because you can test with a subset and invite people in, which is harder with a small user base. But it’s a different operating model. That’s how I’d describe explore versus exploit. On one side you’re asking, how do I optimize the system of delivery and get the most out of it? On the other, how do I learn as much as possible, as fast as possible, to evolve the solution toward what customers need? You want learning in the first model too. It just tends to happen at the edges.

Dave Sharrock [15:09]: Yeah, it’s an interesting relationship. I see it a lot with innovation teams, startups and smaller organizations. Just because of where they are and how their customer base behaves, they’re continually exploring. A much bigger share of their time goes into looking for emergent behavior and new opportunities. Whereas a larger, more stable organization with 10 million customers is about consistency. People know why they come to you. How do you keep giving them that service?

Peter Maddison [15:49]: So where we’re landing is that for a consistent, stable, ongoing service, the Anthropic SDLC seems like the better path. Spec-driven development, much more structured. The other end is almost a vibe-coding version of it. They feel like opposite ends of a spectrum right now, though they aren’t necessarily. They’re different approaches to developing solutions, and I think you need both for different problem sets.

Dave Sharrock [16:29]: And I think one of the key takeaways is that you can have both. In the past, the exploring side was often external. It was a company you’d keep an eye on and buy once the technology matured. You couldn’t really have both under one roof, because doing either one well was hard enough.

Peter Maddison [16:53]: Yes.

Dave Sharrock [16:54]: Exactly.

Three Takeaways and Closing [16:55]

Dave Sharrock [16:55]: I feel like we should try to find three key takeaways.

Peter Maddison [16:58]: Sure. Let’s see. There are different SDLCs. The more structured approach works well when you have large, complex features across large, complex systems and you want stability. It can work in other cases too, but it can feel like a lot, especially if you’re just prototyping. Defining all that context up front can feel like an awful lot of work. But as the system grows and gets more complex, it becomes more and more necessary. That’s the balance between the two. And the other takeaway would be applying the right method to the right problem set. What would you add?

Dave Sharrock [17:58]: What really strikes me is the idea of a point in time, and a continuum you move along. The playbook Anthropic has put together for the AI SDLC, and others we’ve had on the podcast talking about the SDLC components they’re building, show you where things are headed. You may not need all of it at the beginning. You have to know where you are in the cycle, from exploration to growth. From startup, to entrepreneurial phase, to dominant player in a market, and every point in between. There’s a path to follow. These frameworks give clear direction on where that path goes, and then we have to use our judgment about how much to put in place.

Peter Maddison [18:59]: Yes, that makes sense. Awesome. Well, with that, thank you as always, Dave. We look forward to next time. To all our listeners, don’t forget to hit subscribe and tell your friends. You can reach us at feedback@definitelymaybeagile.com. Thanks again, and until next time.

You’ve been listening to Definitely Maybe Agile, the podcast where your hosts, Peter Maddison and Dave Sharrock, focus on the art and science of digital, agile, and DevOps at scale.

spec-driven development, big upfront design, AI SDLC, explore vs exploit,AI software delivery,Feedback Loops,systems thinking,Definitely Maybe Agile,agile podcast,devops podcast,peter maddison,dave sharrock,Agile Transformation,Business Agility,