When AI can generate a working feature in minutes, what does "done" even mean anymore? I've been thinking a lot about how our traditional Definition of Done is quietly becoming obsolete, and I wrote about why that matters. This one might challenge some assumptions you didn't know you had.
We had a sprint review last year, on the profile service refactor, where the team walked through every story. Green checkmarks everywhere. Tests passing, code reviewed, merged to main, deployed to staging. The product owner signed off. Everyone felt good. Two weeks later, a customer emailed us asking why the confirmation email sometimes said "Hi [first_name]" instead of their actual name. Nobody had caught it. It wasn't in any acceptance criteria. The template rendering was technically correct; it just silently fell back to the raw variable when the profile service returned a 206 instead of a 200. Technically done. Actually broken. That story stuck with me, because it perfectly illustrated something I'd been uneasy about for a while: our definition of done was really just a definition of "code complete." And I think most teams are in the same place, even if they don't realize it yet. What "Done" Actually Measures Ask your Scrum team what done means and you'll get the checklist. Reviewed. Tested. Merged. Deployable. Maybe "no critical Sonar issues" if you're fancy. These are fine things to check. I'm not knocking them. But they all measure the same dimension: did a human produce the right artifact correctly? That's it. That's the whole thing. The implicit assumption underneath every item on that checklist is that the hard part was a human writing code, so quality control meant having other humans inspect the code. Peer review catches logic errors. Automated tests catch regressions. CI/CD gates catch broken builds. All of this made sense when the bottleneck was human output. AI just blew that assumption apart. When the Machine Does the Easy Part I've been using Cursor for a while now, probably longer than most people on my team realized at first. I started with Copilot back in early 2023, moved to Claude directly for more complex generation tasks, and then settled into Cursor as my main environment sometime in mid-2023. It's become the thing I open before anything else in the morning. At this point I've used it across enough projects that I've stopped thinking of it as a novelty and started thinking of it as a dependency, which is its own kind of problem worth a separate article. The experience of using it daily is genuinely weird. Not in a bad way, but weird in the sense that the thing I used to spend most of my time on (the actual writing of code) is now the fast part. I can produce a working implementation of something that used to take me an afternoon in about twenty minutes. And that's where the definition of done starts to crack. The checklist was calibrated to slow human output. When it takes a developer two days to write a feature, the review process, the tests, the QA pass, all of that adds up to maybe 30-40% overhead. Annoying, but manageable. When the same developer can generate that feature in an hour, and the AI also writes the unit tests, and the AI also writes a first-pass PR description, the "done" checklist quietly becomes a rubber stamp. You're checking boxes that the machine already filled in for you. The review step exists to catch what the developer might have missed. But if the developer didn't really write it, what exactly are we reviewing? I asked our tech lead this exact question during a retro last quarter. She said, "honestly, I don't know anymore." That was a pretty honest answer. The Stuff the Checklist Was Never Measuring Here's the thing I keep coming back to. The traditional definition of done was always incomplete, even before any of this. AI just made the gap impossible to ignore. There's a whole category of quality that doesn't show up in code review: Does the feature behave correctly under conditions the developer didn't think to test? Not unit tests. Actual unexpected usage. Does it degrade gracefully when a downstream service is slow, or does it just hang? Is the UX actually usable, or did we build exactly what the ticket said and end up with something confusing? Will someone be able to understand this code in 18 months without the original author in the room? Does it produce correct outputs for inputs that weren't in the acceptance criteria? That last one is where AI-generated code gets especially interesting. When I had Cursor generate the initial implementation of a retry-with-backoff wrapper for our Kafka consumer last fall, it looked great. Passed review. Tests green. Shipped. Three weeks later we noticed it was retrying on non-retryable exceptions, because nobody had explicitly listed the exception types we cared about in the ticket, and the AI made a reasonable but wrong assumption. The code was correct per the spec. The spec was incomplete. And the "done" process had no step for catching that gap. Not a great look for any of us, honestly. What I Think "Done" Actually Needs to Cover Now I've been working through this on a couple of teams, and the framing I keep returning to is: done used to mean "produced correctly," but it needs to mean "works correctly in the real world." That sounds obvious. It's not obvious in practice. Here's roughly what I think needs to get added, especially for AI-assisted development. I've been applying these changes using Acceptance Test-Driven Development as the scaffolding, and it's been the most useful forcing function I've found. Write the Acceptance Tests Before Cursor Writes the Code This sounds like a small process tweak. It isn't. The whole point of ATDD is that you define the observable behavior of a feature before you write any implementation. When you're working with AI tooling, that discipline becomes even more important, because Cursor will happily fill in a plausible implementation the moment you describe what you want. If you haven't nailed down the acceptance criteria first, the AI will make assumptions, and those assumptions will pass your unit tests, because it wrote those too. My current habit is to write the Cucumber feature file before I open Cursor for implementation. Something like this, for a checkout flow: Feature: Checkout confirmation email Scenario: Email uses customer first name when profile returns partial response Given a customer with id "u123" exists And the profile service returns a 206 with first_name "Alice" When the customer completes checkout for order "o456" Then a confirmation email is sent to the customer And the email body contains "Alice" And the email body does not contain "[first_name]" Scenario: Email sends even when profile service is unavailable Given a customer with id "u123" exists And the profile service is unavailable When the customer completes checkout for order "o456" Then a confirmation email is sent using the customer's account email And the email body contains a generic greeting That second scenario is the one nobody writes. It's also the one that bites you in production. Once the feature file exists, I use Cursor to generate the step definitions and the implementation. The key difference is that the behavior was specified by a human before the AI touched anything. The acceptance tests are the spec, not an artifact of it. For the step definitions in Java with Cucumber and Spring Boot, the structure looks like this: @SpringBootTest(webEnvironment = SpringBootTest.WebEnvironment.RANDOM_PORT) @CucumberContextConfiguration public class CheckoutStepDefs { @Autowired private TestRestTemplate restTemplate; @MockBean private ProfileServiceClient profileServiceClient; @Autowired private EmailCaptor emailCaptor; private ResponseEntity<Void> checkoutResponse; @Given("the profile service returns a 206 with first_name {string}") public void profileServiceReturnsPartialResponse(String firstName) { ProfileResponse partial = ProfileResponse.builder() .firstName(firstName) .build(); when(profileServiceClient.getProfile("u123")) .thenReturn(ResponseEntity.status(206).body(partial));...