Anthropic published its AI-Native SDLC playbook in August, a blueprint for building software when coding agents write most of the code. It opens with a line I agree with: code is no longer the bottleneck.
The playbook structures the software development lifecycle (SDLC) as a loop of six stages: plan, design, build, test, deploy, and maintain. Each stage ends by committing an artifact, including intent.md, spec.md, plan.md, the diff and its tests, the pull request (PR) with its review findings, and the incident record. Coding agents do the work, and humans stay accountable for judgment.
Two of the playbook’s mechanisms, hooks and a verifier subagent, run in the build stage. PR review and approval gates run in the deploy stage. The testing stage has two separate checks:
A feedback loop checks the coding agent's work on each change. Every session checks its own work before a human sees it, using tests, a build, or a screenshot diff. The sample CLAUDE.md instructs, "If a test fails, fix the code, not the test." A separate verifier subagent, defined in the build stage, can run the final check in a fresh context window so the verdict avoids the assumptions that produced the code.
Both checks assume what the verified state should be is already settled. However, the source of the definition is missing, and there’s no documentation around what happens when the application changes on purpose.
For new functionality, the feedback loop is a one-off check. The session runs its tests or compares a screenshot to the mock, then ends. Turning requirements into durable coverage falls to the developer, or to the coding agent writing the code.
For existing behavior, neither check allows requirements or a test's intent to change. The feedback loop's only guidance is to fix the code, not the test, which is the equivalent of telling a PR it can't update the unit tests it broke on purpose. The evaluation suite grows, but no case is ever revised after the requirements behind it move.
The coding agent that wrote the change ends up deciding what correct means. Every stage of the loop commits an artifact, but they are missing one path and one artifact: an AI-native way to identify the test coverage a change needs, and an independent verdict on whether the application still does what the business intends.
In Anthropic’s playbook, the coding agent's plan.md names "the tests that prove it," and the agent writes those tests alongside the code. Coverage comes from the coding agent's own reading of the change, decided while it builds. The feedback loop confirms the work against those checks, or a screenshot against the mock, and the session ends.
What's missing is a path from requirements to coverage. The intent.md and spec.md already state what the business wants, but no step turns them into test cases the business owns. Coverage grows only as fast, and as well, as each coding session happens to write it.
For business-critical applications, that notion doesn't hold. When you can't afford a production failure, a test case is evidence. In FDA-regulated industries like pharma and medical devices, regulations require a record of every change and who signed off on it. Someone will ask which requirement each test covers, and who approved it when it changed.
The playbook acknowledges this. Its sidebar on legacy systems notes that requirements tools with traceability built in are hard to displace precisely because auditors and regulators already accept them. It treats that as an integration problem to route around. It's also a gap in the loop: the record auditors accept sits outside the artifact chain, and nothing in the chain ties a test back to it.
When the product changes on purpose, yesterday's correct result becomes today's failure. Any system that treats the expected result as fixed will spend its time reporting intended changes as regressions.
A change that breaks an existing test breaks it in one of two ways:
The Anthropic playbook doesn’t account for either of these scenarios. "Fix the code, not the test" is right for a bug fix, and the playbook scopes its test-file hook to that case. Then, the sample CLAUDE.md applies the rule to every change. The evals have the same blind spot: each case is checked for "behavior unchanged," and cases get added or retired, but never revised.
The tests have to change, and the coding agent is left to change them. Columbia's DAPLab studied five leading coding agents and found they "prioritize runnable code over correctness," suppressing errors rather than telling the user something went wrong.
The Anthropic’s playbook's answer is the verifier subagent. A fresh context window removes the session's assumptions from the verdict, but it doesn't change who owns the test case. The subagent, PR review, and code owners all read the tests the coding agent wrote. None of them owns the application's test cases, or decides whether a changed test still checks what it was meant to.
The path starts at the requirements, runs parallel to the coding agent's plan, and meets the diff at the verdict. intent.md and spec.md feed business-owned test cases the same way they feed plan.md.
A test case is a statement of intent, such as "A returning customer can check out with a saved card." When the requirements change, the test cases change with them, before any code does. The person who reviews the spec in the design stage approves them there.
Here, the verdict comes from a verification agent outside the coding session, working from the diff on one side and the test cases on the other. The coding agent never writes or edits a test case. When a test fails after a change, the verdict classifies the failure against its test case:
| Failure type | What happened | Failure type | Failure type |
| Execution drift | A button moved, or a flow gained a screen | Unchanged | Updates how the test runs; the verdict stays pass |
| Intended behavior change | The requirements changed | Already updated from the spec | Updates the test to match, then checks the new behavior |
| Real regression | The application no longer does what the test case says | Unchanged | Fails until the code is fixed |
Two rules keep this framework honest, applied at the release gate the same way the Anthropic playbook applies approval-gate hooks in deploy. Call it the acceptance gate:
A test's expected result changes only when its test case does. A test can change how it runs, but its test case changes only through the requirements.
Separation of duties becomes a property of how the work is routed.
The Anthropic playbook gets a lot right: coding agents move fast, humans hold the judgment, and every step leaves evidence. What it leaves out is that requirements can change, and that someone other than the coding agent has to decide how. A path from requirements to coverage, plus an independent verdict, closes that gap.
That’s the job mabl was built for. Independently verify changes by finding or building coverage for your change, and the system runs your checks, analyzes failures, and adapts coverage. If true independence is required to verify AI-generated code, request a demo of mabl today.