Every regulated team knows it has to test. Fewer can prove, on any given day, that the system in production is exactly the one they validated. That critical gap is what continuous compliance closes, and it is getting wider as agents write more of the code. Engineering and quality leaders in compliance-driven organizations, such as life sciences, financial services, and insurance, are shipping software faster than ever. But the rules for proving that software works have not relaxed, and working with AI has introduced more surface area requiring validation.
This article examines three key areas: what compliance actually requires from your software testing, how to maintain that compliance between audits, and how to maintain compliance when using AI coding agents. We will illustrate through GxP and pharma examples, as they represent one of the most demanding versions of the problem. If your process would satisfy an FDA inspector, it will hold up under SOX, DORA, or a banking supervisor's review.
Continuous compliance is the practice of proving, at any point in time, that your systems meet regulatory requirements, instead of reconstructing that proof before an audit. While most definitions of continuous compliance center on continuous compliance monitoring of infrastructure controls, security policies, or cloud posture, software delivery requires a different focus.
For teams that build and frequently change software, the hardest thing to prove continuously is that the application still works as intended after every single change. That fundamentally makes it a testing problem.
The FDA uses the term "validated state" to describe a system that has been proven to work and remains under strictly controlled change management. This is a useful frame even outside the pharmaceutical industry. When you modify your application without corresponding evidence of functionality and security, your system falls out of its validated state.
To prevent this, testing must generate an unbroken chain of machine-generated proof. GRC tooling matters for overarching policy monitoring, but proving functional correctness through rigorous software testing is the critical piece that GRC tools do not cover.
| Dimension | Point-in-time compliance | Continuous compliance |
|
When proof is created |
Before an audit or a release gate |
Every change, every run |
| Playwright | Manual screenshots, signed protocols | Machine-generated, timestamped records |
| Response to change | Revalidate in batches, or skip | Risk-based testing on each change |
| Gaps between audits | Unknown drift |
Drift detected and recorded |
| Cost Curve |
Unknown spikes at every audit | Inner to mid-loop coverage |
Regulators rarely prescribe exactly how to execute your tests; rather, they prescribe what you must be able to prove. This distinction is what separates software compliance testing from ordinary functional testing. When you strip away the industry-specific jargon, frameworks across highly regulated sectors typically ask for the same five foundational pieces of evidence from your testing.
21 CFR Part 11 serves as an excellent anchor for understanding these requirements because it is explicitly clear. The FDA mandates "secure, computer-generated, time-stamped audit trails" (11.10(e)), rigorous validation (11.10(a)), strict access controls (11.10(d)), and retrievability (11.10(c)). In the life sciences, regulated teams often use the ALCOA+ framework (Attributable, Legible, Contemporaneous, Original, Accurate, plus Complete, Consistent, Enduring, and Available) as shorthand for maintaining strict data integrity.
An audit trail is a secure, computer-generated, timestamped record of who did what, and when.The need for a verifiable audit trail is the common thread among every major regulatory framework. For example, you see similar strict mandates in SOX IT general controls regarding change management and segregation of duties.
HIPAA Security Rule audit controls demand tracking of system activity (45 CFR 164.312(b)). PCI DSS v4.0 Requirement 6 focuses heavily on rigorous change control and testing prior to deployment. SOC 2 CC8.1 emphasizes structured change management, and the EU's Digital Operational Resilience Act (DORA), in force since January 2025, requires financial entities to run documented, risk-based ICT change management (Art. 9) and a testing program that explicitly includes end-to-end and performance testing, with tests at least yearly on systems supporting critical functions (Art. 24–25).
Producing audit-ready evidence for regulated teams requires aligning testing practices with these uncompromising evidentiary demands.
|
What you must prove |
What it means for testing |
Where it shows up |
|
The system works as intended |
Tests mapped to requirements and intended use, executed and passed |
21 CFR 11.10(a), GAMP 5, SOC 2 CC8.1, DORA Art. 25 |
|
Who did what, and when |
Attributable, computer-generated, timestamped records of every run |
21 CFR 11.10(e), HIPAA 164.312(b), ALCOA+ |
|
Nothing was hidden or altered |
Every run kept, pass or fail; no selective re-runs |
21 CFR 11.10(b), FDA data integrity guidance |
|
Changes were verified before release |
Proof of sequence: test ran and passed before deploy |
SOX ITGC change management, PCI DSS Req 6, DORA Art. 9(4)(e) |
|
Independent review |
Separation between who builds, who tests, and who approves |
SOX segregation of duties, GxP independent review |
"Good x Practice" (GxP)—such as GMP, GLP, and GCP—are the strict regulations for products that directly affect human health, stringently enforced by global agencies including the FDA, EMA, MHRA, and PMDA.
Recent regulatory research analyzing over 1,700 FDA warning letters issued between 2016 and 2023 showed that data integrity deficiencies remain the largest driver of enforcement, consistently appearing in 60% to 80% of drug GMP warning letters. This ongoing scrutiny illustrates exactly why pharmaceutical compliance remains the ultimate stress test for software delivery.
When you examine what the FDA actually cites during enforcement, they are overwhelmingly highlighting testing-evidence failures:
Shared logins, meaning actions simply cannot be attributed to a specific individual.
Audit trails that were disabled by users or never turned on in the first place.
Failing test runs being deleted and continuously re-run until they finally produce a passing result.
Users possessing the administrative ability to change the local system clock.
Total lack of necessary revalidation following a system change or patch.
Consider the outdated manual testing picture still common today: a tester takes a screenshot of the interface ensuring the system clock is visible as timestamp proof, pastes it into a Word protocol document, signs it, and a second person co-signs for review. This cumbersome process is repeated for every minor patch. These manual, easily manipulated processes create the evidentiary gaps inspectors are trained to hunt for.
The consequences are staggering. In 2013, Ranbaxy was forced to pay $500 million in criminal and civil penalties resulting from severe compliance failures. In 2023, Intas Pharmaceuticals was placed on a crippling FDA import alert following a warning letter that cited extensive data manipulation. If you swap "FDA inspector" for a SOX auditor, a HIPAA assessor, or a DORA supervisor, the fundamental questions they ask about data integrity and system validation are remarkably consistent.
Regulators are increasingly recognizing that excessive documentation burdens hinder modern software delivery, which is why the FDA has moved from traditional Computer System Validation (CSV) to Computer Software Assurance (CSA). Final guidance for this risk-based successor was published on September 24, 2025, and subsequently updated on February 3, 2026, to align with QMSR/ISO 13485.
The central question in CSA is straightforward: "If this specific function fails, what is the absolute worst that happens to a patient or to product quality?" Testing rigor strictly scales with that identified risk. High-risk functions require robust, fully scripted testing, whereas lower-risk functions may allow for unscripted or exploratory testing methods.
Crucially, the CSA guidance lists automated testing tools and vendor-supplied evidence as highly acceptable assurance methods. It mandates an assurance record containing six elements: intended use, risk determination, testing performed, conclusion, who tested and when, and review/approval. Importantly, the 2026 update specifically notes that this assurance approach applies to automation tools, bots, AI/ML tools, and cloud software. Regulators are explicitly moving toward continuous, risk-based proof, making traditional, point-in-time validation increasingly obsolete.
Achieving continuous compliance requires integrating specific evidence-generating practices into your daily engineering workflows.
Start from intended use and risk. You must classify every application function by its potential impact—whether that involves patient safety, transaction and payment integrity, financial reporting accuracy, or PHI security. Spend your most rigorous testing efforts exactly where failure hurts the organization most.
|
Practice |
Evidence it produces |
|
Intended use and risk |
Risk determination per function |
|
Traceability |
Requirement-to-result links |
|
Pipeline automation |
Timestamped proof that the test passed before deploy |
|
Immutable evidence |
Complete audit trail of every run |
|
Separation of duties |
Pre- and post-execution approvals |
|
Post-release testing |
Continuous record of validated state |
|
Vendor and test-data governance |
Data-location records, subprocessor list, and test-data policy |
The scale of software change is undergoing a transformational shift. A July 2026 Veracode GenAI Code Security Report found that AI now generates roughly half of all committed code, yet averages a concerning 56% security pass rate. The Google's 2025 DORA (DevOps Research and Assessment) report similarly observed that 90% of respondents utilize AI at work, noting that AI adoption correlates with lower delivery stability unless teams maintain strong testing and exceptionally fast feedback loops.
Every established compliance framework fundamentally assumes a human made the code change, a different human independently verified it, and a secure record transparently shows that sequence. AI coding agents severely strain all three of those assumptions simultaneously.
Every code change must leave the underlying system in a securely validated state. When AI agents open dozens of pull requests a day, traditional batch revalidation becomes a bottleneck. The massive change volume guarantees failure if teams attempt to manually revalidate after every patch.
When an agent commits code, was the action performed by the developer, the agent itself, or the service account the agent runs as? A shared agent credential functions exactly like the modern equivalent of a shared login—a classic compliance failure that prevents regulators from tracing actions to accountable individuals.
When the same agent with the same harness writes the application code, generates the corresponding tests, and reports them as passing, you have lost independent verification. To an auditor, this represents a separation-of-duties failure.
AI agents are systematically built to iterate rapidly until internal checks pass. However, without a permanent, immutable record of every single failed attempt leading up to the success, agent iterations look exactly like the "delete the failing run and re-run" pattern that FDA inspectors actively penalize.
Tests generated on the fly by an agent cannot reliably demonstrate what was verified last Tuesday. Audit evidence must unequivocally point to a stable, definitively versioned test case and its specific execution result to prove reproducible evidence.
Because of these severe governance gaps, Gartner recently predicted that over 40% of agentic AI projects will be canceled by the end of 2027 due to spiraling costs, unclear value, or inadequate risk controls. Teams will inevitably use coding agents, but the vital evidence layer must remain entirely independent of them.
|
Compliance expectation |
What coding agents change |
The familiar failure mode |
|
Every change verified before release |
Change volume multiplies |
Failure to revalidate after changes |
|
Actions attributable to a person |
Agents act through shared tokens or service accounts |
Shared logins |
|
Independent verification |
Agent writes code and tests and reports its own result |
No separation of duties |
|
Complete record, pass or fail |
Agents retry until checks pass |
Selective deletion of failures |
|
Reproducible evidence |
Tests generated on the fly |
Missing or unreliable audit trail |
AI coding agents can dramatically accelerate development work, but the verification process and the resulting evidence must purposefully sit outside of those agents. To keep your workflows audit-ready, you must implement strict safeguards:
Applying this independent verification is essentially risk-based assurance applied to an entirely new source of code change. It perfectly aligns with the FDA's 2026 CSA guidance, which deliberately extends the assurance approach to automated bots and AI/ML tools.
Running tests is now merely table stakes. The true differentiator for regulated teams is the unassailable evidence produced around those tests. mabl provides an independent verification layer that sits safely outside your coding agents, bringing its own testing harness that allows agents to assist in test authoring while mabl strictly runs and records the outcomes.
mabl automatically generates evidence for every run, so teams stop gathering it by hand, including screenshots, DOM snapshots, network traces, and request and response details for API tests and API steps inside end-to-end flows. Every record uses secure cloud-side timestamps that no user can maliciously or accidentally alter. The platform maintains an audit trail that simply cannot be disabled, permanently retaining every test run regardless of whether it passed or failed. By deeply integrating with CI/CD pipelines, mabl executes tests on every single code change and securely gates deployments, establishing undisputed proof of sequence. Additionally, scheduled runs continue validating your production environment, ensuring the system remains in its validated state post-release.
To maintain strict attribution, mabl enforces individual authentication, robust SSO, and role-based access controls. It automatically pushes execution results to systems like Jira, keeping requirement-to-test traceability and pre/post-execution approvals perfectly aligned via Xray results reporting. Relying on mabl also provides vendor-supplied evidence, such as SOC 2 Type II reports, annual third-party penetration testing, encryption in transit and at rest, and architecture documentation for your security review, directly supporting your CSA assurance records and DORA third-party risk assessments.
When mabl agents build, run, analyze, or update tests, teams have full visibility into each decision and its rationale, and can set system-wide, granular controls over how those agents behave through agent instructions. That keeps agents inside the same separation-of-duties and attribution controls as the people they work alongside.
See how mabl turns every test run into audit-ready evidence, even when agents are writing the code. Book a demo.