A regression gate goes red. Nobody believes it, so someone reruns it. Once developers are asking whether the problem is their code or the tests, the gate has already lost its authority.
We treated that lost trust as tech debt and set out to win it back. That meant choosing the right user journeys, keeping the gate small, and making every test earn its place in the pipeline.
This describes our PR gate at the time of publishing; our approach remains constant even while the user journeys and tests evolve.
When a team stops trusting a gate, the cost keeps growing. Developers rerun it, override it, and eventually build workflows that route around it.
Coding agents are shipping more changes, which puts more weight on the check before a pull request (PR) merges. That check has to be independent of the agent that wrote the change, and it has to be believable.
We picked the journeys first and the tests second.
To choose the journeys, we looked at 30 days of real usage to see what customers do on our platform. We focused on the most important and most commonly used journeys, and kept the set small but thorough so the gate wouldn't slow developers down.
Then we ranked those journeys by the impact a failure would have on the customer. Our product manager, Krista King, defined three tiers. When our instincts and the usage data disagreed, the tiers followed the data.
| Journey tier | Impact of a failure |
| Tier 1 | The customer can't build, edit, or run their tests, or the results are wrong. |
| Tier 2 | The customer loses time and works around it. |
| Tier 3 | The customer notices, but their work isn't blocked. |
The tiers cut our evaluation work roughly in half. They replaced a vague ask ("give me the tests that should always pass") with a standard we could check.
For each journey, we asked four questions about the tests that cover it. A test went into the gate only when all four answers were yes.
Question two catches tests that look like coverage but aren't: an enabled test that hasn't executed in 90 days isn't coverage. When a test like that was worth keeping, we confirmed it still did its job, fixed it if needed, made sure it was stable, and then turned it on.
A basic test checks that something happened, like a state change. A high-value test proves the outcome held up through a save, a backend run, or a full round trip. Examples from our gate:
I started by pointing a coding agent at the problem: map each journey's requirements against the active tests in our workspace. Its first pass came back with 288 tests. The list was thorough, but as a blocking gate it would almost never pass.
Why a gate should stay focused
Our gate blocks a merge if any test in it fails. That's a policy we chose, and a common one. It also means failure rates multiply.
At 99% reliability per test, a 17-test gate passes cleanly about 84% of the time. A 288-test gate at the same rate passes about 5% of the time. You widen a gate by fixing tests, not by adding more of them.
The agent narrowed the list to 17 tests: one or two per journey, prioritized by tier and chosen for the outcome each one proves. The agent was fast at mapping tests to journeys. Deciding what the gate was for was our job.
Every pull request has to pass that small, fast gate before it can merge. Once it merges and reaches our dev environment, the full regression suite runs against it automatically, well before a release is cut. That suite is roughly 240 user interface (UI) tests plus our API suite. The gate keeps merges fast and trustworthy, and the regression suite casts the wide net, with results triaged daily.
We kept two areas out of the gate on purpose. A test that queried a third-party tracker would let someone else's outage block our merges. Recording tests in the mabl Trainer, our desktop app for building a test by clicking through your application, has its own dedicated suite. Everything else the gate doesn't cover still runs in the full regression suite after every merge.
Even 84% isn't good enough. At that rate, about three in every 20 pull requests would fail for no reason, and developers would lose trust in the gate. They would ignore failures and rerun blindly until the suite passed.
So we set the goal on the gate as a whole: how often is it wrong? A failure that catches a real bug is the gate doing its job. What counts against the gate is a false result, a failure with no product defect behind it.
That bar is steep. For a 17-test gate to give a false result on fewer than one in 100 pull requests, each test can fail falsely only about once in 1,700 runs. I spent two months getting individual tests there, one at a time.
An example: our flow to clear filters failed about once every 10 runs. Before touching the test, I checked the product: it worked, and the test was acting before the section had finished loading. That made it a test-side timing issue.
I could have asked for a product change that exposed an element the test could wait for. I chose the smaller fix: an explicit wait for the section's state before clearing filters.
Other failures pointed the other way. Some duplication tests broke because the product's logic had changed on purpose, so we updated the tests to match. Telling a fragile test apart from an intended product change is what keeps the suite clean.
I applied the same approach to our full UI regression suite: about 240 tests across roughly 52 plans. Some tests run in more than one plan or browser, so each run after a merge adds up to about 300 test runs.
I audited that suite daily for three months. Failures that weren't real product bugs dropped to three or four every two to three days. Today, the entire suite regularly passes in full.
We didn't make the gate blocking on day one. It ran as a reporting-only check for about three weeks before we made it required to merge. Tests went in as they proved themselves: the ones that had already held up went first, and the rest joined once we'd fixed and stabilized them.
Every week, I share a report on our mabl-on-mabl suite (the tests we run against mabl itself). It tracks pass rates, flakiness, and failure categories month over month.
Our engineers can also pull test quality scores into their development environment through the mabl MCP (Model Context Protocol) server. That way they can see which tests are losing reliability without leaving their coding session.
When developers believe the gate, a red result sends them back to their own change first.
That matters more as coding agents write more of the change. The author cannot be the verifier: the coding agent that wrote the code shouldn't be the one that checks it. So the check has to sit outside the agent, and mabl's testing harness is already built for that job. A developer can build a feature with a coding agent, verify it with mabl before opening a pull request, and treat passing tests as a natural approval criterion.
Getting here took usage data, careful curation, and months of unglamorous flake-fixing. Our team decides what belongs in the gate and in the regression suite, and mabl runs both.
To build a gate your team trusts, work through these four steps: