A regression gate goes red. Nobody believes it, so someone reruns it. Once developers are asking whether the problem is their code or the tests, the gate has already lost its authority.

We treated that lost trust as tech debt and set out to win it back. That meant choosing the right user journeys, keeping the gate small, and making every test earn its place in the pipeline.

This describes our PR gate at the time of publishing; our approach remains constant even while the user journeys and tests evolve.

When nobody believes the gate

When a team stops trusting a gate, the cost keeps growing. Developers rerun it, override it, and eventually build workflows that route around it.

Coding agents are shipping more changes, which puts more weight on the check before a pull request (PR) merges. That check has to be independent of the agent that wrote the change, and it has to be believable.

Start with what users do

We picked the journeys first and the tests second.

To choose the journeys, we looked at 30 days of real usage to see what customers do on our platform. We focused on the most important and most commonly used journeys, and kept the set small but thorough so the gate wouldn't slow developers down.

Then we ranked those journeys by the impact a failure would have on the customer. Our product manager, Krista King, defined three tiers. When our instincts and the usage data disagreed, the tiers followed the data.

Journey tier Impact of a failure
Tier 1 The customer can't build, edit, or run their tests, or the results are wrong.
Tier 2 The customer loses time and works around it.
Tier 3 The customer notices, but their work isn't blocked.

The tiers cut our evaluation work roughly in half. They replaced a vague ask ("give me the tests that should always pass") with a standard we could check.

Four questions to ask of the tests behind each journey

For each journey, we asked four questions about the tests that cover it. A test went into the gate only when all four answers were yes.

  1. Is the journey covered on the surface the customer uses?
  2. Does that coverage run?
  3. Does it check the outcome?
  4. Is the coverage proportionate to the journey's tier?


Question two catches tests that look like coverage but aren't: an enabled test that hasn't executed in 90 days isn't coverage. When a test like that was worth keeping, we confirmed it still did its job, fixed it if needed, made sure it was stable, and then turned it on.

What checking the outcome means

A basic test checks that something happened, like a state change. A high-value test proves the outcome held up through a save, a backend run, or a full round trip. Examples from our gate:

  • A changed credential reaches the runtime environment.
  • A failed run is reported as a failure in the results.
  • An edit to a shared flow reaches every test that uses it.

288 tests was the right answer to the wrong question

I started by pointing a coding agent at the problem: map each journey's requirements against the active tests in our workspace. Its first pass came back with 288 tests. The list was thorough, but as a blocking gate it would almost never pass.

Why a gate should stay focused

Our gate blocks a merge if any test in it fails. That's a policy we chose, and a common one. It also means failure rates multiply.

At 99% reliability per test, a 17-test gate passes cleanly about 84% of the time. A 288-test gate at the same rate passes about 5% of the time. You widen a gate by fixing tests, not by adding more of them.

How the agent got to 17

The agent narrowed the list to 17 tests: one or two per journey, prioritized by tier and chosen for the outcome each one proves. The agent was fast at mapping tests to journeys. Deciding what the gate was for was our job.

Every pull request has to pass that small, fast gate before it can merge. Once it merges and reaches our dev environment, the full regression suite runs against it automatically, well before a release is cut. That suite is roughly 240 user interface (UI) tests plus our API suite. The gate keeps merges fast and trustworthy, and the regression suite casts the wide net, with results triaged daily.

What the gate deliberately leaves out

We kept two areas out of the gate on purpose. A test that queried a third-party tracker would let someone else's outage block our merges. Recording tests in the mabl Trainer, our desktop app for building a test by clicking through your application, has its own dedicated suite. Everything else the gate doesn't cover still runs in the full regression suite after every merge.

The unglamorous work of killing flakiness

Even 84% isn't good enough. At that rate, about three in every 20 pull requests would fail for no reason, and developers would lose trust in the gate. They would ignore failures and rerun blindly until the suite passed.

So we set the goal on the gate as a whole: how often is it wrong? A failure that catches a real bug is the gate doing its job. What counts against the gate is a false result, a failure with no product defect behind it.

That bar is steep. For a 17-test gate to give a false result on fewer than one in 100 pull requests, each test can fail falsely only about once in 1,700 runs. I spent two months getting individual tests there, one at a time.

A 1-in-10 failure is still a failure

An example: our flow to clear filters failed about once every 10 runs. Before touching the test, I checked the product: it worked, and the test was acting before the section had finished loading. That made it a test-side timing issue. 

I could have asked for a product change that exposed an element the test could wait for. I chose the smaller fix: an explicit wait for the section's state before clearing filters.

Other failures pointed the other way. Some duplication tests broke because the product's logic had changed on purpose, so we updated the tests to match. Telling a fragile test apart from an intended product change is what keeps the suite clean.

The same discipline across the full suite

I applied the same approach to our full UI regression suite: about 240 tests across roughly 52 plans. Some tests run in more than one plan or browser, so each run after a merge adds up to about 300 test runs.

I audited that suite daily for three months. Failures that weren't real product bugs dropped to three or four every two to three days. Today, the entire suite regularly passes in full.

Turning the gate on in stages

We didn't make the gate blocking on day one. It ran as a reporting-only check for about three weeks before we made it required to merge. Tests went in as they proved themselves: the ones that had already held up went first, and the rest joined once we'd fixed and stabilized them.

Measuring trust so it doesn't quietly erode

Every week, I share a report on our mabl-on-mabl suite (the tests we run against mabl itself). It tracks pass rates, flakiness, and failure categories month over month.

Our engineers can also pull test quality scores into their development environment through the mabl MCP (Model Context Protocol) server. That way they can see which tests are losing reliability without leaving their coding session.

What a trusted gate means for developers

When developers believe the gate, a red result sends them back to their own change first.

That matters more as coding agents write more of the change. The author cannot be the verifier: the coding agent that wrote the code shouldn't be the one that checks it. So the check has to sit outside the agent, and mabl's testing harness is already built for that job. A developer can build a feature with a coding agent, verify it with mabl before opening a pull request, and treat passing tests as a natural approval criterion.

Getting here took usage data, careful curation, and months of unglamorous flake-fixing. Our team decides what belongs in the gate and in the regression suite, and mabl runs both.

Building your own trusted gate: a checklist

To build a gate your team trusts, work through these four steps:

  1. Find a small set of important user journeys. Base the selection on how real users use your product. Keep the set small but thorough enough to cover every journey that matters to the product and the business. Then rank the journeys by the impact a failure would have on the user.
  2. Audit the tests behind those journeys. Make sure the gate holds high-value tests. If a test only checks that a page rendered, update it to check a stateful outcome.
  3. Squash flakes. Measure how often each test fails with no product defect behind it. Failure rates multiply, so a large set of mostly reliable tests will fail most of the time. Fix flaky tests until the gate as a whole is rarely wrong.
  4. Turn it on in stages. Run the gate as an advisory check first, starting with a small number of tests. Make it required only once the data shows developers can trust it.