Clouds
CI/CD·4 min read

Quarantining Flaky Tests Without Losing Trust in CI

A test that fails sometimes is worse than no test at all. Quarantine works only when it is a queue with an owner and an expiry, never a mute button.


The fastest way to destroy a test suite is to let it cry wolf. Once a team has watched a red build turn out to be nothing three days in a row, they start treating red as "probably fine" — and the one real failure slides through behind the noise.

Quarantine is the right tool for this, and it is also the tool most likely to quietly make things worse. Done badly it just moves the flakiness somewhere nobody looks.

Why flakiness is worse than an absent test

A missing test tells you nothing. A flaky test actively lies: it reports a signal with a reliability you cannot state, so the correct engineering response to its output is ignore it and re-run. That response, applied to the whole suite, is how a CI pipeline becomes a ceremony.

The cost is not the occasional wasted minute. It is that every other gate in the pipeline inherits the same skepticism.

The quarantine pattern

The mechanism is simple: detect tests that fail non-deterministically, move them out of the blocking set so the pipeline stays trustworthy, and keep running them so you still see their behaviour.

What makes it work is everything around the mechanism.

Quarantine is a queue, not a bin. A quarantined test needs an owner, a date, and a visible position. If it can sit there indefinitely, you have not fixed anything — you have deleted coverage while keeping the file.

Expiry is mandatory. Every entry gets a deadline, and the deadline fails the build when it passes rather than silently extending. Expiry is what converts "we'll get to it" into a scheduled decision.

The count is a reported metric. Put the quarantine size next to coverage on the same dashboard. Watching it climb is usually enough to get someone to act; watching it stay flat tells you the loop closes.

yaml
# the shape that keeps it honest
quarantine:
  owner: team-payments      # not a person who left last quarter
  expires: 2026-10-20       # hard stop, fails the pipeline on this date
  runs_on: nightly          # still executed, just not blocking

Catching them before they land

The reliable detector is the rerun: run the suite twice on the same commit in the same environment, and any test that disagrees with itself is quarantined automatically before it ever blocks a merge. It costs double runtime for the branch, which you can pay only on the integration job rather than everywhere.

The second detector is simpler and free: track per-test failure rates over the last N runs. Anything in the 2%–98% band is flaky by definition. The middle of the distribution is where the damage lives.

What quarantine will not fix

Timing assumptions, shared state between tests, dependence on wall-clock time or network calls, and parallel suites fighting over one database. Quarantine buys you the room to fix those; it is not the fix.

This is worth being blunt about: if the same test re-enters quarantine twice, the problem is in the test design, and the second exit should come with a rewrite rather than a third expiry date.

The cost you are paying for it

Quarantine is not free in CI time either. Running the blocking suite plus the quarantined set on every merge starts to dominate the job, which is why the rerun detector usually belongs on the integration job only — the same trade-off you make with cache restore order and parallelism, where a slower gate that people trust beats a fast one they have learned to bypass.

The larger ledger entry is downstream. A pipeline with known-flaky tests in it cannot serve as a merge gate for anything risky, because everyone has already been trained that red sometimes means "run it again." Security gates depend on the opposite reflex: when a policy check fails, it has to mean the same thing every time. Quarantine exists to protect that reflex — which only works if the quarantine queue is short enough that nobody has stopped reading it.

Summary

Keep the pipeline green by removing non-deterministic tests from the blocking set — but make quarantine a queue with a named owner, a hard expiry that can fail the build, and a size you report like any other quality metric. The goal is not a quiet dashboard. The goal is that a red build means something again, and that every quarantined test has a scheduled conversation waiting for it.

#ci-cd#testing#flaky-tests#quality#devops

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles