Skip to main content
Flaky TestsCI TrustTest Ownership

Why Your CI Pipeline Is Full of Failed Tests Nobody Trusts

2 September 2026 · OpenCrevo

Ask a developer on almost any mid-sized engineering team what a red CI check means, and you'll get a shrug before an answer. "Probably flaky, just re-run it" is now a reflexive response, not a diagnosis. That reflex is the actual problem. Flaky tests kill CI trust one ignored failure at a time, and once trust is gone, the pipeline stops doing its job even though it's still running every commit. This article explains why that erosion happens, why the usual fixes don't hold, and lays out a concrete re-trust playbook: a flake-rate threshold model, an explicit ownership structure, and a rollout sequence a team can actually run, not just another explanation of why CI trust erodes.

The Real Problem

A pipeline with a 40% chance of a false-red result on any given run isn't a QA inconvenience, it's a broken instrument. If a smoke alarm cries wolf four times out of ten, people stop evacuating when it goes off, and eventually it goes off during a real fire and nobody moves. CI pipelines fail the same way, just slower and quieter.

Here's the pattern almost every team recognizes once it's named: a merge queue backs up because three unrelated PRs all show a red check on the same intermittent test. Someone reruns the job. It goes green. Nobody files anything, because filing something takes longer than clicking rerun, and rerun usually works. Multiply that by every PR, every week, for a year, and you get a team that has fully stopped reading failure output before clicking rerun. The signal is still being produced. Nobody is receiving it anymore.

The cost isn't abstract. A real regression eventually rides through on the back of "it's probably flaky," because the person looking at the failure has no fast way to tell a genuine flake from a genuine bug (Industry consensus: this exact ambiguity is why teams need a structured triage step rather than gut feel, covered in depth in "Test Failure or Product Bug? A Practical Failure-Triage Workflow). By the time someone notices in production, the CI run that should have caught it is three "just rerun it" clicks in the past.

Why This Happens

CI trust doesn't erode from one big failure. It erodes from an accumulation of small, cheap, individually-reasonable decisions:

  • Flaky tests get added faster than they get fixed. Every new feature ships new tests. Almost nobody schedules time to stabilize old ones. The flaky population only grows (Industry consensus).
  • Rerun is free, investigation is not. A developer under a PR deadline has a button that costs one click and usually works, versus an investigation that costs unknown minutes and might not even resolve anything. Rational actors take the cheap option every time, and each individual choice looks reasonable in isolation.
  • No one owns the suite as a whole. Individual engineers own individual tests they wrote. Nobody owns the aggregate flake rate, the average time-to-green, or the trend line. What isn't owned doesn't get fixed; it just gets tolerated (Opinion, but a pattern this article's authors have seen repeatedly across teams that lack a defined test-suite owner).
  • Failures are undifferentiated. A CI dashboard that shows the same red "X" for a genuine regression, a known flake, a stale assertion, and an infrastructure blip gives the reader no reason to treat any of them differently, so they all get the same response: rerun and move on.
  • There's no visible cost to ignoring a failure, until there is. The regression that finally slips through doesn't get attributed back to "we've been ignoring red checks for a year." It gets treated as an isolated incident, and the underlying trust problem survives the postmortem untouched.

None of these five causes is a testing-tool problem. They're organizational: incentives, ownership, and visibility. That's exactly why swapping test frameworks or adding more assertions doesn't fix it.

Common Approaches That Fail

  • "Just write better tests." True in the abstract, useless as a plan. It doesn't say who fixes the 40 already-flaky tests sitting in the suite today, on what timeline, or how progress gets measured (Opinion).
  • Quarantine and forget. Moving flaky tests to a "known flaky" bucket stops them blocking merges, which is necessary, but teams routinely stop there. The quarantine bucket becomes a graveyard nobody revisits, and the underlying defects (in the test or the product) never get fixed (Industry consensus).
  • Mandating zero reruns. Some teams try to force discipline by banning reruns entirely. This looks rigorous on paper and collapses in practice, because it blocks legitimately flaky-but-not-yet-fixed tests from ever going green, which just pushes engineers toward disabling tests outright, a worse outcome than a rerun.
  • A dashboard nobody is required to look at. Visibility tools (flake-rate charts, trend graphs) get built, demoed once, and then ignored, because visibility without an ownership model attached to it doesn't create accountability. A chart nobody is on the hook for looking at changes nothing.
  • Vague "improve test quality" OKRs. A quarterly goal to "reduce flakiness" without a numeric threshold, an owner, and a review cadence is not a plan, it's an aspiration, and aspirations don't survive contact with the next deadline (Opinion).

Practical Solution

Re-establishing CI trust requires treating flake rate as a monitored, owned metric with an explicit SLO, the same way a platform team treats uptime. This is the piece most narrative content on this topic stops short of: a concrete, numeric ownership model.

The flake-rate SLO model

Define flake rate as a rolling metric per test and per suite:

flake_rate = (runs where the test failed then passed on an unchanged retry) / (total runs)

Set tiered thresholds tied to consequence, not vibes:

  • Green: Under 1% flake rate (rolling 30 days), no action, monitor only
  • Watch: 1% to 5% flake rate (rolling 30 days), auto-tagged, added to the weekly stabilization queue, does not block merges
  • Quarantine: 5% to 15% flake rate (rolling 30 days), removed from the required-check gate, owner assigned, 2-week fix-or-delete deadline
  • Red-list: Over 15% flake rate (rolling 30 days), or unresolved past its quarantine deadline, deleted or fully rewritten, not merely skipped indefinitely

The exact percentages are a starting point (Opinion), calibrate them against your own suite's baseline the same way the failure-triage article recommends pulling a month of historical failures before building a workflow. What matters structurally is that every tier has a number, an owner, and a deadline. None of this works as a one-time cleanup; it has to run as a continuous process (Industry consensus: this mirrors how mature platform teams manage error-budget-style SLOs, applied here to test reliability instead of uptime).

The ownership model

Three roles, not one overworked "QA person":

  • Test owner (the engineer who wrote or last touched the test): responsible for triaging a Watch or Quarantine tag within the deadline. This is the person closest to the code, not a separate QA function reading a backlog.
  • Suite owner (a rotating role, weekly or biweekly): responsible for the aggregate flake rate trend, reviewing the quarantine queue, and escalating anything past deadline. This role is what "no one owns the suite as a whole" was missing.
  • Release gatekeeper (usually an engineering manager or tech lead): the person who enforces that Red-list tests actually get deleted or rewritten, not silently left skipped forever. This is the person whose sign-off makes the SLO real instead of decorative.

Without all three, the model degrades back into the dashboard-nobody-looks-at failure mode above.

Implementation

Start with visibility, because you can't gate on a number you're not tracking. A CI job step that classifies a run and writes flake data somewhere queryable:

name: CI

on: [pull_request]

jobs:
 test:
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
 with:
 node-version: 22
      - run: npm ci
      - name: Run tests (attempt 1)
 id: attempt1
 run: npx playwright test
 continue-on-error: true
      - name: Retry once on failure
 id: attempt2
 if: steps.attempt1.outcome == 'failure'
 run: npx playwright test
 continue-on-error: true
      - name: Record flake signal
 if: steps.attempt1.outcome == 'failure' && steps.attempt2.outcome == 'success'
 run: |
 echo "flaky_run,${{ github.sha }},${{ github.run_id }},$(date -u +%Y-%m-%dT%H:%M:%SZ)" >> flake-log.csv
      - name: Fail the job if the retry also failed
 if: steps.attempt1.outcome == 'failure' && steps.attempt2.outcome == 'failure'
 run: exit 1
      - name: Upload flake log
 if: always()
 uses: actions/upload-artifact@v4
 with:
 name: flake-log
 path: flake-log.csv

That artifact feeds a simple aggregation (a scheduled job or a scheduled script reading uploaded artifacts across recent runs) that computes rolling flake rate per test name and applies the tier thresholds above. The mechanism (spreadsheet, database table, or a dedicated internal tool) matters far less than the discipline of actually running it weekly and having the suite owner act on it.

Roll it out in this order, not all at once:

  • Week 1 to 2: instrument only. Add the retry-and-log step above. Don't gate anything yet. Just collect real data on your actual flake population.
  • Week 3: baseline and tier. Pull the data, apply the thresholds, get your real starting numbers. Expect this to be uncomfortable; most teams underestimate their flake rate until they measure it (Industry consensus).
  • Week 4: assign owners. Every test in Watch or Quarantine gets a named test owner. The suite-owner rotation starts here.
  • Week 5 onward: enforce the gate. Quarantine tier tests come out of the required-check list. Their 2-week fix-or-delete clock starts. The release gatekeeper starts reviewing the Red-list queue at a fixed weekly cadence.
  • Ongoing: review the trend, not just the snapshot. A flake rate that's flat or rising after two months means the ownership model isn't actually being enforced, not that the thresholds were wrong.

AI Considerations

An AI agent is well suited to the mechanical layer of this system: classifying a failure as flaky-on-retry, computing rolling flake rates, drafting the weekly quarantine-queue summary, and even opening a stabilization PR for a test with an obvious timing fix (a missing `await`, a race on an unindexed selector). What an AI agent should not be trusted to decide unsupervised is whether a Red-list test gets deleted outright, because that decision can silently remove real coverage for an edge case the model doesn't have enough context to recognize as important (Opinion). Keep the release gatekeeper role human for exactly that reason. The mechanical triage question ("did this fail then pass on an identical retry") is safe to automate; the judgment question ("is this coverage worth keeping in some form") is not.

OpenEvident

If your CI pipeline runs Playwright tests scaffolded with CrevoAI, the sharding and reporting layer for the flake-tracking model above doesn't need to be built from scratch. The OpenEvident GitHub repository publishes `vindicate-actions`, composite GitHub Actions (`setup`, `run-shard`, `merge-reports`) for running CrevoAI-scaffolded Playwright tests in CI with sharded execution and a native job summary, with SHA-pinned dependencies and no third-party data flow (Verified: `openevident-research.md`). A native job summary that clearly shows sharded pass/fail results per run is a meaningful piece of the visibility problem this article describes: it's much harder to reflexively ignore a failure when the summary makes the specific failing shard and test legible at a glance, rather than burying it in raw log output. `[VERIFY API BEFORE PUBLICATION]`: whether `vindicate-actions`' job summary format exposes rolling flake-rate data directly, or only per-run pass/fail; treat that specific capability as unconfirmed until checked against the repo directly.

OpenCrevo Implementation

The ownership model above (test owner, suite owner, release gatekeeper) is a process change, and process changes are usually the part that stalls after the initial enthusiasm of week one. Building the actual gating logic into an existing CI pipeline, wiring the retry-and-log step, standing up the weekly quarantine review, and getting a team to actually enforce fix-or-delete deadlines, is exactly the kind of repeatable, CI-integrated test automation work OpenCrevo's Test Automation service delivers: catching regressions before they reach production through a suite that's actually maintained, not just running. If your team has the CI pipeline but not the discipline around it, OpenCrevo can help design and implement the rollout. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Define flake rate numerically (failed-then-passed-on-retry, over a rolling window) before doing anything else. You cannot gate on a metric you're not measuring.
  • Set explicit tier thresholds (Green, Watch, Quarantine, Red-list) with a consequence and a deadline attached to each, not just a label.
  • Assign all three ownership roles (test owner, suite owner, release gatekeeper) before enforcing any gate. A gate with no owner behind it decays back into a dashboard nobody checks.
  • Give Quarantine-tier tests a hard fix-or-delete deadline, and actually enforce it. An indefinite quarantine bucket is a graveyard, not a fix.
  • Roll out in stages (instrument, baseline, assign, enforce) over roughly a month. Don't gate on day one before you have real data.
  • Review the flake-rate trend weekly, not just the point-in-time snapshot, since a flat or rising trend after enforcement starts is the signal that the model isn't actually being followed.

FAQ

  • Is this worth it for a small team with only a handful of flaky tests? Yes, in a lighter form. A five-person team doesn't need a full rotating suite-owner role, but even an informal version (one person glancing at the flake log weekly, a shared deadline for fixing anything tagged flaky) prevents the reflexive-rerun habit from forming in the first place, which is much easier than reversing it later.
  • How long does this take to implement? The instrumentation (the retry-and-log CI step) is a day of work. Getting the ownership model actually functioning, where deadlines are enforced and the trend is reviewed without prompting, realistically takes 4 to 8 weeks of consistent practice, per the rollout sequence above.
  • Does this replace the need to actually fix flaky tests? No. The SLO model and ownership structure create the pressure and visibility that makes fixing (or deleting) flaky tests actually happen, rather than perpetually deferred. It's the process wrapper, not a substitute for the underlying stabilization work covered in how to reduce flaky Playwright tests in CI/CD.
  • What's the difference between this and just quarantining flaky tests? Quarantine alone removes flaky tests from the merge-blocking path, which is necessary but not sufficient. Without a deadline and an owner attached, quarantine becomes a permanent holding pen. This model adds the deadline, the owner, and the escalation that quarantine alone lacks.
  • How do you measure success after adopting this? Two numbers: the rolling suite-wide flake rate trending down over successive months, and the average time between a real regression landing and someone acting on the CI signal for it trending down as well. The second number is harder to measure directly but is the actual point: a shorter gap means people are trusting and reading red checks again, not just rerunning past them.

Conclusion

A CI pipeline nobody trusts isn't a symptom of bad tests, it's a symptom of an unowned, unmeasured system that everyone has quietly learned to route around. Fixing individual flaky tests helps, but it doesn't rebuild trust on its own, because the underlying incentive (rerun is cheap, investigation is expensive) doesn't change until someone owns the aggregate number and enforces a consequence tied to it. A flake-rate SLO with tiered thresholds and three named ownership roles turns "probably flaky, just rerun it" back into a decision someone is actually accountable for, which is the only way a red check starts meaning something again.

Sources

Playwright documentation on test retries, GitHub Actions continue-on-error and job summary documentation, Google SRE workbook chapter on error budgets and SLOs as a conceptual model adapted here for test reliability rather than uptime, openevident-research.md and opencrevo-research.md in this repository (verified facts on CrevoAI, vindicate-actions, and OpenCrevo's Test Automation service), this program's content-gap-analysis.md (identifying the missing re-trust rollout playbook as this article's original contribution).

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.