Skip to main content
Flaky TestsPlaywright CITest Retries

How to Reduce Flaky Playwright Tests in CI/CD

2 September 2026 · OpenCrevo

Every team with more than a handful of Playwright specs eventually hits the same wall: a test passes locally, passes on a re-run, and fails intermittently in CI for no code-related reason. The usual advice, add retries, use traces, wait for the right locator, is correct but incomplete. It solves flakiness in a suite of 50 tests running on one runner. It does not solve flakiness in a suite of 3,000 tests sharded across 40 parallel CI runners, owned by a dozen different teams, where nobody currently knows whether the flake rate is going up or down. This article covers both layers: the tactical fixes (retries, isolation, tracing) and the part most flaky-test guides skip entirely: how to run Playwright reliably at enterprise scale, with sharding strategy, a flake-rate SLO, and an ownership model that keeps flaky tests from silently rotting a suite nobody trusts.

The Real Problem

A test engineer opens a pull request. Their local run is green. CI comes back red on a test they didn't touch. They re-run the job. It passes. Nobody investigates further, because the deadline is today and the failure "isn't real." This happens a few times a week per team. Multiply it across an organization running Playwright on 40 CI runners with 15 teams contributing specs, and you get a suite where a meaningful fraction of failures are quietly assumed to be noise before anyone reads the log. `(Industry consensus)` Once a team develops the reflex of re-running a failed job instead of reading why it failed, the CI pipeline has stopped being a trust signal and started being a formality.

The tactical causes of Playwright flakiness (race conditions, unawaited network calls, shared test state) are well documented. What's missing from most guides is what happens once a single team's 50 flaky-prone tests become an organization's 3,000-test suite spread across dozens of runners: sharding introduces its own failure modes, nobody owns fixing a flake that isn't "their" feature, and without a measured flake rate, "is this suite getting better or worse" is just an opinion.

Why This Happens

Playwright flakiness in CI has two distinct layers, and conflating them is why so many fixes only half-work.

Layer 1: test-level nondeterminism.

  • Timing and race conditions. A test asserts on an element before the app has finished an async state update. Playwright's auto-waiting handles most of this, but a raw `page.click()` followed immediately by an assertion that depends on a network response Playwright didn't wait for will race.
  • Shared or leftover state. Two tests hit the same seeded database row, or a previous test's `afterEach` didn't clean up a created record, so ordering (which differs between local runs and CI's parallel workers) changes the outcome.
  • Environment differences. CI runners are typically slower, more resource-constrained, and running more parallel processes than a developer's laptop. A test with a tight implicit timing assumption survives locally and fails under CI's CPU contention.

Layer 2: CI-infrastructure nondeterminism, the layer most flaky-test articles skip.

  • Sharding side effects. Splitting a suite across N runners (`--shard=1/4`) means tests that were previously isolated by running sequentially now run concurrently against a shared backend, database, or third-party sandbox. A suite that was "not flaky" pre-sharding can become flaky purely because of the shard count change, not because any spec changed.
  • Runner resource contention. Under-provisioned CI runners (shared CPU, memory limits) produce different failure signatures than a dedicated developer machine, especially for tests involving video/trace recording, which adds CPU and I/O overhead per worker.
  • Network variance to third-party or shared test environments. A staging API shared across every CI job in the org behaves differently under concurrent load from 40 runners than it does when one engineer hits it locally.

Treating both layers as one problem is why "just add retries" doesn't fix Layer 2. Retrying a test that fails because of shard-induced contention on a shared resource doesn't remove the contention, it just makes the failure less visible.

Common Approaches That Fail

  • Blanket retries with no accounting. Setting `retries: 3` in `playwright.config.ts` and calling it solved hides the actual flake rate instead of reducing it. A test that fails 2 out of 3 attempts and passes on the third looks green in the CI summary and is still a real, unfixed problem.
  • Quarantining and forgetting. Skipping a flaky test (`test.skip()` or moving it to a `quarantine` tag) is a legitimate short-term move, but without a review cadence it becomes a one-way door. Quarantined tests accumulate, coverage silently erodes, and nobody notices until a real regression ships through the gap a quarantined test used to cover.
  • Increasing timeouts everywhere. A global timeout bump papers over a genuine race condition by giving it more time to resolve, but it also makes every CI run slower, and it doesn't fix the underlying nondeterminism, it just makes it statistically less likely to surface on any given run.
  • Re-running the whole job until it's green. The most common failure mode of all, because it's free and requires no investigation. It also means CI duration and cost scale with flakiness, and the team learns nothing about which specs are actually unreliable.
  • Treating sharding as purely a speed optimization. Increasing shard count to make CI faster without checking whether the newly-concurrent execution pattern introduces contention on a shared resource is a common way flake rate silently rises right after a CI speed improvement ships.

Practical Solution

Fixing flaky Playwright tests at scale requires three things most guides don't put in one place: a concrete triage decision (fix, quarantine, or delete), a measured flake-rate SLO the organization actually tracks, and an ownership model so a flaky test has a name attached to it, not just a suite.

1. The fix-vs-quarantine-vs-delete decision framework

When a test fails intermittently, don't default to "add a retry." Run it through this decision in order:

  • Can you reproduce the failure with a trace? If `npx playwright show-trace` on the failing CI artifact shows a genuine race condition or a missing wait, this is a fix. Fix it now; it's a real bug in the test.
  • Does the test still cover a requirement no other test covers? If yes, and you can't reproduce the cause immediately, quarantine it (tag it, exclude it from the blocking suite, but keep it running non-blocking in CI so its trend is still visible) with an owner and a deadline, not an indefinite skip.
  • Is the underlying feature deprecated, or is this coverage duplicated elsewhere in the suite? If yes, delete it. A flaky test covering something already covered elsewhere, or something no longer in active use, is pure liability with no offsetting value.

The failure mode to avoid is skipping step 1. Teams under deadline pressure jump straight to quarantine, and quarantine without a deadline becomes permanent deletion by neglect, minus the honesty of actually deleting it.

2. A flake-rate SLO, not a vibe

`(Opinion)` Most organizations running Playwright at scale have no numeric answer to "is our test suite getting more or less reliable this quarter." That should change. Define flake rate per test as reruns-until-pass divided by total runs over a rolling window (e.g. the last 200 executions), and track it per suite and per team, not just as one org-wide number that hides which team's suite is actually the problem.

A workable starting SLO: any test with a flake rate above a threshold (a reasonable starting point is 5 percent over a rolling 200-run window, `(Opinion)`, calibrate to your own suite's tolerance for false CI failures) automatically gets flagged for the fix-vs-quarantine-vs-delete review above, rather than waiting for someone to notice it in conversation. Playwright's built-in JSON reporter gives you the raw per-test pass/fail/retry data to build this; you don't need a separate paid flaky-test dashboard to get a first version of this working.

// playwright.config.ts
import { defineConfig } from '@playwright/test';

export default defineConfig({
 reporter: [
 ['list'],
 ['json', { outputFile: 'test-results/results.json' }],
 ['html', { open: 'never' }],
  ],
 retries: process.env.CI ? 2 : 0,
 workers: process.env.CI ? '50%' : undefined,
 use: {
 trace: 'retain-on-failure',
 video: 'retain-on-failure',
 screenshot: 'only-on-failure',
  },
});

The `json` reporter output includes each test's `status`, `retry` count, and duration. A small script parsing that artifact across recent CI runs (stored, for example, as a build artifact or pushed to a lightweight datastore) is enough to compute a rolling flake rate per test without adopting a new vendor product.

3. An ownership model for flaky tests at scale

This is the part almost never covered in Playwright-specific flaky-test content, because it's an organizational problem, not a Playwright API problem, but it's the actual blocker once a suite passes a few hundred tests written by multiple teams:

  • Every spec file has a named owning team, tracked via a `CODEOWNERS` entry or an equivalent mapping, not "whoever wrote it originally and has since moved teams."
  • A flaky test's quarantine ticket is assigned to the owning team automatically, with a default SLA (for example, 2 sprints) after which an unaddressed quarantined test is escalated to a platform/QA-infra team for a fix-or-delete decision, so it doesn't sit forever.
  • Flake rate is a visible team-level metric, not just a suite-level one. A team whose specs make up a disproportionate share of the suite's flakiness should see that in the same dashboard they see build times and deploy frequency, so it's a normal engineering signal, not a special QA-only concern raised in a separate meeting.
  • A platform or test-infrastructure team owns shared fixtures, seed data, and CI runner configuration, since a flake caused by shared database state or shard contention isn't fairly attributable to whichever team's test happened to fail first.

Without an ownership model, the fix-vs-quarantine-vs-delete framework above has no enforcement mechanism: decisions get made in theory but nobody is accountable for actually making them on a schedule.

Implementation

Isolation and retries, configured deliberately rather than as a blanket setting:

// playwright.config.ts
import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
 fullyParallel: true,
 forbidOnly: !!process.env.CI,
 retries: process.env.CI ? 2 : 0,
 workers: process.env.CI ? 4 : undefined,
 reporter: process.env.CI
    ? [['github'], ['json', { outputFile: 'test-results/results.json' }]]
    : [['list']],
 use: {
 baseURL: process.env.BASE_URL,
 trace: 'retain-on-failure',
 actionTimeout: 10_000,
  },
 projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
  ],
});

Each test gets its own isolated `BrowserContext` (Playwright does this by default per test), which prevents cookie/localStorage bleed between tests, one of the most common causes of "passes alone, fails in the full suite" flakiness. If tests still share state, it's usually because they hit a shared backend database or API, not because of anything in the browser context, which is exactly the Layer 2 problem above.

Sharding across many CI runners, done in a way that doesn't introduce new contention:

# .github/workflows/e2e.yml
name: E2E
on: [pull_request]
jobs:
 test:
 strategy:
 fail-fast: false
 matrix:
 shard: [1, 2, 3, 4, 5, 6, 7, 8]
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
 with:
 node-version: 22
      - run: npm ci
      - run: npx playwright install --with-deps
      - run: npx playwright test --shard=${{ matrix.shard }}/8
 env:
 BASE_URL: ${{ secrets.STAGING_URL }}
      - uses: actions/upload-artifact@v4
 if: always()
 with:
 name: blob-report-${{ matrix.shard }}
 path: blob-report
 retention-days: 7

 merge-reports:
 needs: test
 if: always()
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - uses: actions/download-artifact@v4
 with:
 path: all-blob-reports
 pattern: blob-report-*
 merge-multiple: true
      - run: npx playwright merge-reports --reporter html ./all-blob-reports

Two enterprise-scale details worth calling out explicitly, since they're where sharding-related flakiness actually comes from:

  • Each shard needs its own isolated test data, not a shared row in a shared staging database. If shard 3 and shard 7 both create a user with the same seeded email or hit the same rate-limited endpoint concurrently, you'll see flakiness that looks like a test bug but is actually a sharding/data-isolation bug. Namespacing test data per shard (for example, prefixing created records with the shard index or a run ID) removes an entire category of "worked with 4 shards, flaked with 16" failures.
  • `fail-fast: false` on the matrix is deliberate, not an oversight: killing all shards the moment one fails hides the actual failure pattern across the full suite and makes flake-rate measurement impossible, since you never get a complete picture of which shard's tests failed on a given run.

AI Considerations

`(Opinion)` AI coding agents are increasingly the ones writing and modifying Playwright specs, which changes how flakiness gets introduced. An agent generating a new test from a page recording tends to reproduce the exact timing of the recorded session, including waits that happened to work once but encode no real synchronization logic, for example a hardcoded `page.waitForTimeout(2000)` instead of waiting on a specific network response or DOM state. This is a distinct flakiness source from human-authored tests: an agent optimizing for "the recording passed" rather than "this wait condition is actually correct" produces tests that look complete and pass in isolation, then flake under CI's different timing characteristics.

The mitigation is the same fix-vs-quarantine-vs-delete framework applied earlier and more strictly to AI-generated specs before they merge: a PR review gate that specifically checks for `waitForTimeout` calls and unconditional sleeps in any newly generated spec, since these are disproportionately likely to be the flake source once the suite runs at CI scale rather than as a single local run.

OpenEvident

If your team is generating Playwright specs through an AI coding agent already, it's worth knowing that CrevoAI, the local-first Playwright toolkit published under OpenEvident, runs its codegen and browser-automation MCP tools entirely on the developer's own machine (a VS Code/Cursor extension talks to a local MCP server, `runtime-mcp`, which drives Playwright directly, with no cloud job runner in the loop). The same organization also publishes `vindicate-actions`, a set of composite GitHub Actions specifically for running CrevoAI-scaffolded Playwright tests in CI with sharded execution and a native job summary, described in its README as having "no third-party data flow." For a team building the kind of sharded, multi-runner CI setup described above, that's a concrete existing implementation of shard-plus-merge-report plumbing rather than something to build from scratch. `[VERIFY WITH OPENEVIDENT TEAM]`: whether `vindicate-actions` exposes any built-in flake-rate reporting beyond the merged job summary; the confirmed scope is sharded execution, report merging, and the job summary itself.

OpenCrevo Implementation

Rolling out a flake-rate SLO and an ownership model across an organization with dozens of teams contributing Playwright specs is an implementation problem more than a testing-technique problem: it touches CI configuration, CODEOWNERS mapping, and a reporting pipeline that didn't exist before. OpenCrevo's Test Automation service focuses specifically on repeatable, CI-integrated test suites that catch regressions before they reach production; if the difficult part is standing up that measurement and ownership layer across an existing pipeline rather than fixing any one flaky test, OpenCrevo can help build it into your current CI setup rather than replacing it. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Separate flaky-test causes into test-level (race conditions, shared state) and infrastructure-level (sharding contention, runner resource limits) before choosing a fix.
  • Run every intermittent failure through fix-vs-quarantine-vs-delete before defaulting to a retry or a skip.
  • Give every quarantined test an owner and a deadline; escalate unaddressed quarantines to a platform/QA-infra team instead of letting them sit indefinitely.
  • Track flake rate per test and per owning team, not just a single suite-wide number, using data your `json` reporter output already contains.
  • Namespace test data per CI shard so concurrent shards don't collide on the same seeded record or rate-limited endpoint.
  • Use `fail-fast: false` on sharded CI matrices so a flake-rate measurement reflects the full suite, not whichever shard failed first.
  • Add a PR review gate specifically for unconditional `waitForTimeout` calls in AI-generated specs before they merge.

FAQ

  • Do retries actually help, or just hide the problem? Retries are a legitimate mitigation for a test you've already diagnosed and are actively fixing, so it doesn't block unrelated PRs while the fix lands. Used as a permanent substitute for fixing the underlying cause, they hide the real flake rate instead of reducing it.
  • What's a reasonable flake-rate threshold to start with? There's no universal number; 5 percent over a rolling 200-run window is a reasonable starting point `(Opinion)`, adjusted to how much CI noise your organization can tolerate before people stop trusting red builds.
  • Does sharding make flaky tests worse? Not inherently, but it changes what "isolated" means. Tests that never ran concurrently before sharding can start colliding on shared data or shared external services once they run in parallel across runners, so sharding needs its own data-isolation review, not just a performance one.
  • How long does it take to set up flake-rate tracking? A first version, parsing the existing `json` reporter output for retry counts per test, is a few hours of scripting for a team already running Playwright in CI. A full per-team dashboard with SLA escalation is a bigger, ongoing platform investment.
  • Does an ownership model matter for a small team? Less so below a handful of contributors on one suite, where informal ownership already works. It becomes necessary specifically once multiple teams contribute specs to a shared CI pipeline and "who fixes this" stops having an obvious answer.

Conclusion

Reducing Playwright flakiness in CI is two different jobs stacked on top of each other. The first is the well-covered one: fix the race condition, isolate the state, read the trace. The second is the one most guides skip: at the scale of thousands of tests across dozens of sharded CI runners and multiple contributing teams, flakiness stops being a per-test bug and becomes an organizational measurement and ownership problem. Fixing the first without addressing the second gets you a suite where individual fixes keep landing while the overall flake rate quietly drifts, because nobody is tracking it and nobody specific is accountable for it.

Sources

Playwright official documentation on test retries, sharding, and reporters, GitHub Actions documentation on matrix strategies, OpenEvident's `vindicate` and `vindicate-actions` repositories, OpenCrevo services documentation (`src/data/services.ts` in this repository), industry discussion of CI trust erosion from unmeasured flaky-test rates (`Industry consensus`, synthesized from multiple current QA-engineering and platform-engineering publications, not a single primary source).

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.