Fixing flaky Playwright tests at scale requires three things most guides don't put in one place: a concrete triage decision (fix, quarantine, or delete), a measured flake-rate SLO the organization actually tracks, and an ownership model so a flaky test has a name attached to it, not just a suite.
1. The fix-vs-quarantine-vs-delete decision framework
When a test fails intermittently, don't default to "add a retry." Run it through this decision in order:
- Can you reproduce the failure with a trace? If `npx playwright show-trace` on the failing CI artifact shows a genuine race condition or a missing wait, this is a fix. Fix it now; it's a real bug in the test.
- Does the test still cover a requirement no other test covers? If yes, and you can't reproduce the cause immediately, quarantine it (tag it, exclude it from the blocking suite, but keep it running non-blocking in CI so its trend is still visible) with an owner and a deadline, not an indefinite skip.
- Is the underlying feature deprecated, or is this coverage duplicated elsewhere in the suite? If yes, delete it. A flaky test covering something already covered elsewhere, or something no longer in active use, is pure liability with no offsetting value.
The failure mode to avoid is skipping step 1. Teams under deadline pressure jump straight to quarantine, and quarantine without a deadline becomes permanent deletion by neglect, minus the honesty of actually deleting it.
2. A flake-rate SLO, not a vibe
`(Opinion)` Most organizations running Playwright at scale have no numeric answer to "is our test suite getting more or less reliable this quarter." That should change. Define flake rate per test as reruns-until-pass divided by total runs over a rolling window (e.g. the last 200 executions), and track it per suite and per team, not just as one org-wide number that hides which team's suite is actually the problem.
A workable starting SLO: any test with a flake rate above a threshold (a reasonable starting point is 5 percent over a rolling 200-run window, `(Opinion)`, calibrate to your own suite's tolerance for false CI failures) automatically gets flagged for the fix-vs-quarantine-vs-delete review above, rather than waiting for someone to notice it in conversation. Playwright's built-in JSON reporter gives you the raw per-test pass/fail/retry data to build this; you don't need a separate paid flaky-test dashboard to get a first version of this working.
// playwright.config.ts
import { defineConfig } from '@playwright/test';
export default defineConfig({
reporter: [
['list'],
['json', { outputFile: 'test-results/results.json' }],
['html', { open: 'never' }],
],
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? '50%' : undefined,
use: {
trace: 'retain-on-failure',
video: 'retain-on-failure',
screenshot: 'only-on-failure',
},
});
The `json` reporter output includes each test's `status`, `retry` count, and duration. A small script parsing that artifact across recent CI runs (stored, for example, as a build artifact or pushed to a lightweight datastore) is enough to compute a rolling flake rate per test without adopting a new vendor product.
3. An ownership model for flaky tests at scale
This is the part almost never covered in Playwright-specific flaky-test content, because it's an organizational problem, not a Playwright API problem, but it's the actual blocker once a suite passes a few hundred tests written by multiple teams:
- Every spec file has a named owning team, tracked via a `CODEOWNERS` entry or an equivalent mapping, not "whoever wrote it originally and has since moved teams."
- A flaky test's quarantine ticket is assigned to the owning team automatically, with a default SLA (for example, 2 sprints) after which an unaddressed quarantined test is escalated to a platform/QA-infra team for a fix-or-delete decision, so it doesn't sit forever.
- Flake rate is a visible team-level metric, not just a suite-level one. A team whose specs make up a disproportionate share of the suite's flakiness should see that in the same dashboard they see build times and deploy frequency, so it's a normal engineering signal, not a special QA-only concern raised in a separate meeting.
- A platform or test-infrastructure team owns shared fixtures, seed data, and CI runner configuration, since a flake caused by shared database state or shard contention isn't fairly attributable to whichever team's test happened to fail first.
Without an ownership model, the fix-vs-quarantine-vs-delete framework above has no enforcement mechanism: decisions get made in theory but nobody is accountable for actually making them on a schedule.