End-to-End Testing with Playwright: Lessons from Owning a Test Suite
What I've learned building Playwright E2E coverage for critical user flows — resilient locators, auth setup, test data, flaky tests and running it all in CI.

At my current role I own end-to-end testing: building automated Playwright coverage for the user flows that must never break — sign-in, onboarding, checkout-style flows, and the admin actions that touch money or data.
Coming from mobile development, E2E testing on the web has been refreshing. Playwright is fast, reliable and pleasant to write. But a test suite is only valuable if people trust it, and trust is easy to lose. These are the lessons that made the biggest difference.
Test the flows that cost money when they break
It’s tempting to aim for “test everything”. Don’t. E2E tests are slower and more expensive to maintain than unit tests, so spend them where they pay off.
I start by listing the flows where a bug would cause real damage:
- A user can’t sign up or sign in.
- A payment or subscription fails.
- Data is saved incorrectly or lost.
- A permission check lets someone see what they shouldn’t.
Those get E2E tests first. Edge cases in form validation are usually better covered by unit or component tests.
Use locators that match what users see
The single biggest cause of brittle tests is selecting elements by CSS classes or DOM structure. A designer renames a class and twenty tests fail, even though nothing is broken for users.
Playwright’s role-based locators select elements the way a user (or a screen reader) would find them:
import { test, expect } from '@playwright/test';
test('user can create a customer', async ({ page }) => {
await page.goto('/customers');
await page.getByRole('button', { name: 'New customer' }).click();
await page.getByLabel('Full name').fill('Jane Doe');
await page.getByLabel('Email').fill('jane@example.com');
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByRole('row', { name: /Jane Doe/ })).toBeVisible();
});
A nice side effect: if a button has no accessible name, getByRole can’t find it, and the test pushes the team to fix an accessibility bug. When there’s genuinely no good user-facing handle, I add a data-testid rather than depending on layout.
Sign in once, not in every test
Logging in through the UI at the start of every test is slow and makes every test depend on the login page. Playwright’s setup projects let you sign in once and reuse the session:
// auth.setup.ts
import { test as setup, expect } from '@playwright/test';
setup('authenticate', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill(process.env.E2E_USER!);
await page.getByLabel('Password').fill(process.env.E2E_PASSWORD!);
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page).toHaveURL(/dashboard/);
await page.context().storageState({ path: 'playwright/.auth/user.json' });
});
// playwright.config.ts
export default defineConfig({
projects: [
{ name: 'setup', testMatch: /.*\.setup\.ts/ },
{
name: 'chromium',
use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
dependencies: ['setup'],
},
],
});
The login flow still gets its own dedicated test. Everything else starts already signed in. Make sure playwright/.auth is in .gitignore — it contains live session cookies.
Own your test data
Tests that depend on data someone created by hand in a shared environment will eventually fail for reasons that have nothing to do with the code.
My rules:
- Each test creates what it needs, ideally through the API rather than the UI — it’s much faster. Playwright’s
requestfixture makes this easy. - Use unique values (a timestamp or random suffix in names and emails) so parallel tests and repeated runs don’t collide.
- Clean up in an
afterEachor with a scheduled job that removes test-tagged records.
test('invoice shows the correct total', async ({ page, request }) => {
const res = await request.post('/api/test/invoices', {
data: { customer: `e2e-${Date.now()}`, lines: [{ amount: 1500 }, { amount: 2500 }] },
});
const { id } = await res.json();
await page.goto(`/invoices/${id}`);
await expect(page.getByTestId('invoice-total')).toHaveText('$40.00');
});
Never use fixed waits
await page.waitForTimeout(3000) is the root of most flaky tests. Three seconds is too long on a fast machine and too short on a busy CI runner.
Playwright’s expect assertions auto-retry until the condition is met or the timeout expires, and actions like click() wait for the element to be visible and enabled. Lean on that:
// ❌ flaky
await page.waitForTimeout(3000);
expect(await page.getByText('Saved').isVisible()).toBe(true);
// ✅ waits exactly as long as needed
await expect(page.getByText('Saved')).toBeVisible();
If you need to wait for a network call, wait for that response, not for time to pass:
const saved = page.waitForResponse((r) => r.url().includes('/api/customers') && r.ok());
await page.getByRole('button', { name: 'Save' }).click();
await saved;
Treat flaky tests as bugs
A test that fails 1 time in 20 is worse than no test, because it teaches the team to ignore red builds. When a test flakes, I do one of two things the same day: fix it, or quarantine it (skip it with a linked ticket). I don’t let it sit.
Playwright’s trace viewer makes diagnosis much easier. I enable traces on the first retry in CI:
use: { trace: 'on-first-retry', screenshot: 'only-on-failure' },
The trace captures every action, network request, console log and a DOM snapshot at each step. Most “it passes locally” mysteries take minutes to solve with it.
Common root causes I’ve found: tests sharing data, animations that shift elements mid-click, missing awaits, and assertions that check a value before a background refresh finishes.
Run it in CI on every pull request
Tests that only run on someone’s laptop don’t protect anything. The suite runs in our Azure DevOps pipeline on every pull request, with the HTML report and traces published as build artifacts so a failure can be diagnosed without reproducing it locally. (I wrote about pipeline setup in CI/CD for Flutter with Azure DevOps — the same ideas apply.)
A few settings that help in CI:
retries: 2in CI only, so a single infrastructure hiccup doesn’t block a merge — while the trace still shows you something was flaky.workerstuned to the runner’s CPU count.- Shard long suites across multiple agents with
--shard=1/3,--shard=2/3, and so on.
Takeaways
- Cover the flows where a bug costs money or trust, first.
- Select elements by role, label and text — the way users find them.
- Authenticate once with a setup project; create test data via the API.
- Replace every fixed wait with an auto-retrying assertion.
- Fix or quarantine flaky tests immediately, and use traces to find out why.
A good E2E suite is one the team trusts enough to block a release on. That trust comes from reliability more than coverage.
Comments
Questions, corrections or your own experience — leave a comment below (GitHub sign-in).