Skip to main content
AgentsUse

Browser QA and regression checking

Browser QA Stack

Playwright runs the checks that must pass every time; the Playwright MCP server lets an agent investigate the failures.

3 verified components2 human checkpointsChecked 2026-10-07

The job

  • Re-run critical browser flows after changes and catch regressions
  • Investigate failures with artifacts instead of opinions

Expected output: Repeatable pass/fail checks with traces and screenshots, plus an agent-written failure report a human can confirm in minutes.

Components, and why each one is here

PlaywrightScripted cross-browser checks, screenshots, and traces.

A browser automation framework that drives Chromium, Firefox, and WebKit with a single API for tests, scripts, and AI agents.

Why this component: Deterministic scripts are the cheapest regression net there is.

Playwright MCP ServerAgent access to a live browser through accessibility snapshots.

Microsoft MCP server that automates a real browser through Playwright using accessibility snapshots instead of screenshots.

Why this component: Same project as the library, so debugging uses the same mental model as the tests.

mcp-server-gitRecent history and diffs for the repo under test.

Python MCP server that lets a model inspect and change a local Git repository through a fixed set of Git tools.

Why this component: Grounds the investigation in what actually changed.

Sequence

  1. Write the revenue-critical flows as independent Playwright tests with role and label locators.

  2. Run the suite on every meaningful change; save traces and screenshots on failure.

  3. The agent opens the failing page through Playwright MCP and states a specific claim about what broke.

  4. The agent reads recent commits through the Git MCP server and ties the failure to a change, or says it cannot.

  5. A person confirms bug versus test rot before anything is filed.

Prerequisites

  • Node.js for Playwright and the MCP server
  • The app running locally or on staging
  • An MCP-capable coding agent for investigation

Credentials

  • A test account for flows behind a login, stored outside the repo

Human checkpoints

Human checkpoints
  • Before an agent-filed issue reaches the team: agents blame the app for stale selectors with the same confidence as real bugs.
  • When product behavior changes on purpose: someone has to say so and update the test instead of letting retries hide it.

When it breaks

Selector rot after a redesign

Prefer role and label locators; let the agent propose updates a human approves.

Flaky timing failures

Read the Playwright trace first. Auto-waiting covers most of it; the rest is usually a real race worth seeing.

Agent cannot reproduce the failure

Treat it as unconfirmed. Keep artifacts, rerun the script, escalate with evidence.

Swapping components

Puppeteer with the Chrome DevTools MCP server is the Chrome-first variant of the same shape.

Cost shape and freshness

Playwright and both MCP servers are free and open source. Costs are CI minutes and the model tokens spent investigating failures.

Every component links to its verified profile, checked 2026-10-07. A stack is an editorial recipe, not a tested benchmark: run it on a small job first and keep the checkpoints in place. Spotted an error? Send a correction. Back to the stack index.