Browser QA and regression checking
Browser QA Stack
Playwright runs the checks that must pass every time; the Playwright MCP server lets an agent investigate the failures.
The job
- Re-run critical browser flows after changes and catch regressions
- Investigate failures with artifacts instead of opinions
Expected output: Repeatable pass/fail checks with traces and screenshots, plus an agent-written failure report a human can confirm in minutes.
Components, and why each one is here
A browser automation framework that drives Chromium, Firefox, and WebKit with a single API for tests, scripts, and AI agents.
Why this component: Deterministic scripts are the cheapest regression net there is.
Microsoft MCP server that automates a real browser through Playwright using accessibility snapshots instead of screenshots.
Why this component: Same project as the library, so debugging uses the same mental model as the tests.
Python MCP server that lets a model inspect and change a local Git repository through a fixed set of Git tools.
Why this component: Grounds the investigation in what actually changed.
Sequence
Write the revenue-critical flows as independent Playwright tests with role and label locators.
Run the suite on every meaningful change; save traces and screenshots on failure.
The agent opens the failing page through Playwright MCP and states a specific claim about what broke.
The agent reads recent commits through the Git MCP server and ties the failure to a change, or says it cannot.
A person confirms bug versus test rot before anything is filed.
Prerequisites
- Node.js for Playwright and the MCP server
- The app running locally or on staging
- An MCP-capable coding agent for investigation
Credentials
- A test account for flows behind a login, stored outside the repo
Human checkpoints
- Before an agent-filed issue reaches the team: agents blame the app for stale selectors with the same confidence as real bugs.
- When product behavior changes on purpose: someone has to say so and update the test instead of letting retries hide it.
When it breaks
Selector rot after a redesign
Prefer role and label locators; let the agent propose updates a human approves.
Flaky timing failures
Read the Playwright trace first. Auto-waiting covers most of it; the rest is usually a real race worth seeing.
Agent cannot reproduce the failure
Treat it as unconfirmed. Keep artifacts, rerun the script, escalate with evidence.
Swapping components
Puppeteer with the Chrome DevTools MCP server is the Chrome-first variant of the same shape.
Cost shape and freshness
Playwright and both MCP servers are free and open source. Costs are CI minutes and the model tokens spent investigating failures.
Every component links to its verified profile, checked 2026-10-07. A stack is an editorial recipe, not a tested benchmark: run it on a small job first and keep the checkpoints in place. Spotted an error? Send a correction. Back to the stack index.