Playwright MCP and Regression Testing: Create with the Agent, Replay Without an LLM

Playwright MCP and Regression Testing: Create with the Agent, Replay Without an LLM

Playwright MCP: the Model Context Protocol server published by Microsoft for Playwright (official repository). It lets an AI agent drive a browser by reading the accessibility tree of each page, without screenshots or a vision model. The agent can open a URL, click, fill in a form and read what is displayed.

More and more QA teams hand over their regression testing to an agent such as Claude, connected to Playwright MCP. For exploring an application or reproducing a bug, the tool is remarkable. For a test replayed on every release, it lacks one property: repeating exactly the same journey, because the agent decides every action again on each run. Delta-QA MCP splits the roles differently. Your agent creates the scenario in plain language, then Delta-QA replays it identically, without an LLM, and compares every release to the reference.

The essentials on one page (definition, no-code automation, tool selection grid): Automated regression testing, without code.

Your AI agent already knows how to browse your site. Connect it to Delta-QA: it writes the scenario, and Delta-QA replays it on every release and photographs every step. Try Delta-QA for free →


Why Playwright MCP appeals to so many QA teams

Playwright MCP turns a plain-language request into real actions in a browser. A tester writes "check that the contact form rejects an invalid email," and the agent opens the page, types, submits, then reports back. No script to write, no selector to hunt for: that explains its fast adoption.

The numbers confirm it. The @playwright/mcp package went from 2.2 million npm downloads in September 2025 (npm) to 29.3 million in September 2026 (npm), and its GitHub repository is approaching 38,000 stars. Downloads also include automated installs, so the metric is imperfect. The trend, however, is clear: 13 times as many downloads in one year.

Its strengths are real. Its tools act on the accessibility tree of the page, which Microsoft presents as deterministic tool application, avoiding the ambiguity of screenshot-based approaches. Microsoft recommends it for exploratory automation and for tasks where the agent must keep the browser context. Reproducing a reported bug, checking a page after a fix, discovering a feature: on these one-off tasks, it saves a great deal of time.

Playwright MCP is therefore an excellent exploration tool. What remains to be seen is whether it can become your safety net on every release. For how the protocol itself works, see Playwright and MCP (Model Context Protocol).

A test replayed by an agent is not a regression test

A regression test only has value if it replays the same journey, with the same checks, on every release. That is what lets you attribute a difference to the release, not to the test. An agent driven by an LLM decides every action again on each run: the same prompt can produce a different path, and therefore a different verdict.

This is not a flaw in Playwright; it is the nature of LLMs. In September 2025, the Thinking Machines lab submitted the same request to the Qwen3-235B model 1,000 times, at temperature zero, the setting supposed to be the most stable. It got 80 different responses (Thinking Machines Lab). The Playwright MCP tools run deterministically; the decision to call them, in which order and on which element, does not.

For a QA team, this has three concrete consequences:

  1. A failure does not tell you what changed. When the replay fails, nothing shows whether the site regressed or the agent took a different path.
  2. A success does not prove the same journey was followed. The agent can reach the final page by a detour and conclude that everything works.
  3. Every run costs time and tokens. The agent reasons again at every step, and Microsoft points out that the MCP loads large tool schemas and accessibility trees into the model's context (Playwright MCP README).

Playwright has its own answer: freeze the journey as code. Its test agents explore the application, write up a plan, then turn it into Playwright Test files that a third agent repairs when they break. That is the right path for a development team that versions and maintains its test code. For a QA team that does not code, it means inheriting TypeScript files, selectors and assertions to maintain.

Delta-QA MCP: The agent creates the scenario, Delta-QA replays it

Delta-QA's MCP server keeps what the agent does best, understanding a request in plain language, and removes what it does poorly, replaying identically. Your agent creates the scenario and writes its steps. Delta-QA replays them in its own browser, on your real site, photographs every step and only validates what actually worked.

The server exposes eight tools, which the agent discovers on its own. Six of them are enough to create a test, in this order:

  1. list_workspaces: the agent chooses the workspace where the test is created.
  2. create_scenario: it creates the scenario from a name and a start URL; step 1, opening the page, is set by Delta-QA.
  3. set_scenario_steps: it writes the steps of the journey, out of fourteen gestures (click, type, select, key, scroll, navigate…), each one targeting an element of the page.
  4. capture_scenario: Delta-QA replays the scenario in its browser, on the real site, and takes one photo per step.
  5. get_capture_result: the agent reads the verdict. A failure names the step to fix; the agent fixes it and runs it again.
  6. activate_scenario: after a successful replay, the agent validates the scenario and the capture becomes the reference.

Nothing is validated on trust. An agent can pick the wrong selector or claim that a journey works: until Delta-QA has successfully replayed every step, the scenario stays a draft. The tools, roles and keys are detailed on the Delta-QA MCP server page.

Playwright MCP Delta-QA MCP
Who runs the journey The agent, action by action, on every run Delta-QA, from the steps the agent wrote once
Role of the LLM at replay It decides every action None: the LLM only takes part at creation
What remains after the session The conversation history, or test files to maintain A scenario in Delta-QA, readable and editable without code
Proof of the result The agent's report and its screenshots One photo per step and a verdict that names the failing step
Next release A new agent run, or code to update Identical replay, compared with the reference
Main strength Exploration, bug reproduction, one-off checks Regression tests replayed on every release, without code

Let your agent write your scenarios, keep an identical replay. Delta-QA replays every step on your real site and only validates what actually worked. Try Delta-QA for free →

Example: Create the login journey test with Claude Code

Connecting Claude Code to Delta-QA takes an API key and a single command. Then you describe the journey as you would to a colleague. The agent writes the scenario, Delta-QA replays it and keeps the proof, step by step.

  1. Create your free account and your workspace in Delta-QA.
  2. Create an API key in Settings, Workspace tab, API keys section. Give it a name, the "Can edit" role for an agent that writes scenarios, and an expiry of 30 days, 90 days or 1 year. The key is shown only once, with the ready-made command.
  3. Paste the command into a terminal, replacing dqa_… with your key:
claude mcp add --transport http delta-qa https://api.delta-qa.com/api/mcp --header "Authorization: Bearer dqa_…"
  1. Describe the journey in Claude Code, in plain language. For example: "Create a regression test of the login journey on https://staging.example.com: enter the test account's email, the password, submit, then check that the dashboard is displayed."
  2. Let the agent work. It creates the scenario, writes the steps and asks Delta-QA to replay them. If a step fails, for example a button not found at step 4, the verdict names it: the agent fixes it and runs it again.
  3. Validate. When the replay succeeds, the agent activates the scenario: its capture becomes the reference.

The password, marked as a sensitive value, is encrypted as soon as it arrives, typed at replay and veiled on the photos: the agent itself can no longer read it. The API key only acts within its workspace, with the rights of its role, and can be revoked immediately from the same screen.

Then QA takes over in Delta-QA

Once validated, the scenario created by the agent becomes a Delta-QA scenario like any other. It appears in your workspace, readable and editable without code, as if it had been recorded by browsing. The agent's role stops there: replaying, comparing and deciding go back to the QA team, with all of Delta-QA's tools.

On every release:

  • Replay in one click: the scenario runs on its own, on Chrome, Firefox and WebKit, at the screen widths you choose.
  • Compare: the report shows the reference and the current version side by side, with differences highlighted and classified (critical, warning, minor). Delta-QA's deterministic AI, not an LLM, always gives the same verdict for the same difference.
  • Decide: you approve the change as the new reference, or open the bug from the report.

The report contains what a developer needs to fix the problem: the step concerned, the reference capture, the current capture and the areas that differ. These elements can be passed on as they are to the developer, or to their coding agent, who starts from a precise difference rather than a rough description. After the fix, a new replay of the same scenario confirms that the regression is gone.

Your agent writes the scenario, Delta-QA keeps the proof. Create your API key and have Claude Code write your first regression test. Try Delta-QA for free →

FAQ

What is Playwright MCP?

Playwright MCP is the Model Context Protocol server published by Microsoft for Playwright. It lets an AI agent such as Claude drive a real browser by reading the accessibility tree of the pages: open a URL, click, type, read content. The agent acts from plain-language instructions, without a script.

Can you do regression testing with Playwright MCP?

For a one-off check, yes. For a test replayed on every release, the agent decides every action again on each run, so two runs can follow two different paths. The journey then has to be frozen: as Playwright Test code that the team maintains, or in a tool that replays it identically, such as Delta-QA.

Do you need to code to use the Delta-QA MCP server?

No. You describe the journey in plain language, the agent writes the steps and Delta-QA replays them. The scenario then appears in Delta-QA, readable and editable without code, like a scenario recorded by browsing.

Does the LLM take part when Delta-QA replays the test?

No. The agent only takes part when the scenario is created. Replays are run by Delta-QA's browser and compared with the reference by its deterministic AI, not an LLM: the same input always gives the same verdict.

Which agents can connect to Delta-QA?

Claude Code connects to it with a single command. Any MCP client that accepts a remote "Streamable HTTP" server, with an authentication header, can connect too.

Is the MCP server included in the free account?

Yes. API keys and the MCP server exist on every plan, free account included. What the agent triggers counts towards your plan's quota, as if you had started it yourself.