Local MCP server that exposes Playwright Java (com.microsoft.playwright) through Spring AI MCP
tools over stdio.
The project follows the same stack as the sibling MCP servers:
- Java + Gradle Kotlin DSL
- Spring Boot 4
- Spring AI MCP Server over stdio
- Playwright Java
- Java records for structured MCP responses
@McpToolservices with feature flags underplaywright-mcp.tools.*
This implementation is backed by Playwright Java and keeps its browser automation concepts visible where they matter, but tool descriptions are written for an agent that may only see the MCP contract:
- one default browser session
- element locator
- accessibility snapshot
- locator strategies:
id,role,text,label,placeholder,altText,title,testId,css,xpath
Host-side arbitrary code execution is intentionally not included. Page JavaScript evaluation, storage state, network routing, tracing, downloads, and PDF can be added as separate tool groups.
| Group | Flag | Tools |
|---|---|---|
| Page | PLAYWRIGHT_MCP_TOOLS_PAGE |
pageNavigate, pageWaitForLocator, pageSnapshot |
| Locator | PLAYWRIGHT_MCP_TOOLS_LOCATOR |
locatorClick, locatorFill, locatorType, locatorSelectOption, locatorPress, locatorCheck, locatorText, locatorInspect |
| Screenshot | PLAYWRIGHT_MCP_TOOLS_SCREENSHOT |
pageScreenshot |
All groups are enabled by default.
pageNavigate lazily creates the server-managed default Browser, BrowserContext, and Page, so
a client can start with:
{
"url": "https://example.com"
}for pageNavigate, then inspect:
{
"locator": {
"kind": "css",
"value": "body"
}
}with pageSnapshot.
Navigation accepts absolute HTTP and HTTPS URLs only. If a locator action opens exactly one popup, that popup becomes the current page automatically.
pageSnapshot is the single page-inspection tool. By default it returns just the structural overview
(snapshot): headings, landmarks, links, roles, and locator targets, with grids and tables collapsed
to a one-line summary. The collapse field reports whether collapsing was enabled, how many
data-like nodes were collapsed, and a capped list of collapsed node metadata with the kind, reason,
direct row count, and detected columns. Opt into extra sections with flags (each adds tokens and
time):
includeControls=trueadds thecontrolssection: interactive controls with the detail the structural snapshot omits — role, type, label/placeholder, enabled/disabled, checked, visible/hidden, and a tooltip. Typed field values are not returned.includeTooltips=true(withincludeControls) additionally hovers icon-only buttons that expose only an icon code (e.g.account_balance) and captures their real tooltip.includeGrids=trueadds thegridssection: the rows and columns inside tables and ARIA grids/treegrids (ag-Grid, Angular Material, plain HTML tables). Virtualized grids only keep on-screen rows in the DOM, sorenderedRowCountmay be smaller than the real total.collapseSnapshot=falsereturns the raw Playwright aria snapshot without collapsing data-like tables/grids. Use this only with a narrow root locator or when debugging collapse behavior, because responses can become large.
Bound large pages with the maxControls and maxRows caps, and narrow the read with a root locator
(e.g. nav, aside, main) instead of the default body.
After a menu click or filter change in a single-page app, use pageWaitForLocator (e.g. wait for
role=row to become visible, or for a loading spinner to become hidden) before snapshotting so
asynchronous content has settled.
pageWaitForLocator waits on the first matching element and returns the total count of matches.
found=true with count > 1 is not an error; it means the locator was broad enough to match several
elements. Scope the locator or use nth when uniqueness matters, and prefer a stable status label
for state checks before reading data with locatorText.
For a narrow read from a known element, use locatorText instead of a full pageSnapshot. It reads
the first matching element's compact text, input value, selected option text, aria-label, or
title, and returns the total match count so ambiguous locators are visible. This is useful for
large table rows and status labels.
Use locatorInspect before persisting a locator into generated test code. It reports the total match
count and metadata for the first match, including the raw HTML id, label, accessible name, tag, and
state. When pageSnapshot(includeControls=true) or locatorInspect returns an application-defined
stable HTML id, prefer it as kind=id, then use role/name or label; keep css/xpath as
fallback. Reject ambiguous locators instead of guessing with nth.
Use locatorType for autosuggest and masked inputs that depend on per-key events; ordinary fields
should still use locatorFill. Use locatorSelectOption with exactly one of option value, visible
label, or zero-based index.
pageScreenshot captures the current browser page as a PNG file under
${playwright-mcp.data-dir}/screenshots. It returns the absolute and relative file paths plus page
metadata; it does not return image bytes in the MCP response. Use it for documentation, bug reports,
and step-by-step evidence.
The locator object is the main input shape for element actions and reads:
{
"kind": "role",
"role": "button",
"name": "Save",
"exact": true,
"hasText": null,
"nth": 0,
"visible": true,
"frameSelector": null
}Supported kind values:
id: exact HTMLidattribute invaluecss: CSS selector invaluexpath: XPath selector invaluerole: accessible role inrole, accessible name innametext: visible text invaluelabel: form field label invalueplaceholder: placeholder text invaluealtText: image/media alt text invaluetitle: title attribute text invaluetestId: Playwright test id invalue(data-testidby default)
hasText, visible, and nth narrow the match. frameSelector searches inside an iframe selected
by CSS.
.\gradlew.bat buildFast unit tests:
.\gradlew.bat testInstall Chromium browser binaries for Playwright Java:
.\gradlew.bat installPlaywrightBrowsersIntegration and smoke checks:
.\gradlew.bat integrationTest
.\gradlew.bat smokeTestintegrationTest runs against an in-process local HTTP site and a real headless Chromium. It uses
PLAYWRIGHT_MCP_EXECUTABLE, local Chrome/Edge, or an already-installed Playwright Chromium cache.
If none is available, the test fails with an installation/configuration hint; browser-backed
verification is never silently skipped.
Manual live application exploration:
$env:LIVE_APPLICATION_URL="<application-url>"
$env:LIVE_APPLICATION_USERNAME="<username>"
$env:LIVE_APPLICATION_PASSWORD="<password>"
.\gradlew.bat liveApplicationTestliveApplicationTest is excluded from test, build, integrationTest, and smokeTest. It is a
manual live smoke/exploration check: it opens the configured application URL, logs in with the
configured credentials, opens discovered menu entries, captures compact accessibility snapshots,
and writes build/reports/live-application/report.md with page and form descriptions. Required env
vars are LIVE_APPLICATION_URL, LIVE_APPLICATION_USERNAME, and LIVE_APPLICATION_PASSWORD; the
test is skipped when any of them is missing. Optional env vars: LIVE_APPLICATION_HEADLESS (default
true) and LIVE_APPLICATION_MAX_MENU_ITEMS (default 30).
The server is stdio-only.
| Variable | Default |
|---|---|
PLAYWRIGHT_MCP_DATA_DIR |
${user.home}/.playwright-mcp-server |
PLAYWRIGHT_MCP_BROWSER |
chromium |
PLAYWRIGHT_MCP_CHANNEL |
empty |
PLAYWRIGHT_MCP_EXECUTABLE |
empty |
PLAYWRIGHT_MCP_HEADLESS |
true |
PLAYWRIGHT_MCP_SLOW_MO_MS |
0 |
PLAYWRIGHT_MCP_TOOLS_SCREENSHOT |
true |
PLAYWRIGHT_MCP_VIEWPORT_WIDTH |
1440 |
PLAYWRIGHT_MCP_VIEWPORT_HEIGHT |
1000 |
PLAYWRIGHT_MCP_DEFAULT_TIMEOUT_MS |
30000 |
Logs go to stderr and ${playwright-mcp.data-dir}/logs/playwright-mcp-server.log. stdout is kept
for MCP JSON-RPC.