Playwright's MCP, initially released in March 2025, integrates natively with Claude Code, Cursor, OpenClaw, and Codex to build exploratory automation, self-healing tests, or long-running autonomous workflows.
However, its architecture, built around accessibility tree snapshots, is token-consuming, especially in reasoning loops where the context window rapidly increases, turn after turn. As the context of each visited page stays in the context as the agent progresses, the cost grows, and the agent quickly loses context of the active work after a few navigation steps.
None of this makes Playwright a bad tool. The fact is that Playwright was initially designed for browser testing, not for AI Agents.
How Playwright MCP represents a page to an agent
While Playwright exposes multiple API to inspect (locator(), screenshot()) and interact (.click(), .fill()) with a webpage, its MCP versions primarily rely on accessibility tree snapshots: each button, link, and input with its role, name, and a reference ID the agent uses to act on it.
This snapshot-based architecture is, in theory, a good fit for agents, as text is cheaper than images and refs are more reliable than complex selectors or pixel coordinates.
Here's how your Coding Tool (Codex, Claude Code) or custom agent interacts with Playwright MCP:

Steps 3 and 5 are the most token-consuming; let's dive into what makes them that greedy.
Where the tokens go
Playwright MCP's snapshots architecture, while more efficient than screenshots, produces compounding context window growth and token consumption for 4 main reasons:
1. Snapshot size: A simple login page produces a small snapshot while a dashboard, product listing, or docs page with navigation, footers, and hundreds of links produces a much larger one.
2. Snapshots after every action: Returning the updated page state after each click is what lets the model confirm its action worked. It also means a 15-step task can produce 15 snapshots. By default, Playwright MCP's snapshot mode for responses is set to full (README).
3. Accumulation: Old snapshots don't disappear, they stay in the conversation history, so by step 10 the model is carrying the state of pages it left long ago. This is the part developers notice: Reddit users described context being gone "after 2–3 navigations."
4. Tool schemas: Playwright MCP exposes 25 tools, accounting for 17,857 tool-definition chars added to the agent's context before any instruction is given.
Benchmark: tokens per page and per task
Let's now talk numbers and see how many tokens Playwright MCP consumes on tasks performed on real websites.
We evaluated Playwright MCP's token consumption on 8 websites composed of 4 public website and 4 local ones simulating more complex scenarios.
Four on public sites:
- saucedemo.com: log in, add two items, read the checkout total.
- books.toscrape.com: the three cheapest books in a category that spans two pages. A four-hop Wikipedia navigation, following links only. One fact buried in a 670 KB Node.js docs page.
Four on local test pages, built so they behave the same on every run: A form inside a cross-origin iframe. A button inside a closed Shadow DOM. A clickable icon with no role or accessible name. A 12-page form flow, used to measure how context grows page by page.
We benchmarked the token consumption of Playwright MCP 0.0.82 with its default config, Playwright MCP tuned ( — snapshot-mode none — codegen none, and the agent is told to use browser_find and browser_snapshot with depth), and Stagehand v4, through its Claude Code facade (run / snapshot / screenshot).
Each task ran 3 times per setup, in one randomized order, using Claude Sonnet 5 (*claude-sonnet-5 with thinking disabled*) and driven by the Claude Agent SDK 0.3.282.

Interestingly, the "optimized" Playwright MCP configuration ended up consuming more tokens but was better optimized for prompt caching, leading to a similar cost per task.
Per step, Stagehand v4 sent smaller requests: a median of 6.9k input tokens per model call, against 12.9k–14.1k for Playwright MCP. On the 12-page flow, the picture flips per page. Stagehand's context grew about 1.8k tokens per step, because it re-reads the page after each form submission; MCP grew about 0.55k. But Stagehand batches several actions into one call, so it finished all 12 pages in 27 steps. Both MCP setups hit the 40-step limit in all three runs without finishing.
More evals against multiple models and harnesses are available at https://stagehand.dev/evals.
Why agents lose context
An agent's context window contains all the history of the a ongoing conversation, which means: every token of reasoning, input and outputs tokens as well as tools schemas and their inputs and outputs. A context window naturally grows as the agent performs steps towards its goal.
Pruning the context, also called compaction, comes at a cost, invalidating the prompt cache offered by most LLM providers and therefore, drastically increasing the cost of new agent turns.
An heavy context comes with 2 major drawbacks: of course, cost (but mostly contained through prompt caching) and primarily accuracy. LLMs tend to gradually perform worse as the context grow, leading to hallucinations linked to the inability of the LLM to recall prior instructions.
Why agents miss elements on the page
Accessibility elements are handwritten by humans and sometimes, completely missing or only partial. Therefore, solely relying on these ARIA attributes to construct your agent's vision of a webpage leads to missed elements, like a country road without signs.
While Playwright finds a way to counter this, 2 main technical challenges increases the blindness of agents on the web:
- Cross-origin iframes, common in payment forms, embedded widgets, and auth flows.
- Closed Shadow DOM. Playwright handles open shadow roots, but closed roots are designed to be inaccessible from outside the component.
How to reduce token usage with Playwright today
Before switching tools, these options help:
Use Playwright CLI for coding agents. Microsoft now recommends it: the Playwright MCP README says coding agents "might benefit from using the CLI+SKILLS instead" (README,docs). The CLI keeps page state out of the model's context unless the agent asks for it. Install with npm install -g @playwright/cli@latest.
Tune Playwright MCP. Recent versions include options that cut context:
- — snapshot-mode none stops snapshots from being returned automatically after actions.
- browser_snapshot accepts a filename to save to disk instead of returning inline, and a depth to limit the tree.
- browser_find searches the snapshot for specific text and returns only matching nodes.
- — mobile emulates a mobile device, which usually means lighter pages.
- — caps controls which optional tool groups load, so leave off the ones you don't use.
Microsoft is explicit that MCP still has a place: for long-running autonomous workflows where maintaining continuous browser context outweighs token cost concerns. That's exactly the workload where the cost matters most.
What Stagehand does differently
Stagehand is built for agents while keeping Playwright's interfaces (goto(), click(), locator, screenshot(), and so on).
Stagehand is both optimized for speed and token-efficient execution:
Faster execution is achieved by running Stagehand's core architecture within the browser, as an automatically loaded extension. By doing so, each action gets computer and transmitted directly to the browser, without network in between. This brings up to a 50% faster execution in production deployments.
Token-efficient execution is achieved thanks to core features. First, Stagehand translate each pages to an hybrid accessibility snapshots tree that gives the model only meaningful elements and drop the unnecessary ones. Then, it offers 3 high-level APIs (observe(), act() and extract()) to enable the agent to prompt its way instead of navigating the whole tree. Finally, Stagehand's hybrid accessibility snapshots supports nested and cross-origin iframes as well as closed Shadow DOM.

Stagehand's architecture delivers 50% faster execution and uses 80% fewer tokens than Playwright MCP.
More information on https://docs.stagehand.dev.
When to keep using Playwright
Use Playwright when there's no model in the loop: end-to-end tests, deterministic scripts, CI suites. Its runner, trace viewer, and parallelism are hard to beat, and nothing in this post argues otherwise. Use Playwright CLI when a coding agent needs a browser occasionally alongside a codebase. Reach for Stagehand when the agent itself runs the browser across many steps, on sites you don't control.