Delivered · Study13views
WindTunnel study - ab
webmcp vs playwrite · 50 tasks · 2 harnesses · 2 models
WebMCP delivers substantial efficiency gains—nearly half the cost and tokens, with a third faster execution.
Correctness did not separate the arms beyond the decision margin, and the treatment got there with 49% fewer tokens and 48% lower cost.
Abstract
WebMCP outperforms Playwright across all efficiency metrics in this 50-task benchmark spanning 2 harness/model cells. Cost came in at $0.06 (mean $0.10) versus $0.12 (mean $0.27), a -47.8% improvement. Token consumption was 282.5k (mean 436.3k) against 550k (mean 1.2M), representing -48.6% reduction. Duration improved -35.3%, with a median of 3m 55s (mean 4m 36s) compared to 6m 04s (mean 6m 28s). These gains held consistently across cells, though harness and model effects cannot be isolated separately due to the confounded design.
The result
playwritewebmcp
Best arm overall: webmcp (88.9 vs 79.1 of 100)
Best setup: claude-code / claude-haiku-4-5-20251001 · webmcp (98.5 of 100)
- 1claude-code / claude-haiku-4-5-20251001webmcp98.5best
- 2claude-code / claude-haiku-4-5-20251001playwrite88.2
- 3codex / gpt-5.4-miniwebmcp86.8
- 4codex / gpt-5.4-miniplaywrite76.0
Overall score per arm, 0–100 points (not a pass rate): 75% benchmark pass rate (the benchmark's own graders carry the outcome share — no judge grades or evals on this board) + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.
Quality × efficiency clusters — normalized per task
Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (85% the benchmark's own pass/fail + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.
6 runs not plotted (missing a grade or token count — infra failures included).
Every metric, per harness × model
Benchmark pass rate
- claude-code / claude-haiku-4-5-2025100194% → 98%4.0 pp better
- codex / gpt-5.4-mini92% → 96%4.0 pp better
Task completion
- claude-code / claude-haiku-4-5-20251001100% → 100%even
- codex / gpt-5.4-mini100% → 100%even
Cost per run
- claude-code / claude-haiku-4-5-20251001$0.05 → $0.0350% better
- codex / gpt-5.4-mini$0.19 → $0.0953% better
Tokens per run
- claude-code / claude-haiku-4-5-20251001255.4k → 176.6k31% better
- codex / gpt-5.4-mini726.2k → 295.1k59% better
Duration per run
- claude-code / claude-haiku-4-5-202510013m 45s → 3m 30s7% better
- codex / gpt-5.4-mini7m 43s → 3m 59s48% better
playwrite
webmcp
The publisher has not described this setup — what each group's tools are, who built them, and how the tasks were chosen.
Distributions
Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.
Cost per run
Token composition — average per run
Run outcomes
How each run ended — whether the agent finished, not whether its answer passed. Whether answers passed is in the results above.
Statistics
| Metric | Arm | n | Mean | Median (pooled) | Min | Max | Std dev |
|---|---|---|---|---|---|---|---|
| Cost | base | 98/100 | $0.27 | $0.08 | $0.01 | $7.90 | $0.88 |
| treat | 96/100 | $0.10 | $0.03 | $0.01 | $0.61 | $0.14 | |
| Duration | base | 100/100 | 6m 28s | 6m 44s | 1m 03s | 22m 09s | 3m 43s |
| treat | 100/100 | 4m 36s | 3m 49s | 1m 09s | 10m 40s | 2m 36s | |
| Total tokens | base | 98/100 | 1.2M | 360.3k | 34.4k | 34.2M | 3.9M |
| treat | 96/100 | 436.3k | 178.1k | 26.1k | 2.6M | 585.8k | |
| Output tokens | base | 98/100 | 5.4k | 2.5k | 224 | 114.6k | 13.3k |
| treat | 96/100 | 2.3k | 1.1k | 339 | 10.3k | 2.7k | |
| Cache-read tokens | base | 98/100 | 1.2M | 336.1k | 21.8k | 32.3M | 3.7M |
| treat | 96/100 | 404k | 171.9k | 10.6k | 2.5M | 548.1k | |
| Cache-write tokens | base | 98/100 | 9k | 2.1k | 0 | 59.3k | 13.9k |
| treat | 96/100 | 3.7k | 3.4k | 0 | 39.7k | 6.8k | |
| Turns | base | 100/100 | 16.5 | 8 | 2 | 98 | 19.3 |
| treat | 100/100 | 9.9 | 8 | 3 | 81 | 9 |
n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.
What each setup did
Derived from each run's recorded tool calls — not from a model's description of the run — and aggregated per setup, so a behavior seen across several runs is stated once with its rate. Open a finding to see the runs behind it, each linked to its journey at the step where it happened.
codex · openai · gpt-5.4-mini · playwrite3 findings
- 5 of 50 runs · 5 completed
- 5 of 50 runs · 5 completed
- 5 of 50 runs · 5 completed
Task text withheld
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Tasks
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Runs
Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.
Every run, filterable… of 200Show runsHide runs
| Harness | Model | Task | Arm | Completed | Pass | Tokens | Cost | Duration |
|---|
Loading runs…
Methodology
What each metric means
- Completed
- The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
- Denominator: Terminal runs, excluding those killed by our own infrastructure.
- Quality index
- A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
- Denominator: Graded runs — runs the study's grader scored.
- Infra-excluded
- A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
- Denominator: Reported as a count beside every affected panel.
- A/B arms
- Every task runs twice per harness × model cell — once as playwrite, once as webmcp — on the same prompt, same model, cold start for both arms. Both arms are real configurations.
- One run per task
- Every task ran once per cell and arm, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
- Sample size
- 2 cells × 50 tasks × 2 arms = 200 runs.