Measured · Study5views

WindTunnel study - test

WindTunnel · 50 tasks · 2 harnesses · 2 models

WindTunnel is the winner on this benchmark.

Ranked shootout, no baseline arm: cells ranked by the composite score — the benchmark's own graders, then efficiency — claude-code / claude-haiku-4-5-20251001 leads.

Abstract

WindTunnel delivered consistent efficiency gains across all 50 tasks. The treatment achieved a median cost of $0.07 (mean $0.17), median token usage of 275.6k (mean 713.5k), and median duration of 5m 27s (mean 4m 55s). Performance was stable across the 2 harness/model cells, with claude-code / claude-haiku-4-5-20251001 leading the ranking. The single-arm shootout design measures treatment performance directly without a baseline arm for comparison.

The result

Color by

claude-codecodex

Best setup: claude-code / claude-haiku-4-5-20251001 (98.1 of 100)

  1. 1claude-code / claude-haiku-4-5-2025100198.1best
  2. 2codex / gpt-5.4-mini85.1

Overall score per cell, 0–100 points (not a pass rate): 75% benchmark pass rate (the benchmark's own graders carry the outcome share — no judge grades or evals on this board) + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.

Leaderboard

Cells ranked by the composite score — outcome (the lenses that graded) × efficiency — absolute values, no baseline arm. Tokens, cost and duration are per-cell medians, the efficiency the score folds.

#HarnessModelScorePass rateQualityEvalsGoalTokensCostDurationRuns
1claude-codeclaude-haiku-4-5-2025100198.198%−−88%176.7k$0.035m 02s50/50
2codexgpt-5.4-mini85.194%−−94%350.8k$0.104m 58s50/50

Pass rate

94.9% (95% CI 89.7–98.7)over 76 graded runs
CellPass rateTasksRunsRan asMatcher
claude-code/claude-haiku-4-5-2025100197.3% (95% CI 91.9–100)3737 (+13 unmeasured)with my setuptext_predicate
codex/gpt-5.4-mini92.3% (95% CI 84.6–100)3939 (+11 unmeasured)with my setuptext_predicate

Each run's answer was extracted from its response and matched against the benchmark's own key — no model in the loop. Two-stage: the mean over repeats per task, then the mean over tasks, with the interval taken over tasks.

Quality × efficiency clusters — normalized per task

Every graded run, standardized WITHIN its task so difficulty cancels out: → right = fewer tokens than the field on the same task, ↑ up = higher quality index (85% the benchmark's own pass/fail + 15% outcome) than the field. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.

Label
Task
Harness
Model
89 of 89 runs match
Show
Color by

claude-codecodex

better · cheaperworse · pricier
← pricier than the fieldefficiency (σ, per task)cheaper than the field →

11 runs not plotted (missing a grade or token count — infra failures included).

Every metric, per harness × model

Benchmark pass rate

  1. claude-code / claude-haiku-4-5-2025100198%
  2. codex / gpt-5.4-mini94%

Task completion

  1. codex / gpt-5.4-mini94%
  2. claude-code / claude-haiku-4-5-2025100188%

Cost per run

  1. claude-code / claude-haiku-4-5-20251001$0.03
  2. codex / gpt-5.4-mini$0.10

Tokens per run

  1. claude-code / claude-haiku-4-5-20251001176.7k
  2. codex / gpt-5.4-mini350.8k

Duration per run

  1. codex / gpt-5.4-mini4m 58s
  2. claude-code / claude-haiku-4-5-202510015m 02s

Your setup

webmcp

The publisher has not described this setup — what each group's tools are, who built them, and how the tasks were chosen.

Distributions

Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.

Cost per run

Your setup
code/claude-haiku-4-5-20251001
codex/gpt-5.4-mini
0$3.02

Token composition — average per run

InputOutputCache readCache write
code/claude-haiku-4-5-20251001 · treat
215.3k
codex/gpt-5.4-mini · treat
1.2M

Statistics

MetricArmnMeanMedian (pooled)MinMaxStd dev
Costtreat90/100$0.17$0.04$0.01$3.02$0.41
Durationtreat100/1004m 55s4m 58s2.2s12m 13s2m 50s
Total tokenstreat90/100713.5k179.1k26.1k13.2M1.8M
Output tokenstreat90/1003.3k1.2k39643.4k6.3k
Cache-read tokenstreat90/100666.5k172.3k10.6k12.5M1.7M
Cache-write tokenstreat90/1003.9k0061.1k8.3k
Turnstreat91/10010.293447

n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.

What each setup did

Derived from each run's recorded tool calls — not from a model's description of the run — and aggregated per setup, so a behavior seen across several runs is stated once with its rate. Open a finding to see the runs behind it, each linked to its journey at the step where it happened.

Harness
Model
Task
Pick a finding or a setup to list the runs behind it.
claude-code · anthropic · claude-haiku-4-5-20251001 · Your setup1 finding
  • 6 of 50 runs · 0 completed
codex · openai · gpt-5.4-mini · Your setup4 findings
  • 4 of 50 runs · 4 completed
  • 4 of 50 runs · 4 completed
  • 4 of 50 runs · 4 completed
  • 3 of 50 runs · 0 completed

Task text withheld

Task text withheld — this benchmark is guarded and its tasks are not republished here.

Tasks

Runs

Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.

Every run, filterable… of 100Show runs
Task
Harness
Model
Outcome
Status
… of 100 runs match
Sort
HarnessModelTaskCompletedPassTokensCostDuration

Loading runs…

Methodology

What each metric means
Completed
The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
Denominator: Terminal runs, excluding those killed by our own infrastructure.
Quality index
A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
Denominator: Graded runs — runs the study's grader scored.
Infra-excluded
A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
Denominator: Reported as a count beside every affected panel.
Single arm
Every task runs once per harness × model cell — a shootout with no baseline arm. The readout is absolute (quality, success, tokens, cost) and the cells are ranked into a leaderboard.
One run per task
Every task ran once per cell, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
Sample size
2 cells × 50 tasks × 1 arm = 100 runs.