Do agent upgrades deliver what they promise? We measure.

Each study takes one subject's promise — a skill, a tool policy, an add-on — and measures it across real harness × model cells: A/B studies run every task under both arms and read the promise off the deltas; shootouts run one arm per cell and rank the field.

WindTunnel

WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.

+0ppcompletion rate(+0% relative)
100%100%completion rate
claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 50 tasks13views

Not deliveredBenchmarkSep 28, 2026

WindTunnel

WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.

Top of the board claude-code / claude-haiku-4-5-20251001

claude-code — claude-haiku-4-5-20251001codex — gpt-5.4-mini

2 cells · 50 tasks5views

MeasuredBenchmarkSep 28, 2026

WindTunnel

WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.

Top of the board claude-code / claude-haiku-4-5-20251001

claude-code — claude-haiku-4-5-20251001

1 cell · 50 tasks2views

MeasuredBenchmarkSep 28, 2026

ora Agent Smoke (5 tasks)

ora's own five-task smoke corpus: short, deterministic, text-only items that every harness and model can attempt. It exists to prove the benchmark loop end to end — not to rank frontier capability.

Results pending — runs still completing.

claude-code — claude-haiku-4-5-20251001

1 cell · 1 task2views

InconclusiveBenchmarkSep 27, 2026

MBPP

Mostly Basic Python Problems — short entry-level tasks, each with three assertions that decide the verdict.

Top of the board claude-code / claude-haiku-4-5-20251001

claude-code — claude-haiku-4-5-20251001

1 cell · 3 tasks1views

MeasuredBenchmarkSep 27, 2026

HELM

Stanford's holistic evaluation framework — many scenarios and metrics driven by HELM's own runner, reported as the framework reports them.

Results pending — runs still completing.

claude-code — claude-sonnet-5

1 cell · 1 task1views

InconclusiveBenchmarkSep 27, 2026

ora Agent Smoke (5 tasks)

ora's own five-task smoke corpus: short, deterministic, text-only items that every harness and model can attempt. It exists to prove the benchmark loop end to end — not to rank frontier capability.

Results pending — runs still completing.

claude-code — claude-haiku-4-5-20251001

1 cell · 1 task2views

InconclusiveBenchmarkSep 27, 2026

MBPP

Mostly Basic Python Problems — short entry-level tasks, each with three assertions that decide the verdict.

Top of the board claude-code / claude-haiku-4-5-20251001

claude-code — claude-haiku-4-5-20251001

1 cell · 3 tasksNo views yet

MeasuredBenchmarkSep 27, 2026

HELM

Stanford's holistic evaluation framework — many scenarios and metrics driven by HELM's own runner, reported as the framework reports them.

Results pending — runs still completing.

claude-code — claude-sonnet-5

1 cell · 1 task1views

InconclusiveBenchmarkSep 27, 2026

GSM8K

Grade-school word problems that need several arithmetic steps — the classic multi-step reasoning check.

Top of the board openai-compatible / accounts/fireworks/models/gpt-oss-120b

openai-compatible — accounts/fireworks/models/gpt-oss-120b

2 cells · 3 tasks1views

MeasuredBenchmarkSep 27, 2026

GSM8K

Grade-school word problems that need several arithmetic steps — the classic multi-step reasoning check.

Top of the board claude-consumer / claude-sonnet-5

claude-consumer — claude-sonnet-5

1 cell · 3 tasksNo views yet

MeasuredBenchmarkSep 27, 2026