Do agent upgrades deliver what they promise? We measure.
Each study takes one subject's promise — a skill, a tool policy, an add-on — and measures it across real harness × model cells: A/B studies run every task under both arms and read the promise off the deltas; shootouts run one arm per cell and rank the field.
WindTunnel
WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.
2 cells · 50 tasks13views
WindTunnel
WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.
Top of the board claude-code / claude-haiku-4-5-20251001
2 cells · 50 tasks5views
WindTunnel
WindTunnel, by nekuda, measures how well AI agents complete real tasks on the web. Across eight realistic sites (shops, booking, ticketing, a CRM, a learning platform, a blog, a directory and a members-only app), agents look up answers or take actions, and every task is scored automatically against the site itself.
Top of the board claude-code / claude-haiku-4-5-20251001
1 cell · 50 tasks2views
ora Agent Smoke (5 tasks)
ora's own five-task smoke corpus: short, deterministic, text-only items that every harness and model can attempt. It exists to prove the benchmark loop end to end — not to rank frontier capability.
Results pending — runs still completing.
1 cell · 1 task2views
MBPP
Mostly Basic Python Problems — short entry-level tasks, each with three assertions that decide the verdict.
Top of the board claude-code / claude-haiku-4-5-20251001
1 cell · 3 tasks1views
HELM
Stanford's holistic evaluation framework — many scenarios and metrics driven by HELM's own runner, reported as the framework reports them.
Results pending — runs still completing.
1 cell · 1 task1views
ora Agent Smoke (5 tasks)
ora's own five-task smoke corpus: short, deterministic, text-only items that every harness and model can attempt. It exists to prove the benchmark loop end to end — not to rank frontier capability.
Results pending — runs still completing.
1 cell · 1 task2views
MBPP
Mostly Basic Python Problems — short entry-level tasks, each with three assertions that decide the verdict.
Top of the board claude-code / claude-haiku-4-5-20251001
1 cell · 3 tasksNo views yet
HELM
Stanford's holistic evaluation framework — many scenarios and metrics driven by HELM's own runner, reported as the framework reports them.
Results pending — runs still completing.
1 cell · 1 task1views
GSM8K
Grade-school word problems that need several arithmetic steps — the classic multi-step reasoning check.
Top of the board openai-compatible / accounts/fireworks/models/gpt-oss-120b
2 cells · 3 tasks1views
GSM8K
Grade-school word problems that need several arithmetic steps — the classic multi-step reasoning check.
Top of the board claude-consumer / claude-sonnet-5
1 cell · 3 tasksNo views yet