Inconclusive · Study1views

acceptance row2 1790523436773

HELM · 1 task · 1 harness · 1 model

This benchmark's result hinges on infrastructure factors rather than measurable performance differences.

Every run so far was cut short by infrastructure, so the shootout measures nothing yet.

Abstract

This benchmark evaluated a single task across one harness/model cell in a single-arm shootout design. No baseline arm was run, making traditional efficiency comparisons unavailable. The treatment figures represent the pooled shootout results. The ranking across cells identifies claude-code / claude-sonnet-5 as the leader. The benchmark could not be decided on performance grounds; the outcome was determined on infrastructure considerations.

The result

No composite yet — nothing on this board was graded. A benchmark's own checks, the evals lens (pre-registered checks) or the judge panel (goal achievement and rubric quality) each score here.

Leaderboard

Cells ranked by rubric quality, then evals pass rate, then completion rate, then cost — absolute values, no baseline arm. Tokens, cost and duration are per-cell medians, the efficiency the score folds.

#HarnessModelScoreQualityEvalsGoalTokensCostDurationRuns
1claude-codeclaude-sonnet-5−−−−−−−0/1

Distributions

Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.

Duration per run

Your setup
code/claude-sonnet-5
04m 51s

Statistics

MetricArmnMeanMedian (pooled)MinMaxStd dev
Durationtreat1/14m 51s4m 51s4m 51s4m 51s−

n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.

Task text withheld

Task text withheld — this benchmark is guarded and its tasks are not republished here.

Tasks

Runs

Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.

Every run, filterable… of 1Show runs
Sort
HarnessModelTaskCompletedQualityTokensCostDuration

Loading runs…

Methodology

What each metric means
Completed
The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
Denominator: Terminal runs, excluding those killed by our own infrastructure.
Quality index
A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
Denominator: Graded runs — runs the study's grader scored.
Infra-excluded
A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
Denominator: Reported as a count beside every affected panel.
Single arm
Every task runs once per harness × model cell — a shootout with no baseline arm. The readout is absolute (quality, success, tokens, cost) and the cells are ranked into a leaderboard.
Answered by
The agent harness in each cell. The benchmark's own runner built every prompt and computed every metric; each question was answered by one agent run — the model inside its harness, with the setup shown above — so these numbers are not comparable with the benchmark's bare-model leaderboard. Per-token log-probabilities, sampling parameters and token-efficiency metrics are not measured this way.
One run per task
Every task ran once per cell, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
Sample size
1 cell × 1 task × 1 arm = 1 runs.