Inconclusive · Study1views
acceptance row2 1790523436773
HELM · 1 task · 1 harness · 1 model
This benchmark's result hinges on infrastructure factors rather than measurable performance differences.
Every run so far was cut short by infrastructure, so the shootout measures nothing yet.
Abstract
This benchmark evaluated a single task across one harness/model cell in a single-arm shootout design. No baseline arm was run, making traditional efficiency comparisons unavailable. The treatment figures represent the pooled shootout results. The ranking across cells identifies claude-code / claude-sonnet-5 as the leader. The benchmark could not be decided on performance grounds; the outcome was determined on infrastructure considerations.
The result
No composite yet — nothing on this board was graded. A benchmark's own checks, the evals lens (pre-registered checks) or the judge panel (goal achievement and rubric quality) each score here.
Leaderboard
Cells ranked by rubric quality, then evals pass rate, then completion rate, then cost — absolute values, no baseline arm. Tokens, cost and duration are per-cell medians, the efficiency the score folds.
| # | Harness | Model | Score | Quality | Evals | Goal | Tokens | Cost | Duration | Runs |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-code | claude-sonnet-5 | − | − | − | − | − | − | − | 0/1 |
Distributions
Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.
Duration per run
Statistics
| Metric | Arm | n | Mean | Median (pooled) | Min | Max | Std dev |
|---|---|---|---|---|---|---|---|
| Duration | treat | 1/1 | 4m 51s | 4m 51s | 4m 51s | 4m 51s | − |
n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.
Task text withheld
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Tasks
Task text withheld — this benchmark is guarded and its tasks are not republished here.
Runs
Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.
Every run, filterable… of 1Show runsHide runs
| Harness | Model | Task | Completed | Quality | Tokens | Cost | Duration |
|---|
Loading runs…
Methodology
What each metric means
- Completed
- The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
- Denominator: Terminal runs, excluding those killed by our own infrastructure.
- Quality index
- A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
- Denominator: Graded runs — runs the study's grader scored.
- Infra-excluded
- A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
- Denominator: Reported as a count beside every affected panel.
- Single arm
- Every task runs once per harness × model cell — a shootout with no baseline arm. The readout is absolute (quality, success, tokens, cost) and the cells are ranked into a leaderboard.
- Answered by
- The agent harness in each cell. The benchmark's own runner built every prompt and computed every metric; each question was answered by one agent run — the model inside its harness, with the setup shown above — so these numbers are not comparable with the benchmark's bare-model leaderboard. Per-token log-probabilities, sampling parameters and token-efficiency metrics are not measured this way.
- One run per task
- Every task ran once per cell, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
- Sample size
- 1 cell × 1 task × 1 arm = 1 runs.