Inconclusive · Study2views

acceptance row5 skills 1790524032999

ora Agent Smoke (5 tasks) · 1 task · 1 harness · 1 model

Ora Agent Smoke delivered flawless execution on this task.

Ranked shootout, no baseline arm: cells ranked by the composite score — rubric quality and goal achievement, then efficiency — claude-code / claude-haiku-4-5-20251001 leads.

Abstract

The treatment with ora Agent Smoke achieved a median cost of $0.01 (mean $0.01), consumed 34.5k tokens (mean 34.5k), and completed in 16.5s (mean 16.5s). Quality was perfect at 100/100 on the rubric. The benchmark ran a single harness/model cell (claude-code / claude-haiku-4-5-20251001) in shootout mode with no baseline arm for comparison. The treatment was decided the winner on quality grounds.

The result

Color by

claude-code

  1. 1claude-code / claude-haiku-4-5-20251001100.0best

Overall score per cell, 0–100 points (not a pass rate): 30% goal + 30% quality + 25% efficiency (cost · tokens · duration, vs the board's best arm). Absent components renormalize. Best at the top — the board's best setup is tagged.

Leaderboard

Cells ranked by the composite score — outcome (the lenses that graded) × efficiency — absolute values, no baseline arm. Tokens, cost and duration are per-cell medians, the efficiency the score folds.

#HarnessModelScorePass rateQualityEvalsGoalTokensCostDurationRuns
1claude-codeclaude-haiku-4-5-20251001100.0100%100−100%34.5k$0.0116.5s1/1

Pass rate

100%over 1 graded run
CellPass rateTasksRunsRan asMatcher
claude-code/claude-haiku-4-5-20251001100%11with my setupnormalized

Each run's answer was extracted from its response and matched against the benchmark's own key — no model in the loop. Two-stage: the mean over repeats per task, then the mean over tasks, with the interval taken over tasks.

Quality × efficiency clusters — normalized across the study

Every graded run, standardized across the study — each task ran once, so there is no field per task to compare against: → right = fewer tokens than the study's average run, ↑ up = higher quality index (70% rubric pass rate + 15% intent + 15% outcome) than average. Every run scored alike on quality here, so the vertical axis separates nothing — read the map left to right. The field pools every harness and model, so a setup's left–right position largely reflects its own token habits. Dots are runs; each harness logo is a harness × model × arm centroid — hover it for the model and averages. Up-right wins.

Show
Color by

claude-code

better · cheaperworse · pricier
← pricier than the fieldefficiency (σ, per task)cheaper than the field →

Every metric, per harness × model

Benchmark pass rate

  1. claude-code / claude-haiku-4-5-20251001100%

Task completion

  1. claude-code / claude-haiku-4-5-20251001100%

Goal achievement

  1. claude-code / claude-haiku-4-5-20251001100%

Rubric quality /100

  1. claude-code / claude-haiku-4-5-20251001100

Cost per run

  1. claude-code / claude-haiku-4-5-20251001$0.01

Tokens per run

  1. claude-code / claude-haiku-4-5-2025100134.5k

Duration per run

  1. claude-code / claude-haiku-4-5-2025100116.5s

Rubric criteria passed

  1. claude-code / claude-haiku-4-5-20251001100%

Your setup

skills

The publisher has not described this setup — what each group's tools are, who built them, and how the tasks were chosen.

Distributions

Every completed run is one dot — the spread the averages hide. Click a dot to replay that run's journey.

Rubric quality per run

Your setup
code/claude-haiku-4-5-20251001
0100

Token composition — average per run

InputOutputCache readCache write
code/claude-haiku-4-5-20251001 · treat
34.5k

Efficiency frontier

Loading
Your setup

Statistics

MetricArmnMeanMedian (pooled)MinMaxStd dev
Rubric qualitytreat1/1100100100100−
Costtreat1/1$0.01$0.01$0.01$0.01−
Durationtreat1/116.5s16.5s16.5s16.5s−
Total tokenstreat1/134.5k34.5k34.5k34.5k−
Output tokenstreat1/1284284284284−
Cache-read tokenstreat1/131.3k31.3k31.3k31.3k−
Cache-write tokenstreat1/13k3k3k3k−
Turnstreat1/12222−

n = runs carrying the fact / completed runs in the arm; every statistic runs over present facts only — a missing fact is never counted as 0. Std dev is the sample form (n−1), withheld below n = 2. Every figure here POOLS all runs. The typical-run figures in the abstract, the head-to-head decision and the per-setup panels take each task's median first and then the median across tasks, so a task that ran more often never outweighs one that ran once; the two medians can differ. A promise study's headline (“cut average tokens”) compares per-arm means.

What each setup did

Derived from each run's recorded tool calls — not from a model's description of the run — and aggregated per setup, so a behavior seen across several runs is stated once with its rate. Open a finding to see the runs behind it, each linked to its journey at the step where it happened.

Pick a finding or a setup to list the runs behind it.
claude-code · anthropic · claude-haiku-4-5-20251001 · Your setup1 finding
  • 1 of 1 run · 1 completed

Task text withheld

Task text withheld — this benchmark is guarded and its tasks are not republished here.

Tasks

Runs

Every run is inspectable — open one to replay the agent's journey step by step, with the analyst's read underneath.

Every run, filterable… of 1Show runs
Sort
HarnessModelTaskCompletedQualityTokensCostDuration

Loading runs…

Methodology

What each metric means
Completed
The run finished and the harness returned a response. It does NOT mean the answer was correct — a completed run can score zero on quality.
Denominator: Terminal runs, excluding those killed by our own infrastructure.
Goal achievement
The run fully achieved the task's goal: the judge panel passed every pre-registered rubric criterion (a rubric score of 100).
Denominator: Graded completed runs.
Rubric quality
The share of pre-registered acceptance criteria the judge panel passed, as a 0-100 score. Partial credit is possible.
Denominator: The criteria count frozen before any run — never the judge's returned count.
Quality index
A composite used only in the quality x efficiency plot, built from whichever grader scored this study's runs — the judge's rubric, the pre-registered checks or the benchmark's own verdict — plus whether the run's own outcome was success. The map's caption names the grader and its weights.
Denominator: Graded runs — runs the study's grader scored.
Pairwise verdict
A blind, position-debiased comparison of the two arms' final answers. It sees answer text only — cost and latency are measured separately.
Denominator: Task-cell pairs where both arms produced a response.
Infra-excluded
A run killed by our own infrastructure. It is a missing measurement, never a loss for the arm, and is excluded from every rate denominator.
Denominator: Reported as a count beside every affected panel.
Single arm
Every task runs once per harness × model cell — a shootout with no baseline arm. The readout is absolute (quality, success, tokens, cost) and the cells are ranked into a leaderboard.
Pre-registered rubrics
Each task's acceptance criteria are written at task generation, before any run exists, so grading can never be shaped by the results.
Judge panel
A panel of independent judges (claude-opus-5, gpt-5.5, claude-fable-5), each at provider-default sampling (claude-opus-5, gpt-5.5 and claude-fable-5 do not accept a temperature setting), scores every successful response against its rubric; a criterion passes only when a strict majority of the panel passes it. Each judge sees only the task, the rubric, and the response — never which arm produced it, never token counts.
One run per task
Every task ran once per cell, in one phrasing. Run-to-run variation is therefore not measured: a single task's difference can be one lucky or unlucky attempt. Every row of the reference table treats the TASKS as the unit: its intervals resample the tasks and its p-values come from a paired test over them, so they describe how much the result depends on which tasks were drawn — not how it would change if the same tasks were run again. With few tasks no difference can reach significance (five tasks cannot go below p = 1/16).
Sample size
1 cell × 1 task × 1 arm = 1 runs.