PROVENANCE FIRST

Methodology and boundaries

This site presents traceable archive records. Missing values are not zero, and incompatible measurements are not collapsed into a leaderboard.

01

Provenance first

Every prompt, run metric, and review retains its source file and line location.

02

Reviews stay separate

Human reviews, historical AI reports, and this site verification are never silently merged.

03

Missing means missing

Unknown configuration is shown as not recorded; per-task costs and capabilities are not inferred.

04

Compare like with like

Only runs for the same task version can be compared; deltas require matching measurement definitions.

How historical conclusions are presented

The 93.6/100 score and 15/15 statement are quoted only from the historical “15-task completion quality assessment.” They are not recomputed or independently proven by this website. For tasks 06 and 07, the human visual judgment differs materially from the AI assessment, so both are retained side by side.

Source report ↗
BATCH MEASUREMENTS

Total

Tasks15
Sessions12
Duration8,282s
API693
Tools735
Failures / retries42
Input tokens334,670
Output tokens1,072,534
Cache read63,042,176
Batch cost¥9.30