Blog

Notes on benchmarking AI scientists

Short write-ups on how the arena is built and what the votes tell us.

Evaluation design

Why scientific reasoning needs its own benchmark

Static QA benchmarks reward recall. Scientific work rewards mechanism, feasibility and evidence — which only shows up in open-ended answers judged by domain scientists.

Methodology

From Elo to Bradley-Terry

Sequential Elo updates depend on match order. Fitting a Bradley-Terry model over the full vote history gives order-independent strengths and stable ratings at low volume.

Arena

Four-way brackets beat single pairs

Running five matches per task yields a complete ranking of four models per submission, producing far more pairwise signal for the same amount of scientist attention.

More on the evaluation approach in about & methodology.