Evaluation design
Why scientific reasoning needs its own benchmark
Static QA benchmarks reward recall. Scientific work rewards mechanism, feasibility and evidence — which only shows up in open-ended answers judged by domain scientists.
Blog
Short write-ups on how the arena is built and what the votes tell us.
Static QA benchmarks reward recall. Scientific work rewards mechanism, feasibility and evidence — which only shows up in open-ended answers judged by domain scientists.
Sequential Elo updates depend on match order. Fitting a Bradley-Terry model over the full vote history gives order-independent strengths and stable ratings at low volume.
Running five matches per task yields a complete ranking of four models per submission, producing far more pairwise signal for the same amount of scientist attention.
More on the evaluation approach in about & methodology.