Evaluation infrastructure for AI scientists

Scientific Arena

Evaluate AI models on real scientific reasoning tasks, from hypothesis generation to experiment design.

Benchmark tasks
Models tracked
Judged comparisons
0
Line illustration of a neuron network merging into a molecular lattice and a spectral trace

How it works

A reproducible pipeline from research question to ranking

01

Submit a scientific task

Researchers contribute real questions with background, evidence files, expected output and a rubric.

02

Run frontier models

The same task is dispatched to OpenAI, Anthropic, Google and open-source models through one provider interface.

03

Blind pairwise judging

Scientists see two anonymised answers side by side and pick the stronger one — no model names, no branding.

04

Rankings that update live

Every vote adjusts Elo overall and per domain and capability, alongside structured 1–5 rating dimensions.

Scientific capabilities evaluated

Six axes of scientific competence, judged separately

Literature synthesis

Integrating dispersed findings into a coherent, sourced account.

Mechanistic reasoning

Building causal chains that respect known biology and chemistry.

Hypothesis generation

Proposing novel, falsifiable explanations worth testing.

Experiment design

Specifying feasible protocols, controls and powered readouts.

Data interpretation

Reading real datasets without overclaiming or ignoring confounds.

Research planning

Sequencing a programme of work under real resource constraints.

Leaderboards

Current overall standing

Full leaderboard
RankModelProviderEloW / L
No comparisons recorded yet — be the first to judge in the arena.

For researchers and AI labs

For researchers

Contribute the questions you actually care about, define the rubric that matters in your field, and see which systems hold up under domain expertise rather than leaderboard folklore.

  • — Author tasks with evidence attachments and expected outputs
  • — Judge blindly and record five structured rating dimensions
  • — Track your own evaluation history and agreement over time

For AI labs

A provider-neutral harness with per-domain and per-capability breakdowns, so regressions in mechanistic reasoning are not hidden by gains in fluency.

  • — One model interface for OpenAI, Anthropic, Google and open weights
  • — Programmatic endpoints for models, generation and evaluation
  • — Domain and capability slices of the Elo table