Scientific Arena
Evaluate AI models on real scientific reasoning tasks, from hypothesis generation to experiment design.
- Benchmark tasks
- —
- Models tracked
- —
- Judged comparisons
- 0

How it works
A reproducible pipeline from research question to ranking
Submit a scientific task
Researchers contribute real questions with background, evidence files, expected output and a rubric.
Run frontier models
The same task is dispatched to OpenAI, Anthropic, Google and open-source models through one provider interface.
Blind pairwise judging
Scientists see two anonymised answers side by side and pick the stronger one — no model names, no branding.
Rankings that update live
Every vote adjusts Elo overall and per domain and capability, alongside structured 1–5 rating dimensions.
Scientific capabilities evaluated
Six axes of scientific competence, judged separately
Literature synthesis
Integrating dispersed findings into a coherent, sourced account.
Mechanistic reasoning
Building causal chains that respect known biology and chemistry.
Hypothesis generation
Proposing novel, falsifiable explanations worth testing.
Experiment design
Specifying feasible protocols, controls and powered readouts.
Data interpretation
Reading real datasets without overclaiming or ignoring confounds.
Research planning
Sequencing a programme of work under real resource constraints.
Leaderboards
Current overall standing
| Rank | Model | Provider | Elo | W / L |
|---|---|---|---|---|
| No comparisons recorded yet — be the first to judge in the arena. | ||||
For researchers and AI labs
For researchers
Contribute the questions you actually care about, define the rubric that matters in your field, and see which systems hold up under domain expertise rather than leaderboard folklore.
- — Author tasks with evidence attachments and expected outputs
- — Judge blindly and record five structured rating dimensions
- — Track your own evaluation history and agreement over time
For AI labs
A provider-neutral harness with per-domain and per-capability breakdowns, so regressions in mechanistic reasoning are not hidden by gains in fluency.
- — One model interface for OpenAI, Anthropic, Google and open weights
- — Programmatic endpoints for models, generation and evaluation
- — Domain and capability slices of the Elo table