External pipelines · RISE project catalogue
NatureBench
A cross-discipline benchmark (arXiv:2606.24530) of 90 tasks distilled from peer-reviewed Nature-family papers across six scientific domains, asking whether AI coding agents can match — or surpass — the published state of the art. Each task is a containerized package (task brief, the paper's dataset, a held-out test set with hidden ground truth, an automated evaluator) built by NatureGym, an automated Claude-Code-skills pipeline that converts a published paper into an executable Docker task. Sits in the RISE evaluation-infrastructure layer alongside AstaBench, MLGym, and EconCS Bench, but targets empirical scientific ML with executable, SOTA-anchored scoring.
Where it sits
left: what it builds on · right: what builds on it · pale: exampleContributed by FrontisAI
How studies reach it
No published study reaches it yet.
Disciplines it reaches
No study reaches it yet.
Solid: published studies. Light: examples.
Computed from the records on this site: what each study, template and specialist names as used, which study extends which, and who contributed what. 0 studies in total.
What it does
Scores agents against each source paper's *reported SOTA* (Surpass-SOTA rate) rather than mere reproduction, with an information firewall that strips the source method from the task brief so agents must discover solutions. Ships a full harness with built-in adapters for Claude Code, Codex CLI, and Gemini CLI, a post-hoc validity judge, and a public leaderboard: across twelve harness-model configurations the best reaches a 17.8% Surpass-SOTA rate, with failures dominated by method-selection errors.
- Focus
- analysis
- Inputs
- nature-family-papers, coding-agent-configurations
- Outputs
- containerized-task-packages, surpass-sota-scores, leaderboard
- Architecture
- tool-use
- Maintained by
- FrontisAI
- Started
- 2026
Description
Data model- Discipline
- General
- Method family
- not specified
- Design
- not specified
- Research stage
- Data analysisCode generation
- Contributors
- FrontisAI
- Usage
- not used in published research yet
- Source
- RISE project catalogue · projects/landscape · @4c17bae
- Record
- pipeline:naturebench · JSON
Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.
Bring it into the standard
A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.