External pipelines · RISE project catalogue
LifeSciBench
OpenAI's expert-authored benchmark (announced 2026-06-17) for measuring how well AI models support real-world life-science research. 750 free-response tasks span seven workflows — evidence handling, analysis, design/optimization, scientific reasoning, validation/operations, translation, and scientific communication — across seven biological domains, each pairing a scientific prompt, supporting artifacts, and an expert-written grading rubric. Sits in the RISE evaluation-infrastructure layer alongside AstaBench and EconCS Bench, but targets biomedical research capability.
Where it sits
left: what it builds on · right: what builds on it · pale: exampleContributed by OpenAI
How studies reach it
No published study reaches it yet.
Disciplines it reaches
No study reaches it yet.
Solid: published studies. Light: examples.
Computed from the records on this site: what each study, template and specialist names as used, which study extends which, and who contributed what. 0 studies in total.
What it does
Grades models against ~19,020 rubric criteria (~25 per task) decomposing each expected answer into individual claims, calculations, decisions, justifications, and caveats — authored by 173 PhD-level biotech/pharma scientists and validated by 453 independent expert reviewers (97% doctorate-holding, >96% agreement). Roughly 53% of tasks attach real scientific artifacts (genomic sequences, chemical structures, figures, tables, PDFs), and the best model (GPT-Rosalind) clears only 36.1% of tasks — with a sharp drop from text-only to artifact-bearing tasks.
- Focus
- end-to-end
- Inputs
- task-prompts, scientific-artifacts
- Outputs
- rubric-scores, model-pass-rates
- Maintained by
- OpenAI
- Started
- 2026
Description
Data model- Discipline
- Biomedical
- Method family
- not specified
- Design
- not specified
- Research stage
- Literature synthesisResearch designData analysisDrafting
- Contributors
- OpenAI
- Usage
- not used in published research yet
- Source
- RISE project catalogue · projects/landscape · @4c17bae
- Record
- pipeline:lifescibench · JSON
Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.
Bring it into the standard
A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.