OpenAI has released a new evaluation benchmark, LifeSciBench, designed to measure AI systems’ capabilities in real scientific research settings. According to Odaily, LifeSciBench is built on 750 expert-written tasks spanning seven research workflows and seven biology domains.
The tasks were contributed by 173 researchers with PhDs and experience in biotechnology or the pharmaceutical industry. The benchmark focuses on assessing complex research abilities, including evidence integration, experimental design, data analysis, scientific reasoning, and research communication, rather than single fact-based questions.
More than 79% of the tasks involve multi-step reasoning, with an average of about four reasoning steps per task. The benchmark also includes 1,062 real research-related data attachments, such as papers, charts, sequence data, and structure files.