LifeSciBench: GPT-Rosalind 36.1%, artifact gap 17pts

openailifesci-benchlifescibenchlife-sciencesbenchmarkgpt-rosalind+13
LifeSciBench 16:9 illustration card
Image: OpenAI / LifeSciBench announcement (June 17, 2026)

On 2026-06-17, OpenAI published LifeSciBench — a 750-task, 1,062-artifact, 19,020-criterion life-sciences evaluation built with 173 PhD-level scientists and validated by 453 independent expert reviewers (OpenAI, Introducing LifeSciBench, 2026-06-17). GPT-Rosalind, a new life-sciences model introduced on a request-access basis, reaches 36.1% exact pass rate vs 25.7% for GPT-5.5. The under-reported finding: GPT-Rosalind drops from 45.1% on text-only tasks to 28.1% on tasks with artifacts — a 17-point drop.

What it is

LifeSciBench contains 750 expert-authored tasks across seven biological domains and seven workflows: evidence handling, analysis, design and optimization, scientific reasoning, validation and operations, translation, and scientific communication (OpenAI, Introducing LifeSciBench, 2026-06-17).

The benchmark reports two metrics: pass rate (70% rubric threshold) and score (average rubric reward with partial credit). The rubric-reward number is the more honest capability signal.

The artifact-handling gap

The most useful number in the announcement is buried in the “Where AI systems still fall short” section:

Both models drop the same amount. OpenAI confirms frontier models struggle at extracting information from complex figures or large sequence files (OpenAI, Introducing LifeSciBench, 2026-06-17).

For AI builders shipping agents that touch real lab data — PDFs, sequence files, gel images, instrument exports — this is the number to plan around. The 36.1% headline is a ceiling on text-only performance.

Where the gains are

The improvement is concentrated in talking-about workflows (communication, translation), not designing-and-analyzing workflows.

Partial-credit pattern

In roughly 14% of tasks, models earned substantial rubric credit despite failing the exact-pass threshold. GPT-Rosalind had 109 tasks with pass rates below 20% while earning at least 50% rubric reward — models can identify relevant evidence but miss key constraints.

Why it matters

The 19,020-criterion rubric is the real contribution. LifeSciBench asks “is the model useful in a research meeting?” rather than “did the model get the right answer?” The seven-workflow taxonomy maps cleanly onto how biotech R&D teams talk about their work.

Risks and caveats

  1. Self-report by the model owner. OpenAI is both benchmark publisher and model owner. No third-party reproduction exists as of 2026-06-20. No public code or task data is referenced in the announcement.
  2. 36.1% is a 70% rubric threshold, not a clean “right answer.” The rubric-reward number (44.7% vs 29.1% on expert-useful outputs) is the more honest signal.
  3. Scientific Communication gain is on n=9 tasks. Too small to generalize.
  4. Artifact handling is the load-bearing weakness. The 17-point drop is the number most useful to AI builders.
  5. Real research is iterative; LifeSciBench is not. OpenAI names deployment studies in live workflows as the proof point the benchmark cannot supply.

What to watch

Verdict

LifeSciBench shows frontier models improving rapidly at scientific prose and bench-to-bedside translation, and not measurably at design-and-analysis work. The 17-point artifact-handling drop is the number to plan around. It is a self-report by the model owner, GPT-Rosalind access is gated, and the strongest gain is on nine tasks (OpenAI, Introducing LifeSciBench, 2026-06-17).