analysis

1 article

← All topics

LifeSciBench: GPT-Rosalind 36.1%, artifact gap 17pts
openailifesci-benchlifescibenchlife-sciencesbenchmark+14

OpenAI published LifeSciBench, a 750-task life-sciences evaluation. GPT-Rosalind hits 36.1% vs GPT-5.5's 25.7%, but drops 17 points on tasks with artifacts.