Gamow Labs introduced LabBench, 20 held-out tasks from real wet-lab records in drug discovery and genomics, where agents must commit to the next experimental step. GPT-6 Astra and Claude Opus 5.5 tied, but Astra finished tasks in a median 5 minutes versus Opus's 35. Agents interpreted evidence well but passed only 21% of criteria requiring choosing or ranking, and no agent passed any of 13 criteria on which experiment should come first.
No score is assigned. Sources and their independence are shown in the citation chain below.