← Back to the wire

LabBench: Can AI agents decide what experiment to run next?

AchievementBenchmarkSep 28, 2026

Gamow Labs introduced LabBench, 20 held-out tasks from real wet-lab records in drug discovery and genomics, where agents must commit to the next experimental step. GPT-6 Astra and Claude Opus 5.5 tied, but Astra finished tasks in a median 5 minutes versus Opus's 35. Agents interpreted evidence well but passed only 21% of criteria requiring choosing or ranking, and no agent passed any of 13 criteria on which experiment should come first.

Receipt № 21461 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

GPT-6 AstraModelClaude Opus 5.5Model
Canonical: https://gamowlabs.com/labbench-benchmarking-ai-wet-lab-decisions.html