Anthropic says Claude Fable 5.1 scored 52.6% on the new Terminal-Bench-Science 0.1 benchmark, up from 24.7% for Fable 5 and 22.4% for GPT-5.6. The author tested Fable 5.1's five reasoning levels by generating SVG pelicans: low and medium skipped visible reasoning, while max produced the best result at 65,927 output tokens, 13 minutes 54 seconds, and $3.30. An animated version was created by piping the max output back at high effort.
No score is assigned. Sources and their independence are shown in the citation chain below.