← Back to the wire

ATLAS-Finance: Evaluating AI Agents Inside a Bank

AchievementBenchmarkSep 14, 2026

Claude Opus 5 achieved the highest pass rate, 12.3%, on ATLAS-Finance, a new benchmark of 100 expert-level tasks set in 13 realistic financial firm environments. Claude Fable 5.1 and GPT-6 Astra scored 12.0% and 11.3%, while the other eight models tested fell below 10%. Common failures included applying wrong financial logic, omitting required scope, and failing to propagate correctly calculated values downstream. Tasks take human experts 15-30 hours and are graded against expert-authored rubrics.

Receipt № 19271 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

Claude Opus 5ModelClaude Fable 5.1ModelGPT-6 AstraModel
Canonical: https://joinhandshake.com/research/benchmarks/articles/atlas-finance-evaluating-ai-agents-inside-a-bank/