Researchers evaluated agentic LLM systems on 72 breast cancer cases using 1,147 case-specific rubrics generated via Asymmetric Information Rubric Generation. The best-performing configuration, Claude Opus 4.8 with the D&C+SA pipeline, achieved a global score of 0.594 ± 0.025. Tool use and increased agent autonomy produced mixed results across clinical domains and disease stages. Oncologist-led error analysis identified persistent failures including incorrect recommendations, citation errors, and overconfidence, concluding the systems remain insufficient for unsupervised clinical use.
No score is assigned. Sources and their independence are shown in the citation chain below.