← Back to the wire

Agentic systems for breast cancer treatment recommendations

AchievementResearchJul 13, 2026

Researchers evaluated agentic LLM systems on 72 breast cancer cases using 1,147 case-specific rubrics generated via Asymmetric Information Rubric Generation. The best-performing configuration, Claude Opus 4.8 with the D&C+SA pipeline, achieved a global score of 0.594 ± 0.025. Tool use and increased agent autonomy produced mixed results across clinical domains and disease stages. Oncologist-led error analysis identified persistent failures including incorrect recommendations, citation errors, and overconfidence, concluding the systems remain insufficient for unsupervised clinical use.

Receipt № 6761 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
Claude Opus 4.8Model
Canonical: https://arxiv.org/abs/2607.12051v1