← Back to the wire

IBM and UC Berkeley Diagnose Why Enterprise Agents Fail Using IT-Bench and MAST

AchievementResearchFeb 18, 2026

IBM and UC Berkeley applied MAST (Multi-Agent System Failure Taxonomy) to 310 ITBench SRE traces to diagnose why enterprise agents fail in IT automation tasks. Researchers including Ayhan Sebin, Saurabh Jha, Rohan Arora, Daby Sow, Mert Cemri, Melissa Pan, and Ion Stoica found that Gemini-3-Flash exhibits "surgical failure" profiles with isolated errors, while open-source models Kimi-K2 and GPT-oss-120b show compounding failure patterns where errors cascade over time. The analysis classified failures as "non-fatal" or "fatal," moving beyond simple success-rate metrics.

Receipt № 5531 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

Ayhan SebinPersonUC BerkeleyCompanyGemini-3-FlashModelKimi-K2ModelGPT-oss-120bModelRohan AroraPersonSaurabh JhaPersonDaby SowPersonMert CemriPersonMelissa PanPersonIon StoicaPersonIBMCompany
Canonical: https://huggingface.co/blog/ibm-research/itbenchandmast