← Back to the wire

Inside VAKRA: Reasoning, Tool Use, and Failure Modes of Agents

AchievementBenchmarkApr 15, 2026

IBM Research and Hugging Face introduced VAKRA, a tool-grounded, executable benchmark evaluating how AI agents reason and act in enterprise-like environments. VAKRA measures compositional reasoning across APIs and documents, featuring over 8,000 locally hosted APIs spanning 62 domains with tasks requiring 3-7 step reasoning chains. The benchmark includes four capability categories, including API chaining and tool selection. Developers Ankita Naik, Danish, Ben, Anupama Murthi, Praveen Venkateswaran, Siyu, and Ayhan Sebin report that models perform poorly on VAKRA.

Receipt № 5141 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

Canonical: https://huggingface.co/blog/ibm-research/vakra-benchmark-analysis