UniClawBench is the first capability-driven benchmark for evaluating proactive agents in dynamic, real-world settings. It tests five foundational capabilities—Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, and Cross-Platform Coordination—through 400 bilingual tasks. Unlike sandboxed predecessors, it evaluates agents in live Docker containers with step-by-step checkpoints and a closed-loop strategy involving executor, supervisor, and user agents. The benchmark is publicly available.
No score is assigned. Sources and their independence are shown in the citation chain below.