AdaPlanBench, introduced by Jiayu Liu, is a dynamic interactive benchmark evaluating whether LLM agents can adaptively plan under progressively revealed world and user constraints. Built on 307 household tasks, it reveals hidden constraints only when agents propose violating plans, requiring iterative revision. Experiments across ten leading LLMs show the best model reached only 67.75% accuracy, with performance degrading as constraints accumulated and user constraints posing a particularly large challenge.
No score is assigned. Sources and their independence are shown in the citation chain below.