← Back to the wire

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

AchievementBenchmarkJul 10, 2026

AdaPlanBench, introduced by Jiayu Liu, is a dynamic interactive benchmark evaluating whether LLM agents can adaptively plan under progressively revealed world and user constraints. Built on 307 household tasks, it reveals hidden constraints only when agents propose violating plans, requiring iterative revision. Experiments across ten leading LLMs show the best model reached only 67.75% accuracy, with performance degrading as constraints accumulated and user constraints posing a particularly large challenge.

Receipt № 4631 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
AdaPlanBenchModelJiayu LiuPerson
Canonical: https://arxiv.org/abs/2606.05622