Sunshine Jiang introduced Prompt-Driven Exploration (PDE), a strategy where a vision-language model analyzes rollout video and rewrites prompts to guide reinforcement learning policies. PDE applies posterior sampling at the prompt level, enabling global behavioral changes that action-space noise cannot achieve. Experiments showed PDE allowed RL to learn successful policies from zero-reward starts and improved sample efficiency across manipulation and reasoning tasks.
No score is assigned. Sources and their independence are shown in the citation chain below.