← Back to the wire

Prompt-Driven Exploration

AnnouncementResearchJul 9, 2026

Sunshine Jiang introduced Prompt-Driven Exploration (PDE), a strategy where a vision-language model analyzes rollout video and rewrites prompts to guide reinforcement learning policies. PDE applies posterior sampling at the prompt level, enabling global behavioral changes that action-space noise cannot achieve. Experiments showed PDE allowed RL to learn successful policies from zero-reward starts and improved sample efficiency across manipulation and reasoning tasks.

Receipt № 5941 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
Sunshine JiangPerson
Canonical: https://arxiv.org/abs/2607.08837