Yikai Wang proposes distributionally robust regret optimization (DRRO) for RLHF to address reward over-optimization, where proxy rewards improve as true quality deteriorates. DRRO pessimizes worst-case regret rather than worst-case value, using a Wasserstein ambiguity set over reward laws. The method decomposes into promptwise regret problems with closed-form adversaries, yielding a practical policy-gradient algorithm that adds a sampled bonus to GRPO-style training. Experiments show DRRO mitigates over-optimization more effectively than existing baselines.
No score is assigned. Sources and their independence are shown in the citation chain below.