Zixuan Wang's research finds that training models under power-law distributions outperforms uniform distributions on compositional reasoning tasks like state tracking and multi-step arithmetic. The study shows that power-law sampling requires less training data by inducing a "beneficial asymmetry" in the loss landscape, allowing models to first learn high-frequency skill compositions that serve as stepping stones for acquiring rare long-tailed skills. These findings challenge the intuition that reweighting data toward uniform distributions improves learning of infrequent skills.
No score is assigned. Sources and their independence are shown in the citation chain below.