SLATE achieves a 7.0% relative improvement over Search-R1 on a 7B model and 30.7% on a 3B model across seven QA benchmarks. Chris Samarinas presents the method, which addresses credit assignment in reinforcement learning for retrieval-augmented language models through truncated step-level sampling that provably reduces advantage estimate variance by up to a factor of T, plus dense decomposed process rewards separately evaluating reasoning, query, and answer quality via an LLM judge.
No score is assigned. Sources and their independence are shown in the citation chain below.