A paper on boundary-aware safety training found that tuning Qwen3-8B on political refusal data raised in-distribution refusal from 9.47% to 84.75%, while over-refusal on XSTest rose from 2.00% to 74.00%. The authors argue topic-level safety is too blunt, propose harmful-benign prompt pairs to measure refusal boundaries, and report unsafe-response rates scored by LlamaGuard-3 falling from 26.26% to 0.14% across three benchmarks.
No score is assigned. Sources and their independence are shown in the citation chain below.