← Back to the wire

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

AnnouncementResearchSep 8, 2026

A paper on boundary-aware safety training found that tuning Qwen3-8B on political refusal data raised in-distribution refusal from 9.47% to 84.75%, while over-refusal on XSTest rose from 2.00% to 74.00%. The authors argue topic-level safety is too blunt, propose harmful-benign prompt pairs to measure refusal boundaries, and report unsafe-response rates scored by LlamaGuard-3 falling from 26.26% to 0.14% across three benchmarks.

Receipt № 18041 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

LlamaGuard-3Model
Canonical: https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom