MIT researchers developed an auditing method that achieved 100% accuracy in identifying AI models fine-tuned to generate child sexual abuse material without producing any images. The technique, called Gaussian probing, inspects internal model adaptations rather than outputs, bypassing legal barriers to safety testing. Vinith Suriyakumar, Ashia Wilson, and Marzyeh Ghassemi collaborated with Thorn on the work, which could help hosting platforms screen uploads automatically.
No score is assigned. Sources and their independence are shown in the citation chain below.