The Oversight Board tested 10 commercial LLMs from Anthropic, DeepSeek, Google, Meta, and OpenAI through interfaces provided by Google and Microsoft, finding models were more than twice as likely to refuse criticizing restrictive political regimes. Refusal rates averaged 34% for restrictive jurisdictions versus 14% for permissive ones. Models cited inconsistent or nonexistent policies, and sometimes referenced local laws, potentially misleading users about the actual reasons behind their behavior.
No score is assigned. Sources and their independence are shown in the citation chain below.