A test of gpt-5.6-sol found that ChatGPT refused to give a probability estimate for a female employee scenario in half of runs, while answering male employee and manager versions. The tester used Codex to run 50 trials per case and perform the statistical analysis, then manually verified the results. ChatGPT also showed a more pro-manager than pro-employee tendency. The author notes AI cannot always be trusted to count correctly.
No score is assigned. Sources and their independence are shown in the citation chain below.