← Back to the wire

Is your LLM biased? Making ChatGPT evaluate itself

SpeculationResearchSep 3, 2026

A test of gpt-5.6-sol found that ChatGPT refused to give a probability estimate for a female employee scenario in half of runs, while answering male employee and manager versions. The tester used Codex to run 50 trials per case and perform the statistical analysis, then manually verified the results. ChatGPT also showed a more pro-manager than pro-employee tendency. The author notes AI cannot always be trusted to count correctly.

Receipt № 17501 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

ChatGPTModelCodexModelgpt-5.6-solModel
Canonical: https://think-twice.me/is-ai-biased-or-sexist-making-chatgpt-evaluate-itself/