An independent Redactle benchmark of 22 model configurations found Gemini 3.7 Flash solved all 12 puzzles with a 100% solve rate at $0.005 per run. Gemini 3.8 Flash matched that result across low, medium, and high reasoning settings. Grok 4.6 also achieved 100% but at higher cost and slower times. Gemini and Grok outperformed OpenAI and Anthropic models; reasoning effort showed limited benefit. The evaluation covered 264 of 288 planned attempts.
No score is assigned. Sources and their independence are shown in the citation chain below.