Google has announced the first double-blind evaluation of a proprietary frontier AI model, testing Gemini Flash Lite against confidential benchmarks in a privacy-preserving environment. Partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, the pilot uses Confidential Space in Google Cloud's Confidential Computing portfolio so evaluators cannot see model weights and Google cannot see test prompts. The cryptographic approach aims to prevent benchmark contamination, where models inflate scores by seeing questions in advance.
No score is assigned. Sources and their independence are shown in the citation chain below.