Piloting the world's first double-blind AI evaluations
| Source: Google DeepMind Blog
Tags: Google DeepMind, Gemini, benchmark contamination, AI safety, confidential computing, AI evaluation, Singapore AI Safety Institute
Google DeepMind and partners including Singapore's AI Safety Institute ran the first cryptographically enforced double-blind evaluation of a frontier AI model, using Confidential Computing to prevent both benchmark contamination and IP exposure — a direct technical response to growing distrust of self-reported AI benchmarks.
Details
The AI industry faces a credibility problem with benchmarks: if a model has seen test questions during training, its scores become unreliable. Today, Google DeepMind announced the first technically enforced double-blind evaluation of a proprietary frontier AI model, piloted with a Gemini Flash Lite model. The system uses Google Cloud's Confidential Space, a confidential computing environment that cryptographically ensures neither side can see the other's secrets: the evaluator cannot access Gemini's model weights, and Google cannot see the evaluator's test prompts. This eliminates the historical tradeoff where one party had to simply trust the other. Partners include the Singapore AI Safety Institute, OpenMined (a privacy-preserving ML org), AVERI, and MLCommons. The framework was developed in response to policymaker and enterprise demands for verifiable, trustworthy benchmarks as AI capabilities scale — particularly relevant for AI Safety Institutes globally that require independent evaluation without IP exposure. The approach does not retroactively solve contamination in existing benchmarks, but provides infrastructure for future evaluations that can be trusted by third parties without compromising proprietary weights or test integrity. This matters most for regulatory contexts where independent third-party audit is becoming a compliance requirement.