AI benchmarks have a trust problem and Google wants to fix it

| Source: THE DECODER

Tags: Google DeepMind, benchmark contamination, Gemini, AI safety, Singapore AI Safety Institute, Confidential Space, AI evaluation

Google DeepMind has launched the first cryptographic double-blind evaluation of a proprietary frontier AI model, using Google Cloud's Confidential Space so neither the evaluator sees model weights nor Google sees test questions — directly attacking AI benchmark contamination.

Details

Benchmark contamination — where models train on the very questions they're later tested on — has been an open and growing problem in AI evaluation. High scores become suspect when there's no guarantee the model hasn't seen the questions. Google DeepMind says it has a solution. In a pilot with the Singapore AI Safety Institute, DeepMind ran what it calls the first double-blind evaluation of a proprietary frontier model, testing a Gemini Flash Lite variant against confidential benchmarks. The setup uses Google Cloud's Confidential Space: test prompts stay cryptographically locked from Google, while model weights stay locked from evaluators. Neither party sees what they shouldn't. Previously, external evaluations forced an uncomfortable tradeoff: share test prompts (contamination risk) or share model weights (IP risk). The ARC-AGI benchmark evaluation of Anthropic's Fable 5 illustrated this — Anthropic's 30-day data retention policy delayed the assessment. The new approach theoretically eliminates the dilemma. DeepMind frames cybersecurity and government evaluations as the biggest beneficiaries, where test data sensitivity makes contractual safeguards insufficient. A technical report details the methodology. If adopted more broadly, this could raise the bar for how much practitioners can trust AI benchmark results.