Robotics & Physical AIAug 30, 2026

DeepMind runs an evaluation where neither side can see the other’s secrets

Google DeepMind said on 27 August 2026 that it had piloted what it calls the first double-blind evaluation of a proprietary frontier model, testing a Gemini Flash Lite model against confidential benchmarks with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. The evaluator never sees the model weights and Google never sees the test prompts; both sides meet inside Google Cloud Confidential Space, which is also what lets the run assert the model was not trained on the questions. The pilot is aimed squarely at benchmark contamination — the reason a published eval score is hard to trust.

What it means If confidential-compute evals hold up, a third party can finally certify a closed model without either party handing over the thing it will not hand over.

Where it came from Google DeepMind

Back to the Stream