This system prevents benchmark contamination—a problem where an AI accidentally learns from test questions during training—by keeping the evaluation data hidden from the model and the model's inner workings hidden from the evaluators.
This approach addresses a critical trust gap in the AI industry where high test scores may result from a model "peeking" at exam questions rather than actual capability.
By moving beyond traditional legal contracts to technical safeguards, Google aims to provide policymakers and enterprises with verifiable proof that a model’s safety and performance benchmarks are accurate.
This is particularly vital for highly sensitive sectors like cybersecurity, where inflated or inaccurate reliability metrics could lead to significant infrastructure risks.
The technical mechanism relies on Confidential Space, a tool within Google Cloud’s confidential computing portfolio that allows two parties to share data for processing without either side seeing the other's proprietary information.
This removes the historic trade-off where evaluators had to share their secret prompts or model owners had to hand over their intellectual property, known as model weights.
This pilot is intended to establish a new standard for independent oversight, allowing government bodies and safety institutes to rigorously stress-test advanced AI without compromising data security.