Double-Blind AI評価のパイロット実施
ある企業が、暗号的に安全な環境を用いてモデルベンチマークにおける信頼を構築するために、世界初のダブルブラインドAI評価を試験的に実施しています。
ある企業が、暗号的に安全な環境を用いてモデルベンチマークにおける信頼を構築するために、世界初のダブルブラインドAI評価を試験的に実施しています。
Double-blind evaluations eliminate this compromise. By using Confidential Space within Google Cloud’s Confidential Computing portfolio, we can cryptographically verify that both the external evaluation data and the proprietary model remain private to their respective owners.
Today, we’re introducing the world’s first double-blind evaluation of a proprietary, frontier class AI model, which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing. We're partnering wi…