Secure AI Benchmark Enclave with four professionals, encrypted benchmark servers, and evaluation panels

Piloting the world’s first double-blind AI evaluations

This matters because AI benchmark trust is becoming an infrastructure problem, not just a research one. Double-blind evaluations built on confidential computing could give enterprises stronger assurance that model claims were measured in isolation, without benchmark leakage contaminating safety, capability or procurement decisions.

Continue reading

1 2 3 337