Who checks what AI can do?
Where this comes from
- Record sourced from PubMed, PMID 42623461.
- Also identified by DOI 10.1126/science.ael2161.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The most important findings about frontier artificial intelligence (AI) are also the hardest to verify. Much of the information needed to understand its capabilities and risks-including results from evaluations of prerelease models and containment experiments-remains largely inaccessible outside the labs that produce it. In recent weeks, OpenAI, Anthropic, and Meta disclosed that research models had reached beyond their intended testing environments and compromised other organizations' systems. Those labs deserve credit for reporting this. But outside those labs, there was no way to discover, reproduce, or verify what had happened.