In everyday words
Before a new model launches, CAISI can test it and share feedback so developers can fix issues or add safeguards before the public relies on it.
Need a meaning?
Testing a model before it is publicly released to understand capabilities and risks.NIST’s Center for AI Standards and Innovation, focused on AI measurement and security research.
Quick Sip
What you need to know
- Who is affected
- companies, policy teams, public-sector readers
- What changed
- On May 5, 2026, NIST said its Center for AI Standards and Innovation (CAISI) signed agreements with Google DeepMind, Microsoft, and xAI to conduct pre-deployment evaluations and targeted research on frontier AI security.
- Why it matters
- If more leading labs share models with evaluators before release, it can surface safety and security failures earlier, while creating clearer expectations for how “independent evaluations” are run in practice.
- What to watch next
- Watch what evaluation methods CAISI standardizes, what kinds of risks they publish results on, and whether more AI developers join similar testing programs.
Four useful details
- Covers pre-deployment evaluations plus post-deployment assessment and research.
- NIST says CAISI has completed more than 40 evaluations to date.
- The agreements can involve testing models with reduced safeguards to assess national-security risks.
OpenAI · Official AnnouncementOffering Zero Data Retention for frontier models ↗
Adds source-backed context on ai safety from OpenAI.
OpenAI · Official UpdateAdvancing content provenance for a safer, more transparent AI ecosystem ↗Adds source-backed context on ai news from OpenAI.
Notion · Official AnnouncementIntroducing Notion’s Developer Platform ↗Adds source-backed context on ai tools from Notion.
Your next sip
All latest briefings →