AI at WorkAI Policy Source checked

NIST expands CAISI agreements for pre-deployment frontier AI testing

Before a new model launches, CAISI can test it and share feedback so developers can fix issues or add safeguards before the public relies on it.

Original source ↗
Start here

In everyday words

Before a new model launches, CAISI can test it and share feedback so developers can fix issues or add safeguards before the public relies on it.

Need a meaning?

What you need to know

Who is affected
companies, policy teams, public-sector readers
What changed
On May 5, 2026, NIST said its Center for AI Standards and Innovation (CAISI) signed agreements with Google DeepMind, Microsoft, and xAI to conduct pre-deployment evaluations and targeted research on frontier AI security.
Why it matters
If more leading labs share models with evaluators before release, it can surface safety and security failures earlier, while creating clearer expectations for how “independent evaluations” are run in practice.
What to watch next
Watch what evaluation methods CAISI standardizes, what kinds of risks they publish results on, and whether more AI developers join similar testing programs.
Four useful details
  • Covers pre-deployment evaluations plus post-deployment assessment and research.
  • NIST says CAISI has completed more than 40 evaluations to date.
  • The agreements can involve testing models with reduced safeguards to assess national-security risks.
Your next sip

Continue reading

All latest briefings →
Previous briefing · AI at Work OpenAI updates GPT‑5.5 Instant with fewer hallucinations and new… May 6, 2026 · 2 min Next briefing · AI at Work Google’s ReasoningBank aims to help agents learn from past runs May 5, 2026 · 2 min