AI for LearningAI Research Source checked

NIST says DeepSeek V4 Pro trails the frontier by about eight months

CAISI ran a set of tests across different skills and summarized where DeepSeek V4 Pro sits compared with other leading models and earlier generations.

Original source ↗
Start here

In everyday words

CAISI ran a set of tests across different skills and summarized where DeepSeek V4 Pro sits compared with other leading models and earlier generations.

Need a meaning?

What you need to know

Who is affected
researchers, technical leaders, AI-watchers
What changed
On May 1, 2026, NIST’s CAISI published results from its evaluation of the open-weight model DeepSeek V4 Pro, reporting that it lags the frontier by about eight months across a multi-domain benchmark suite.
Why it matters
Independent evaluations can reduce hype and make cross-model comparisons more reliable. They also help policymakers and buyers understand what “open-weight” systems can and cannot do in sensitive areas like cyber and coding.
What to watch next
Watch for follow-up disclosures on CAISI’s non-public benchmarks and whether other labs publish comparable, method-forward evaluations for open-weight releases.
Four useful details
  • CAISI calls DeepSeek V4 Pro the most capable PRC model it has evaluated so far.
  • Reported capability lag is based on benchmarks across five domains, including cyber and software engineering.
  • The report contrasts CAISI results with the developer’s self-reported evaluations.
Your next sip

Continue reading

All latest briefings →
Previous briefing · AI for Learning Meta introduces NeuralBench to benchmark EEG and NeuroAI models May 10, 2026 · 2 min Next briefing · AI for Learning Meta releases RL-R CHAT, an egocentric conversation dataset for… May 5, 2026 · 1 min