AI at WorkAI Tools Source checked

Paper: today’s AI tests miss reasoning

The researchers built a new way to check whether AI can handle truly new problems by recombining what it already knows. On their new test, the same systems that did well on a usual “new questions” split did much worse. When the researchers made the new test easier in certain ways, scores rose back toward usual levels.

Original source ↗
Start here

In everyday words

The researchers built a new way to check whether AI can handle truly new problems by recombining what it already knows. On their new test, the same systems that did well on a usual “new questions” split did much worse. When the researchers made the new test easier in certain ways, scores rose back toward usual levels.

Need a meaning?

What you need to know

Who is affected
Teams choosing AI tools for unpredictable workflows, People who design or buy AI evaluations and scorecards, Researchers building tests meant to reflect real-world problem solving
What changed
A new arXiv paper argues many “systematic generalization” tests use simplifying setups that make evaluation easier but incomplete. The authors introduce TranSGrid, a unified task designed to require deductive, inductive, and abductive reasoning together. Across seven Transformer models on 4,800 TranSGrid problems, results were much worse than on a standard held-out test set.
Why it matters
At work, teams often rely on AI evaluation scores to judge readiness for new, unfamiliar situations. This paper suggests some common tests may overstate real-world flexibility because they simplify the kind of reasoning required. It also suggests that changing just one simplifying assumption can make a hard test look “normal,” which can mislead selection decisions.
What to watch next
Whether other groups reproduce TranSGrid-style results, and whether future evaluations keep inductive and abductive demands instead of simplifying them away.
Four useful details
  • New TranSGrid test combines deductive, inductive, and abductive reasoning in one task.
  • Seven Transformers scored far lower on TranSGrid than on a held-out test split.
  • Making actions or goals simpler pushed scores back toward “normal” levels.
Your next sip

Continue reading

All latest briefings →
Next briefing · AI at Work OpenAI process for reporting misalignment Sep 17, 2026 · 1 min