AI at WorkAI Tools Source checked

Checking which machine-read rules still hold

Some software reads statutes first and turns them into “if X, then Y” rules. Different software can disagree, so some rules may be unsafe to trust. This paper proposes a check: only accept a rule when it still holds after many simulated “reading mistakes.” The authors warn it works best when calibrated to each chapter, or made tolerant to errors.

Original source ↗
Start here

In everyday words

Some software reads statutes first and turns them into “if X, then Y” rules. Different software can disagree, so some rules may be unsafe to trust. This paper proposes a check: only accept a rule when it still holds after many simulated “reading mistakes.” The authors warn it works best when calibrated to each chapter, or made tolerant to errors.

Need a meaning?

What you need to know

Who is affected
Compliance and legal operations teams using automated statute reading tools, Policy teams summarizing or mapping statutory requirements at scale, Risk and audit teams reviewing machine-produced legal logic, Vendors building software that converts statutes into structured rules
What changed
Researchers studied how two independently written statute “extractors” disagree when turning laws into structured facts. On Missouri statutes, they diverged on whether numeric thresholds exist, with a reported false-negative rate of 0.43. They propose a “survival certificate” that only approves a rule if it keeps holding under simulated extractor disagreement, and attach evidence spans and a minimal counterexample.
Why it matters
If your workplace relies on tools that pre-read laws, small extraction errors can flip a compliance rule. This work offers a way to label which extracted rules are likely to survive known extraction noise. It also shows the approach can become uninformative when a single, global error setting is reused across different chapters.
What to watch next
Whether teams can practically calibrate error rates per chapter, and how often “certified” rules remain useful across new statute sets.
Four useful details
  • Two statute-reading tools disagreed on numeric thresholds; reported false-negative rate was 0.43.
  • A “survival certificate” only approves rules that survive 1,000 simulated disagreement trials at a strict lower-bound threshold.
  • A global error setting made most held-out chapters fall below an “informativeness” floor; per-chapter calibration was advised.
Your next sip

Continue reading

All latest briefings →
Next briefing · AI at Work Astra meets a critical cyber capability bar Sep 2, 2026 · 1 min