AI at WorkAI Agents Source checked

When AI code judges don’t have a basis

If one AI is asked to decide which of two code answers is correct, it may sound confident even when it has no real proof. The paper shows a multi-step “check each claim” approach can still fail for code, because it may not get different evidence for each option. A practical fix is for the judge to sometimes say, “I can’t tell from the evidence,” based on signals...

Original source ↗
Start here

In everyday words

If one AI is asked to decide which of two code answers is correct, it may sound confident even when it has no real proof. The paper shows a multi-step “check each claim” approach can still fail for code, because it may not get different evidence for each option. A practical fix is for the judge to sometimes say, “I can’t tell from the evidence,” based on signals...

Need a meaning?

What you need to know

Who is affected
Software teams using AI to compare or review code, Managers relying on AI-generated code review decisions, Teams building workflows where AI checks other AI outputs
What changed
Researchers studied when an AI system that judges code is actually supported by evidence. They argue the evidence must be independent from the code being judged and must differ between the two code options. In code judging, that second condition can fail. Testing a published multi-step judge (MARCH) on two code-judging comparisons, they found it often rated both options equally good.
Why it matters
At work, teams may use AI to review or compare code changes. This paper suggests some multi-step “verification” methods can still produce confident-sounding judgments without a real basis. The authors’ main contribution is a way to detect, without extra labels, when the judge lacks support and should refuse to decide. That can reduce misplaced confidence in automated code review.
What to watch next
Whether teams adopting AI-based code review add a “refuse to decide without evidence” step, especially for head-to-head comparisons.
Four useful details
  • In tests, the multi-step judge often said both code options were equally good.
  • A log-based check let the judge skip comparisons it couldn’t support.
  • The paper’s focus is detecting “no basis,” not claiming a better judge overall.
Your next sip

Continue reading

All latest briefings →
Previous briefing · AI for Learning Hugging Face posts Holo4 update Sep 28, 2026 · 53 sec Next briefing · AI at Work Secure server memory for Private AI Compute Sep 27, 2026 · 59 sec