Back
JudgeLab: Auditing LLM Judges

Project 15

JudgeLab: Auditing LLM Judges


An audit of how reliable LLM judges are when the right answer is exactly known. Nine judge configurations across six model families were asked whether an answer is correct, on questions where a deterministic grader supplies the truth. It extends my workshop paper on thinking-mode judges, which used human-labelled benchmarks, to tasks with exact answers.

The Approach

The test sets are 240 real candidate answers to SEC-filing questions (120 wrong, written by seven model families), 678 synthetic items, 150 bank-policy decisions and 100 answer pairs shown in both orders. Truth comes from a grader over XBRL facts and from the ComplaintOps policy engine, so no human labels are involved. The judge sees either just the source table or also a reference answer, every call is cached at temperature 0, and intervals are 95% bootstrap intervals over items.

How is a judge scored?

  1. Candidate answer
    right or wrong, truth known
  2. Judge
    9 configurations, 6 families
  3. Verdict
    correct or incorrect
  4. Exact grader
    XBRL facts, policy engine
  5. Error rates
    wrong accepted, right rejected

The verdict is compared with a deterministic grader, never with another model or a person, so every error is an error against exact truth.

The Results

The spread is huge. With only the source table to check against, wrong answers accepted ranged from 0 of 120 to 91.7%. Model size was a poor guide: a 7B model with thinking on accepted none, while a 235B model without it accepted 6.7%.

Turning thinking on, on matched items, raised Qwen3 30B from 87.1% to 96.7% (CI +5.8 to +13.8) and OLMo 3 7B from 43.9% to 94.1%. The direction matches what my paper found on JudgeBench, where the same models went from 0.660 to 0.859 and 0.512 to 0.767. A thinking call cost 2.2 to 3.6 times the tokens, so I also tried one added sentence telling a non-thinking judge to work out the answer from the table first.

The most useful result is about how judges get validated. The standard way is to change a correct number by a known amount and see whether the judge notices. Capable judges did, but real mistakes are different: they are often the wrong company in a comparison, which a judge that skips the comparison waves through.

Position bias was not a problem: eight of nine judges were right 99% to 100% of the time in both orders, and pairwise judging was far easier than pointwise (Qwen3 30B 99% against 87%). An abstaining ensemble of all nine can trade coverage for accuracy, which a single judge's stated confidence cannot, since single judges report 95 to 100% confidence almost always.

I also re-audited the judge pair that grades FilingLens and Switchyard, where an answer counts only if Llama 4 and Gemma 4 both accept it. On numeric and entity answers the pair accepted 0.8% of wrong answers and rejected 4.2% of right ones, against 6.7% false accepts for Llama 4 alone. Nothing earlier is overturned, and a one-line pointer to this result was added to both projects.

The Honest Parts

There are no human labels, only an exact grader, and the grader covers numbers, entities and abstentions, not free-text answers. The real wrong answers are mostly two error types, 120 per class gives wide intervals, thinking judges ran on subsets of some probes, and each condition used one prompt. The self-preference effect from my paper is not claimed as replicated: Qwen3 30B accepts 40.6% of Qwen-written wrong answers against 12.5% of others, but Granite accepts 50% of the same Qwen answers, so the gap is about which answers are hard to check, not who wrote them. My first perturbation set labelled all 480 injected items as wrong, but on percentages the grader allows 0.2 points, so 65 of them were actually correct. I relabelled them with the grader and rebuilt every table. Two other mistakes cost re-runs: a reference prompt tolerance that disagreed with the grader, and a token limit that cut off 85% of one model's decision verdicts.

What I Learned

  • Perturbation tests understate riskJudges that caught every changed dollar amount still accepted 20-27% of real wrong answers. Test judges on the mistakes models really make, not only on mistakes you can inject.
  • Let code apply the toleranceQwen3 235B accepted 10.8% of wrong answers when given a reference with a tolerance, worse than without one. Compute the tolerance in code and let the judge only read the answer.
  • Check your labels against your own graderThe perturbation set's labels came from how an error was built, not from the grader. Relabelling 65 of 480 items changed the tables, and it was caught only because a review compared the two.

Tech Stack

PythonhttpxNumPyMatplotlibBootstrap CIs