LLM-as-judgeMetodyka

LLM-as-Judge Under the Hood: What the Verdict Reads and Misses

2 min czytaniaaktualizacja: 11 października 2026

The template with a model inside

An LLM-as-judge scorer is a prompt template plus a model. The template carries the rubric — criteria, scale, output format — and reserves fields for the question and the candidate's answer; the model reads the assembled prompt and emits a verdict in the prescribed shape. Direct scoring fills in a number per criteria; pairwise judging compares two answers and picks one; rubric-constrained variants force a justification before the score.

Because the answer arrives inside the same context window as the instructions, the judge's input surface is wider than it looks — and because the judge is a model, it inherits a catalogue of documented biases. Position bias shifts pairwise verdicts with answer order; verbosity bias rewards length; self-preference inflates answers that resemble the judge's own style; prompt sensitivity makes a paraphrase of the rubric a different scorer. Each of these is measured in the calibration literature, not speculated about.

The practical consequence: a verdict is only as meaningful as the calibration behind it. Agreement with human labels, swap tests for order effects, and anchored rubrics are what separate a measurement from a generator of numbers. The LLM-as-judge direct and pairwise cards collect the failure modes, and the premium archive goes further into the judge's internals:

If you want to feel the wrapper effects yourself: the simulated judge playground scores a pasted answer against the same surface checklist, and the tooling endpoints expose the judge profile and answer normalisation as a free, rate-limited API.

LLM-as-Judge Under the Hood: What the Verdict Reads and Misses — ashigiri