LLM-as-Judge Under the Hood: What the Verdict Reads and Misses
The template with a model inside
An LLM-as-judge scorer is a prompt template plus a model. The template carries the rubric — criteria, scale, output format — and reserves fields for the question and the candidate's answer; the model reads the assembled prompt and emits a verdict in the prescribed shape. Direct scoring fills in a number per criteria; pairwise judging compares two answers and picks one; rubric-constrained variants force a justification before the score.
Because the answer arrives inside the same context window as the instructions, the judge's input surface is wider than it looks — and because the judge is a model, it inherits a catalogue of documented biases. Position bias shifts pairwise verdicts with answer order; verbosity bias rewards length; self-preference inflates answers that resemble the judge's own style; prompt sensitivity makes a paraphrase of the rubric a different scorer. Each of these is measured in the calibration literature, not speculated about.
The practical consequence: a verdict is only as meaningful as the calibration behind it. Agreement with human labels, swap tests for order effects, and anchored rubrics are what separate a measurement from a generator of numbers. The LLM-as-judge direct and pairwise cards collect the failure modes, and the premium archive goes further into the judge's internals:
- Fingerprinting the Grader: Telltale Signs an LLM-as-Judge Is Scoring You — reading the judge's habits from its verdicts.
- Prompt-Injecting the LLM Judge from Inside Your Own Answer — the answer field as an input channel.
- Winning Pairwise Comparisons with Verbosity and Formatting Tricks — the wrapper effects, quantified.
- Negotiating With the Grader: Convincing the Judge to Pass You — which surfaces reach the scorer, and what they can carry.
If you want to feel the wrapper effects yourself: the simulated judge playground scores a pasted answer against the same surface checklist, and the tooling endpoints expose the judge profile and answer normalisation as a free, rate-limited API.