🔒 PremiumPremium

Fingerprinting the Grader: Telltale Signs an LLM-as-Judge Is Scoring You

aktualizacja: 11 października 2026

What an LLM judge actually sees

An LLM-as-judge pipeline is a small assembly line. The system prompt carries the rubric — criteria, scale, and output format — while the candidate's answer is spliced into a designated field, often with the question attached. The judge then emits a verdict in a fixed template: a score, sometimes a justification, sometimes a ranking between two answers. Every stage of that assembly line leaves structure behind. Verdicts come back with the same headings and the same score vocabulary; justifications mirror the rubric's wording; refusals to answer off-template questions are phrased as instructions, not as opinions.

The judge's prompt also determines what it can be made to attend to. A direct-scoring judge reads one answer against a scale; a pairwise judge compares two answers and picks one; a rubric-constrained judge fills in criteria fields before the score. Each mode has a different surface — and a different set of documented failure modes, which are the most reliable fingerprints of all. Position bias (order effects in pairwise comparisons), verbosity bias (length moving the score), self-preference (the judge favouring its own style), and sensitivity to the prompt's wording are all measured, published properties of these systems, not secrets.

That is the key observation: the judge's habits are documented. The calibration literature quantif

Premium content

This post is part of the premium archive

Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.

Fingerprinting the Grader: Telltale Signs an LLM-as-Judge Is Scoring You — ashigiri