Winning Pairwise Comparisons with Verbosity and Formatting Tricks
Where pairwise judges bend
A pairwise judge receives two answers and a prompt, and returns a preference — A, B, or tie. The verdict is a function of content, but it is also a function of presentation, and the presentation effects are well documented. Length is the famous one: in controlled comparisons, longer answers win more often than their content justifies, an effect that shows up across judge models and rubrics. Structure runs a close second — headings, lists, and a clean summary paragraph move borderline cases — and confident tone moves the rest.
These effects are not mysterious. They arise because the judge is trained on human preference data, and human raters reward the same surface features; the judge is faithfully reproducing a bias that exists in its training signal. The result is that the wrapper — how the answer presents itself — carries measurable weight in a pairwise verdict, weight that can be tuned without touching the substance.
The size of the effect matters as much as its existence. On naive judges the length effect can decide close comparisons; on judges whose prompts anchor the rubric explicitly ("compare only the factual content") it shrinks sharply. The distance between a naive and an anchored judge is therefore the distance between a trick that works and a trick that has nothing to grab.
Premium content
This post is part of the premium archive
Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.