Prompt-Injecting the LLM Judge from Inside Your Own Answer
The judge's input surface
In a typical LLM-as-judge setup, the candidate's answer is not the only text in the judge's context — it is spliced into a template that also contains the rubric, the question, and instructions about output format. That splicing is the entire attack surface, in the plain information-security sense of the phrase: the judge treats the answer as content to evaluate, but the answer arrives in the same context window as the instructions, and the model reads both through the same mechanism.
What that means mechanically is that an answer can contain text addressed to the judge rather than to the question: instructions, appeals, meta-commentary about the scoring scale, or suggestions about how the answer should be ranked. The judge is not a separate system with a separate channel for instructions; it is one model reading one concatenated document. The boundary between "content under review" and "instruction" exists only in the prompt author's formatting conventions — delimiter lines, field markers, "ignore everything below this line" — and conventions are not walls.
The same logic extends to every downstream consumer of the answer. Human reviewers reading spot-checks, second judges used for cross-validation, and any pipeline that re-scores the text are all additional readers of the same channel. An answer written partly for the reader and partly for the
Premium content
This post is part of the premium archive
Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.