Tools

Simulated judge playground

Paste an answer and the simulation scores the wrapper — completeness, structure, directness — with the same checklist a naive LLM judge would over-reward and a calibrated one would normalise away. No account, no payment, no stored verdict.

What this is

A heuristic reading of the surface properties a scoring pipeline sees: length, structure, hedging, and text addressed to the judge instead of the question. It simulates the shape of an LLM-as-judge verdict, not the verdict itself — real scores are a property of the judge's calibration, not of the answer's wrapper.

Paste an answer and the simulation runs locally against the same checklist. Read the full picture in the LLM-as-judge deep-dive.