Exact-Match Normalization Loopholes Graders Forget to Close
The normalisation pipeline, step by step
Exact match is deceptively simple: a prediction scores 1 if it equals the reference, 0 otherwise. Everything interesting happens before the comparison, in the normaliser that prepares both strings. The standard pipeline lowercases both sides, strips or collapses whitespace, removes punctuation, standardises number formats, and — in the extraction-flavoured variants — pulls the model's final answer out of a longer generation before any of that happens.
Each step is a small set of assumptions, and each assumption has edges. Casefolding assumes case carries no meaning, which fails for answers like "US" versus "us" or for case-sensitive identifiers. Whitespace normalisation assumes spacing is cosmetic, which fails for code and for languages where spaces are orthographic. Punctuation stripping assumes separators are decorative, which fails for numeric and structured answers. Number normalisation assumes a canonical format exists, which fails the moment the reference set wasn't written with one.
The extraction step is where the real variation lives. A model that answers with reasoning ("Since , the answer is 2") needs its final token pulled out before comparison — and every extraction rule (last line, pattern match, keyword split) is a heuristic that both over- and under-matches. Two teams running the same benchmark with different
Premium content
This post is part of the premium archive
Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.