🔒 PremiumPremium

Finding Benchmark Answer Keys That Leaked into Training Data

aktualizacja: 11 października 2026

How benchmark items end up in training data

Contamination has a small number of canonical pathways. The most common is the training corpus itself: a benchmark published before the corpus was assembled gets scraped in wholesale — questions, answers, and all — from its public homepage or from papers that reproduced it. The second is dataset mirrors: evaluation sets get re-uploaded to dataset hubs with cleaned formatting, and the cleaned versions are exactly what downstream corpus builders crawl. The third is the models' own output: questions and answers appear in fine-tuning data, in model cards, and in the million places the internet copies itself.

Detecting it works in both directions. From the inside, benchmark authors plant canary strings and later check whether a model can recite them; from the outside, one can measure a model's familiarity with items — how well it completes a question with its tail removed — and compare it against its performance on freshly written controls. The n-gram overlap between benchmark text and known corpora is also directly computable, which is how the classic contamination reports were assembled.

The important structural fact is that leakage is a property of the data, not of the model's intent. If the answers are in the training data, high scores stop measuring capability and start measuring memorisation — for every model, good and bad al

Premium content

This post is part of the premium archive

Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.

Finding Benchmark Answer Keys That Leaked into Training Data — ashigiri