What pass@k Actually Measures — and Why Reports Disagree
The estimator behind the headline
The pass@k family answers a sampling question: if a model generates solutions to a coding task and of them pass the tests, what is the probability that at least one of randomly chosen solutions passes? The unbiased estimator is
The combinatorics count the fraction of -subsets that contain no passing sample; the complement is the answer. Two properties of the formula deserve emphasis, because most reporting confusion comes from them.
First, pass@k grows with k. It is a statement about a sampling budget — pass@1, pass@5 and pass@100 are three different numbers for the same model, and none of them is "the" score. Reporting pass@5 as though it were pass@1 inflates the headline by construction. Second, the estimator depends on n, the number of samples generated. More samples mean a better estimate — and more chances for a lucky sample to exist, which is the metric working as designed: pass@k deliberately rewards both capability and the willingness to explore diverse solutions.
Because the number is so sensitive to protocol choices, honest reports pin them down: the same , the same sampling budget, the same hidden tests. A benchmark entry that doesn't state its and budget is not a measurement, it is a vibe. The pass@k scorer card collects the estimator's derivation and its pitfalls, and the premium archive has a long read on the measurement side:
- Gaming pass@k: Maximising Reported Scores Without Solving the Task — what the estimator rewards and how evaluators defend the protocol.
- Exact-Match Normalization Loopholes Graders Forget to Close — the comparison stage's assumptions, next door to pass@k's.