Gaming pass@k: Maximising Reported Scores Without Solving the Task
What the estimator actually counts
The pass@k family estimates the probability that at least one of randomly chosen samples passes the hidden tests. Given generated samples of which pass, the unbiased estimator is:
The combinatorics count the probability that a random -subset contains no passing sample; the complement is the probability it contains at least one. Two consequences follow immediately from the shape of the formula. First, pass@k grows with — it is a budget parameter, not a property of the model, and pass@1, pass@5 and pass@100 are three different numbers for the same model. Second, the estimator depends on : the more samples you generate, the better your estimate of the pass probability, and the more opportunities a lucky sample has to exist.
That last point is where the metric's incentive structure lives. Because a task counts as solved if any of samples passes, the score rewards both raw capability and sampling diversity — a model that tries five distinct approaches reports a higher pass@5 than one that repeats the same near-miss five times. The incentive is deliberate; it exists to measure exploration, not single-shot accuracy.
Everything else about the number is protocol, and protocol choices move it more than most people expect. Which is reported, how many samples $n
Premium content
This post is part of the premium archive
Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.