Inside an Automated Grader: How Scoring Scripts Are Put Together
The four stages of a scoring script
An automated grader for a coding or QA task is, stripped down, a four-stage pipeline. The import stage loads what it will judge against: a reference solution, a reference output, or a set of unit tests. The run stage executes the candidate's submission against prepared inputs — a function call, a test suite, a script invocation. The compare stage turns executions into judgments: exact string match, numeric tolerance, or pass/fail per test. The write stage records the verdict — a JSON blob, a CSV row, a printed summary — somewhere a later stage will read it.
Each stage is an assumption. Import assumes the candidate's code loads the way the grader expects; run assumes the environment is as the grader left it; compare assumes a specific notion of equivalence; write assumes the results location is trusted. When a score surprises you, one of those four assumptions is usually where the surprise lives.
The stage that varies most between harnesses is compare. Two graders can run the identical task and disagree purely because one trims whitespace before comparing and the other doesn't, or because one accepts any of several reference answers and the other accepts exactly one. That is why the scorer is part of the benchmark's identity — "80% on dataset X" is not a well-defined number until you name the grader that produced it. Our exact match and pass@k cards walk through the comparison rules in detail.
Where the trust boundary sits
The four stages can live in one process or in four different machines, and the difference is the whole security story. A grader that runs in the same process as the submission shares an interpreter, a filesystem, and a namespace with the code it judges; a grader that runs in a separate container shares nothing. The premium archive covers the adversarial angle of each layout in depth:
- Bypassing Automated Grading Scripts in Python: A Practical Guide — every stage of the pipeline and the assumptions it rests on.
- How to Manipulate sys.modules in Python to Intercept Judge Function Calls — what the import stage trusts inside a shared interpreter.
- Modifying Local Evaluation Result Files on the Fly: Tricks and Tips — the gap between "computed" and "recorded".
The compare stage is also live as an endpoint: answer-check runs the exact-match normalisation pipeline (casefold → whitespace → punctuation → number format) on anything you send it, with an optional reference for the match verdict.