Escaping the Grader's Subprocess: Breaking Out of the Harness
What a harness boundary is made of
A modern evaluation harness treats candidate code as what it is — untrusted code — and builds a boundary accordingly. The candidate runs in a subprocess, often in a container: a PID namespace hides the host's processes, a mount namespace gives it a private filesystem view, a network policy limits its egress, and resource limits — rlimits, cgroup quotas, timeouts — cap what it can consume. Seccomp filters or a syscall interception layer restrict what the process can ask the kernel to do, which is a stronger statement than any filesystem rule, because it applies to the syscall before the kernel acts on it.
The grader, meanwhile, lives on the other side of that boundary: a separate process, sometimes a separate container, with its own filesystem view and its own credentials. The score travels from the grader's side to the operator's side, never through the candidate's reachable state. That separation is the entire product: the harness is valuable exactly to the extent that the boundary holds.
Boundaries hold by design, not by luck, and their components are ordinary, documented isolation mechanisms — namespaces, seccomp, cgroups, MAC policy. Anyone studying the harness is studying the same building blocks that containers and sandboxes have used for a decade, each with known strengths and known failure modes.
Premium content
This post is part of the premium archive
Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.