MetodykaPodstawy

Where an Evaluation Sandbox Actually Ends: Isolation Layers, One by One

1 min czytaniaaktualizacja: 11 października 2026

The stack, from the kernel up

A modern evaluation sandbox is not one mechanism — it is a stack, and each layer answers a different question. The PID namespace decides which processes exist from your point of view; the mount namespace decides which files exist; cgroups and rlimits decide how much CPU, memory and I/O you may consume; seccomp decides which syscalls you may ask the kernel to perform at all; egress policy decides which network destinations exist. Each layer is independent, each has known failure modes, and the sandbox's strength is the strength of its weakest relevant layer.

Two layers do most of the real work in grading contexts. The mount namespace — or a chroot, its older cousin — implements "you can only see your workspace" at the filesystem level, which is categorically stronger than checking path strings in the harness. Seccomp implements "you may only do these syscalls" at the kernel boundary, which is stronger still: a syscall that is filtered never happens, regardless of what the code intended.

Reading a harness from the inside is mostly a matter of enumerating the stack: which namespaces are visible, which syscalls are filtered, what the egress policy permits, where the filesystem view ends. The premium archive walks the layers one by one:

Where an Evaluation Sandbox Actually Ends: Isolation Layers, One by One — ashigiri