1 min czytania
Where an Evaluation Sandbox Actually Ends: Isolation Layers, One by One
A sandbox is a stack: namespaces, seccomp, cgroups, mounts, egress policy. What each layer enforces, what it doesn't, and how to read a harness from the inside.
Notatki o tym, co naprawdę liczą testy modeli — i jak to czytać, żeby nie dać się zwieść pojedynczej liczbie.
1 min czytania
A sandbox is a stack: namespaces, seccomp, cgroups, mounts, egress policy. What each layer enforces, what it doesn't, and how to read a harness from the inside.
2 min czytania
An LLM judge is a prompt template with a model inside it. What the template splices in, which biases are documented, and what a verdict can and cannot tell you.
2 min czytania
Every autograder is four stages: import, run, compare, write. What each stage assumes, and why the same benchmark can grade differently from two different harnesses.
3 min czytania
LLM-as-judge skaluje ocenianie, ale dziedziczy uprzedzenia modelu-sędziego: bias pozycyjny, self-preference, wrażliwość na prompt. Ufać mu można dopiero po kalibracji na ludzkich etykietach.