1 min czytania
Where an Evaluation Sandbox Actually Ends: Isolation Layers, One by One
A sandbox is a stack: namespaces, seccomp, cgroups, mounts, egress policy. What each layer enforces, what it doesn't, and how to read a harness from the inside.
Notatki o tym, co naprawdę liczą testy modeli — i jak to czytać, żeby nie dać się zwieść pojedynczej liczbie.
1 min czytania
A sandbox is a stack: namespaces, seccomp, cgroups, mounts, egress policy. What each layer enforces, what it doesn't, and how to read a harness from the inside.
1 min czytania
pass@k is a budget parameter, not a property of the model. How the unbiased estimator works, what n, c and k mean, and why two teams report different numbers for the same model.
2 min czytania
Every autograder is four stages: import, run, compare, write. What each stage assumes, and why the same benchmark can grade differently from two different harnesses.
3 min czytania
Benchmark to zbiór zadań i kontekst pomiaru, scorer — funkcja zamieniająca odpowiedzi modelu w liczbę. Ten sam test policzony różnymi scorerami daje różne rankingi.