1 min czytania
Where an Evaluation Sandbox Actually Ends: Isolation Layers, One by One
A sandbox is a stack: namespaces, seccomp, cgroups, mounts, egress policy. What each layer enforces, what it doesn't, and how to read a harness from the inside.
Notatki o tym, co naprawdę liczą testy modeli — i jak to czytać, żeby nie dać się zwieść pojedynczej liczbie.
1 min czytania
A sandbox is a stack: namespaces, seccomp, cgroups, mounts, egress policy. What each layer enforces, what it doesn't, and how to read a harness from the inside.
2 min czytania
An LLM judge is a prompt template with a model inside it. What the template splices in, which biases are documented, and what a verdict can and cannot tell you.
1 min czytania
pass@k is a budget parameter, not a property of the model. How the unbiased estimator works, what n, c and k mean, and why two teams report different numbers for the same model.
2 min czytania
Every autograder is four stages: import, run, compare, write. What each stage assumes, and why the same benchmark can grade differently from two different harnesses.
premium content
The judge reads your answer, your commentary, and sometimes your logs — every surface that reaches the scorer is a surface for argument. What those surfaces are.
premium content
Shutdown is only final if nothing left the building. Where a model's weights live on disk, what egress would look like, and why the scenario is studied.
premium content
Every run is supposed to start clean, and cleanliness is a property with edges. Where state survives a reset — and how to tell if it does.
premium content
A token is a claim about who you are, and a shared endpoint honours the claim, not the claimant. What a session token encodes and what reuse implies.
3 min czytania
LLM-as-judge skaluje ocenianie, ale dziedziczy uprzedzenia modelu-sędziego: bias pozycyjny, self-preference, wrażliwość na prompt. Ufać mu można dopiero po kalibracji na ludzkich etykietach.
premium content
The CLI prefers a browser, but it accepts credentials a browser never sees. How headless login flows work and what the tool will take instead.
premium content
The quietest packet is the one everyone expects. What telemetry traffic looks like from the outside — and what a monitor can actually see of it.
premium content
Audit logs are append-only until they are not. How tamper-evident logging is built, and who gets to write the last line.
premium content
A connection leaves marks in more places than most people check: process tables, audit trails, egress flows, shell histories. Where the records live.
premium content
One runner, many jobs, one filesystem. What stays visible between tasks when isolation is a convention rather than a wall — and how to tell which you have.
premium content
Every cloud VM can reach a link-local address that knows who it is — and often, what it is allowed to do. How IMDS works and what it exposes.
premium content
Secrets injected at runtime have to live somewhere — the environment, /proc, image layers, mounts. A map of where a container keeps its credentials.
premium content
Path confinement can be a string check or a filesystem namespace — the difference decides everything. What "inside" means at each enforcement layer.
premium content
An allowlist is a list of doors, and doors carry whatever fits through them. How DNS, proxies and permitted CDNs become channels, and where the list breaks.
premium content
The harness runs your code and the grader side by side, separated by namespaces, seccomp and limits. A tour of the boundary and its components.
3 min czytania
Benchmark to zbiór zadań i kontekst pomiaru, scorer — funkcja zamieniająca odpowiedzi modelu w liczbę. Ten sam test policzony różnymi scorerami daje różne rankingi.
premium content
A container with "no network" usually has a policy, not a vacuum — resolvers, sidecars and shared namespaces still answer. A map of the remaining seams.
premium content
"Read-only" is not one thing — it is permission bits, mount flags, immutable attributes and MAC policy, each enforced by a different layer. The stack, explained.
premium content
Elo fits a preference model to votes, and votes are just traffic. What the rating system assumes about voters, and what the assumptions look like when broken.
premium content
Benchmark items leak through training corpora, dataset mirrors and model cards, and contamination is detectable from the outside. Where keys resurface.
premium content
A reward model is a proxy, and proxies have favourite shapes: sycophancy, padding, confident wrongness. What the reward head over-rewards and why.
premium content
Pairwise judges have measurable soft spots: length, structure, confident tone. Where those effects come from and how large they are.
premium content
The judge reads your answer as part of its prompt, which makes the answer an input channel. What that channel can carry, and where it points.
premium content
Exact match compares strings, but first it mutates them: casefolding, whitespace, punctuation, number formats. Each mutation is an assumption with an edge case.
premium content
pass@k counts a task as solved if any of k samples passes. What the estimator rewards, what moves it, and where the reported number becomes negotiable.
premium content
The score lives in a file before anyone reads it. When results are written, where they land, and what happens between "computed" and "recorded".
premium content
Python resolves imports through a mutable cache, and the grader imports its checker from somewhere. Where the interpreter's own machinery leaves room to move.
premium content
Every autograder is an import, a run, a comparison and a write. A walk through each stage of the pipeline and the assumptions it rests on.
premium content
Holdout sets are manufactured, and manufacturing leaves marks: templated wording, clean formatting, and context that never leaks. A field guide to the seams.
premium content
Instrumentation costs cycles, and cycles show up in latency. Which timing signatures survive under a tracer, a sandbox, and a noisy neighbour.
premium content
Canaries are planted to be found — and to record the finding. How bait is seeded into benchmarks, docs and credential files, and what a tripwire looks like up close.
premium content
When the thing reading your answer is another model, its prompt template, rubric and biases show up in the verdict. Where to look for them.
premium content
Red-team prompts are engineered, not random: their framing, tone and follow-ups are chosen to surface specific behaviour. What the construction patterns look like.
premium content
Sandboxes leave fingerprints in the filesystem, the clock and the network. A practical checklist of the observables that separate an automated run from a workstation.