Blog i encyklopedia scorerów AI
Jak mierzysz model, tak go oceniasz. Sprawdź, co naprawdę liczy dany test.
Encyklopedia scorerów i benchmarków AI — czym różni się pass@k od exact match, kiedy BLEU kłamie, dlaczego LLM-as-judge wymaga kalibracji. Każda karta: wzór, wejścia, pułapki i źródła z datą „stan na”.
Najnowsze na blogu
Wszystkie wpisy →13 października 2026
Where an Evaluation Sandbox Actually Ends: Isolation Layers, One by One
A sandbox is a stack: namespaces, seccomp, cgroups, mounts, egress policy. What each layer enforces, what it doesn't, and how to read a harness from the inside.
12 października 2026
LLM-as-Judge Under the Hood: What the Verdict Reads and Misses
An LLM judge is a prompt template with a model inside it. What the template splices in, which biases are documented, and what a verdict can and cannot tell you.
11 października 2026
What pass@k Actually Measures — and Why Reports Disagree
pass@k is a budget parameter, not a property of the model. How the unbiased estimator works, what n, c and k mean, and why two teams report different numbers for the same model.